Safety helmet detection system based on multi-scale image generation and analysis
The safety helmet detection system, which utilizes multi-scale image generation and analysis, addresses the robustness and accuracy issues in safety helmet detection under complex backgrounds. It achieves high-precision detection of safety helmets across multiple scales, thereby enhancing the robustness and engineering practicality of the detection system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-31
AI Technical Summary
Existing safety helmet detection technologies are susceptible to multi-scale changes, background interference, and weak feature factors in complex environments, making it difficult to achieve high-precision and high-robust detection.
A safety helmet detection system based on multi-scale image generation and analysis is adopted, including an image layering module, a difference modeling module, a consistency evaluation module, an interference judgment module, and a target detection module. Through multi-scale feature decomposition, key point position difference analysis, feature consistency evaluation, and background interference elimination, the robustness and accuracy of detection are improved.
The system achieves high robustness and high accuracy in detecting multi-scale safety helmets in complex scenarios, overcoming the effects of factors such as changes in lighting, occlusion, and cluttered backgrounds, thereby enhancing the model's generalization ability and engineering practicality.
Smart Images

Figure CN121767928A_ABST
Abstract
Description
Technical Field
[0002] This invention relates to the field of safety helmet detection technology, and in particular to a safety helmet detection system based on multi-scale image generation and analysis. Background Technology
[0004] With increasingly stringent safety management requirements in high-risk work environments such as construction and industrial production, automatic helmet detection technology has become a crucial component of intelligent monitoring systems. Traditional helmet detection methods are mostly based on general object detection frameworks, such as YOLO and SSD deep learning models. While these methods offer high detection accuracy in specific scenarios, they are susceptible to factors like lighting variations, occlusion, interference from similarly colored objects, and scale diversity in complex backgrounds, leading to frequent false positives and false negatives. Especially in multi-scale scenarios, the target features in small helmets or images taken at long distances are weak and difficult to distinguish from background noise, severely impacting detection robustness.
[0005] In recent years, some studies have attempted to introduce multi-scale analysis mechanisms to enhance feature representation capabilities. However, most methods only perform scale adaptation through image pyramids or feature pyramids, lacking refined modeling of key spatial location and intensity features at different scales, and failing to effectively suppress the misleading effect of background interference on key feature matching. Furthermore, existing methods generally neglect the dynamic characteristics of structural consistency and positional deviation between the standard template and the target object, limiting further improvements in feature discrimination capabilities.
[0006] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0008] The main objective of this invention is to provide a safety helmet detection system based on multi-scale image generation and analysis, which aims to solve the technical problem that existing safety helmet detection technologies are easily affected by multi-scale changes, background interference, and weak features in complex backgrounds, making it difficult to achieve high-precision and high-robust detection.
[0009] To achieve the above objectives, the present invention provides a safety helmet detection system based on multi-scale image generation and analysis, the system comprising an image layering module, a difference modeling module, a consistency evaluation module, an interference determination module, and a target detection module;
[0010] The image layering module is used to acquire multi-scale images of the scene to be tested and the standard safety helmet image, and denot them as the multi-scale image to be analyzed and the standard multi-scale image; multi-scale feature decomposition is performed on the multi-scale image to be analyzed and the standard multi-scale image to be analyzed, and each feature component to be analyzed and each standard feature component are extracted.
[0011] The difference modeling module is used to analyze the differences between the feature components to be analyzed and the standard feature components regarding the coordinates of the corresponding key points of each feature. It constructs the variation coefficient of the difference between the key point positions of the feature components to be analyzed and the standard feature components, and obtains the background interference factor of the feature components to be analyzed by combining the average level of the difference between the feature components to be analyzed and the standard feature components regarding the coordinates of the corresponding key points of the feature.
[0012] The consistency evaluation module is used to determine the feature consistency of the feature vector to be analyzed based on the fluctuation of the difference between the feature vector to be analyzed and the standard feature vector with respect to the intensity of each feature key point.
[0013] The interference determination module is used to construct the superimposed interference feature value of each feature vector to be analyzed based on the background interference factor and feature consistency of each feature vector to be analyzed. Based on the superimposed interference feature value of all feature components to be analyzed, the effective safety helmet feature vector in the multi-scale image to be analyzed is extracted, and the target feature after background interference is eliminated is obtained through feature fusion.
[0014] The target detection module is used to detect safety helmets in the scene image under test by using a classifier combined with target features after background interference removal, and to obtain the existence and category results of the safety helmets.
[0015] Optionally, the extraction of each feature component to be analyzed and each standard feature component includes:
[0016] Construct the feature vector to be analyzed and the standard feature vector, and use the feature components of the feature vector to be analyzed after multi-scale feature decomposition as the feature components to be analyzed.
[0017] Multi-scale feature decomposition is performed on the standard feature vector, and the discrete coefficients of each feature component of the standard feature vector are calculated. The average of the discrete coefficients of all feature components of the standard feature vector is used as the judgment threshold. Feature components with discrete coefficients greater than or equal to the judgment threshold are all recorded as standard feature components.
[0018] Optionally, the construction of the feature vector to be analyzed and the standard feature vector further includes:
[0019] All feature data in the multi-scale image to be analyzed are arranged in ascending order of scale to form the feature vector to be analyzed, and all feature data in the standard multi-scale image are arranged in ascending order of scale to form the standard feature vector; the feature data includes the coordinates of feature key points, edge intensity values and color channel data at each scale.
[0020] Optionally, the formula for calculating the coefficient of variation of the key point position difference between the feature component to be analyzed and the standard feature component is as follows:
[0021] ;
[0022] In the formula, Let be the coefficient of variation of the keypoint position difference between the i-th eigenvector to be analyzed and the j-th standard eigenvector, norm() be the exponential normalization function, and M be the number of elements in the keypoint position difference vector between the i-th eigenvector to be analyzed and the j-th standard eigenvector. and These are the k-th and (k-1)-th elements in the key point position difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector, respectively.
[0023] Optionally, the construction of the key point location difference vector further includes:
[0024] For each feature vector to be analyzed and each standard feature vector, the coordinates corresponding to the positions of all feature key points in the feature vector to be analyzed and the standard feature vector are arranged according to the key point number to form a vector, and are denoted as the key point coordinate vector of the feature vector to be analyzed and the key point coordinate vector of the standard feature vector.
[0025] The difference vector between the feature vector to be analyzed and the standard feature vector with respect to the key point coordinate vector is denoted as the key point position difference vector between the feature vector to be analyzed and the standard feature vector.
[0026] Optionally, the formula for calculating the background interference factor of the feature component to be analyzed is:
[0027] ;
[0028] In the formula, Let N be the background interference factor for the i-th eigenvector to be analyzed, and N be the number of standard eigenvectors. Let be the mean of the elements in the key point position difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector.
[0029] Optionally, the formula for calculating the feature consistency of the feature vector to be analyzed is:
[0030] ;
[0031] In the formula, Let be the feature consistency of the i-th feature vector to be analyzed, and exp() be an exponential function with the natural constant as the base. Let be the information entropy of the feature intensity difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector. Here, the intensity values of all feature keypoint positions in the i-th feature vector to be analyzed and the j-th standard feature vector are sorted according to the keypoint index to form the keypoint intensity vector of the i-th feature vector to be analyzed and the keypoint intensity vector of the j-th standard feature vector. The difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector with respect to the keypoint intensity vector is denoted as the feature intensity difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector.
[0032] Optionally, the superimposed interference feature value of each feature vector to be analyzed is the ratio of the background interference factor to the feature consistency of each feature vector to be analyzed.
[0033] Optionally, the extraction of the effective safety helmet feature vector further includes:
[0034] Threshold segmentation is performed on the superimposed interference feature values of all feature vectors to be analyzed to obtain the segmentation threshold. Feature vectors to be analyzed whose superimposed interference feature values are lower than the segmentation threshold are recorded as valid safety helmet feature vectors.
[0035] In this invention, a safety helmet detection system based on multi-scale image generation and analysis achieves high robustness and high accuracy in detecting multi-scale safety helmets in complex scenarios by constructing a multi-stage collaborative processing framework that includes image layering, difference modeling, consistency evaluation, interference determination, and target detection. First, the image layering module decomposes the input image into multiple scales, enhancing the system's ability to perceive targets at different scales and improving small target detection performance. The difference modeling module analyzes the changing trends and average deviations of key point position differences to construct a background interference factor, effectively quantifying the degree of interference on spatial structures and suppressing the influence of non-rigid deformation and background clutter. The consistency evaluation module assesses the similarity between the target and the standard template from the perspective of feature intensity response, improving the reliability of feature matching. The interference determination module fuses the background interference factor and feature consistency to construct superimposed interference feature values, adaptively selecting effective feature vectors with low interference and strong discriminative power, and obtaining a pure target representation through feature fusion, significantly improving feature quality. Finally, the target detection module achieves accurate classification and localization based on the optimized target features. The overall system overcomes the problems of false detection and missed detection in complex environments such as changes in lighting, occlusion, and cluttered backgrounds of traditional methods, enhances the generalization ability and engineering practicality of the model, and is suitable for a variety of industrial safety monitoring scenarios. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the structure of the first embodiment of the safety helmet detection system based on multi-scale image generation and analysis of the present invention;
[0038] Figure 2 This is a flowchart illustrating the specific steps involved in extracting each feature component to be analyzed and each standard feature component in the safety helmet detection system based on multi-scale image generation and analysis according to the present invention.
[0039] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0042] In one embodiment, such as Figure 1 As shown, a safety helmet detection system based on multi-scale image generation and analysis is provided. The system includes an image layering module, a difference modeling module, a consistency evaluation module, an interference determination module, and a target detection module.
[0043] The image layering module is used to acquire multi-scale images of the scene to be tested and the standard safety helmet image, and denot them as the multi-scale image to be analyzed and the standard multi-scale image; multi-scale feature decomposition is performed on the multi-scale image to be analyzed and the standard multi-scale image to be analyzed, and each feature component to be analyzed and each standard feature component are extracted.
[0044] The image layering module can be a processing unit used to decompose the input image at multiple scales to generate images of different resolution levels. This enhances the system's ability to perceive targets at different scales (especially small-sized safety helmets) and improves the foundation for feature extraction in multi-scale scenes. In this embodiment, the image layering module can perform layer-by-layer downsampling or frequency band separation on the original image using multi-scale representation methods such as image pyramids, Gaussian pyramids, or wavelet transforms. Furthermore, the image layering module can include, but is not limited to, one or more of Gaussian image pyramid modules, Laplacian image pyramid modules, and wavelet multi-scale decomposition modules. The scene image to be tested can be a field monitoring image containing potential safety helmet targets, which can be used as the system input source to provide visual information about the area to be detected. In an exemplary embodiment, the scene image to be tested can be acquired in real time by cameras deployed at construction sites or industrial areas. The standard safety helmet image can be a reference image of a safety helmet worn correctly, without interference, at a frontal angle, and using it as a template benchmark for structural and strength feature comparison with the image to be tested. For example, the standard safety helmet image can be a pre-collected and labeled standard sample, which is then stored in the system template library after standardization processing.
[0045] The multi-scale image to be analyzed can be a set of images with different resolutions obtained by multi-scale decomposition of the scene image under test. It can be used to provide multi-granular visual representations for subsequent feature extraction, especially enhancing the discernibility of small targets at low resolution levels. In a specific embodiment, the multi-scale image to be analyzed can include, but is not limited to, low-frequency approximation sub-images, high-frequency detail sub-images, and mesoscale transition sub-images. The standard multi-scale image can be a set of multi-resolution templates obtained by processing a standard safety helmet image using the same multi-scale decomposition strategy. It can be used for feature alignment and comparison with the multi-scale image to be analyzed at the same scale level. Further, the standard multi-scale image can include, but is not limited to, standard low-frequency sub-images, standard high-frequency sub-images, and standard mesoscale sub-images. Multi-scale feature decomposition can be the process of decomposing an image into multiple feature components according to frequency or scale dimensions. It can be used to separate structural and detail information in an image, facilitating independent modeling at different scales. In an exemplary embodiment, multi-scale feature decomposition can employ wavelet packet decomposition, Contourlet transform, and non-subsampled contourlet transform.
[0046] The feature components to be analyzed can be local feature representations with specific scale and orientation characteristics extracted from the multi-scale image to be analyzed. They can be used to carry the structural and strength information of the target under test at a specific scale and are used for comparison with standard feature components. For example, the feature components to be analyzed can include, but are not limited to, edge-dominated feature components, texture-dominated feature components, and shape-dominated feature components. The standard feature components can be the feature representation of an ideal safety helmet at the corresponding scale extracted from a standard multi-scale image. They can be used as a comparison benchmark to measure the structural consistency and strength similarity of the feature components to be analyzed. In a specific embodiment, the standard feature components can include, but are not limited to, standard edge feature components, standard texture feature components, and standard shape feature components.
[0047] Acquiring multi-scale images of the scene to be tested and the standard safety helmet image can be achieved by applying a multi-scale decomposition algorithm to the original input image to generate a set of sub-images with different resolutions or frequency bands. Furthermore, this operation can be implemented by constructing a Gaussian pyramid for layer-by-layer downsampling and using discrete wavelet transform for frequency band separation, thus providing the system with multi-granular visual input and enhancing its adaptability to scale changes. Multi-scale feature decomposition of the multi-scale image to be analyzed and the standard multi-scale image can be performed by extracting directional or structural local feature components at each multi-scale image level. Furthermore, this operation can be implemented by using Gabor filter banks to extract multi-directional texture components and using Contourlet transform to obtain edge and contour components, thus separating structural and detailed information in the image, facilitating subsequent key point extraction and comparison. Extracting each feature component to be analyzed and each standard feature component can be achieved by collecting the corresponding feature representations of the test and standard images according to scale levels from the multi-scale feature decomposition results. Furthermore, this operation can be achieved by extracting features by scale index pairing and automatically matching corresponding components through feature clustering, thereby establishing scale-aligned feature pairs and providing a basis for cross-scale comparison.
[0048] The difference modeling module is used to analyze the differences between the feature components to be analyzed and the standard feature components regarding the coordinates of the corresponding key points of each feature. It constructs the variation coefficient of the difference between the key point positions of the feature components to be analyzed and the standard feature components, and obtains the background interference factor of the feature components to be analyzed by combining the average level of the difference between the feature components to be analyzed and the standard feature components regarding the coordinates of the corresponding key points of the feature.
[0049] The difference modeling module can be a processing unit used to analyze the deviation patterns between the features to be tested and the standard template at key point spatial locations and to quantify the degree of interference. It can be used to construct a background interference factor to suppress the misleading effect of structural shifts caused by occlusion, deformation, or background clutter on feature matching. In this embodiment, the difference modeling module can extract the corresponding coordinates based on the key point detection algorithm and then calculate the statistical characteristics (such as variance and mean) of the difference sequence to model the interference intensity. Furthermore, the difference modeling module can receive the feature components to be analyzed and the standard feature components output from the image layering module, providing a background interference factor for the interference determination module.
[0050] The coordinates corresponding to the location of key features can be the spatial coordinates of salient structural points that have a semantic correspondence with the standard feature components to be analyzed. These coordinates can be used as the basic unit for structural comparison, quantifying spatial deformation and positional shift. In a specific embodiment, the coordinates corresponding to the location of key features can be extracted through corner detection (such as Harris), saliency detection, or deep keypoint regression networks. For example, the coordinates corresponding to the location of key features may include, but are not limited to, the coordinates of the brim corner, the center coordinates of the top of the hat, and the center coordinates of a safety sign.
[0051] The coefficient of variation of keypoint position differences can be a statistical measure describing the trend of changes in the positional deviations of multiple keypoints across scale or spatial dimensions. It can be used to reflect the non-rigidity of the overall structural deformation and to distinguish between real targets and interference. In an exemplary embodiment, the coefficient of variation of keypoint position differences can be obtained by calculating the slope, variance, or mean of the first-order difference sequence of each keypoint coordinate difference. The average level of the coordinate differences corresponding to the feature keypoint positions can be a measure of the mean of all corresponding keypoint position deviations, which can be used to characterize the overall spatial offset and serve as a basic indicator of background interference intensity. Furthermore, this indicator can be obtained by calculating the arithmetic mean of the Euclidean or Manhattan distances between each keypoint.
[0052] Background interference factor can be a quantitative index of interference constructed by combining the variation trend of key point position differences with the average level. It can be used to numerically express the degree to which the current feature component is affected by background clutter, occlusion, or deformation. In a specific embodiment, the background interference factor can be obtained by fusing the variation coefficient and the average level through a weighted linear combination or a nonlinear mapping function. Analyzing the difference in coordinates between the feature component to be analyzed and the standard feature component with respect to the corresponding coordinates of each feature key point position can be done by calculating the coordinate offset of each pair of corresponding key points and analyzing its distribution pattern among multiple key points. Furthermore, this operation can be achieved by fitting the offset vector field and calculating its divergence, statistically analyzing the consistency ratio of the offset direction, etc., thereby revealing the non-rigid characteristics of structural deformation and distinguishing between real targets and interference objects.
[0053] Constructing the variation coefficient of the key point position difference between the feature component to be analyzed and the standard feature component can be achieved by calculating its rate of change or volatility index based on the key point position difference sequence. Further, this operation can be implemented by calculating the standard deviation of the first difference, fitting a linear trend and taking the absolute value of the slope, thereby quantifying the dynamic characteristics of structural offset and aiding in determining whether it is caused by background interference. Combining the average level of the coordinate differences between the feature component to be analyzed and the standard feature component regarding the corresponding feature key point positions, the background interference factor of the feature component to be analyzed can be obtained. This can be achieved by fusing the variation coefficient and the average deviation into a single interference index through function mapping. Further, this operation can be implemented through weighted linear combination: α × average deviation + β × variation coefficient, or normalization after mapping with a nonlinear activation function, thereby comprehensively assessing the degree of interference to the spatial structure by integrating static offset and dynamic changes.
[0054] The consistency evaluation module is used to determine the feature consistency of the feature vector to be analyzed based on the fluctuation of the difference between the feature vector to be analyzed and the standard feature vector with respect to the intensity of each feature key point.
[0055] The consistency evaluation module can be used as an evaluation unit to measure the similarity between the target and the standard template in terms of intensity response at key points. It can improve the reliability of identifying real safety helmet targets during feature matching and reduce mismatches caused by changes in illumination or color interference. In this embodiment, the consistency evaluation module can quantify consistency by comparing the intensity value fluctuations (such as standard deviation and correlation coefficient) between the feature vector to be analyzed and the standard feature vector at each key point. Furthermore, the consistency evaluation module can rely on the feature vectors provided by the image layering module and output the feature consistency to the interference determination module.
[0056] The feature vector to be analyzed can be a vector representation composed of the intensity values of key points in the feature components to be analyzed. It can be used to carry the response intensity information of the target at a specific scale for consistency evaluation. For example, the feature vector to be analyzed can include, but is not limited to, grayscale intensity vectors, gradient magnitude vectors, and local binary mode response vectors. The standard feature vector can be a reference vector composed of the intensity values of corresponding key points in the standard feature components. It can be used as a benchmark for intensity comparison and for calculating feature consistency. In a specific embodiment, the standard feature vector can include, but is not limited to, standard grayscale intensity vectors, standard gradient magnitude vectors, and standard LBP response vectors. The intensity of the feature key point can be the pixel response value or local feature response value at the location of the key point. It can be used to reflect the saliency or stability of the target at that location and for measuring the template matching quality. Further, the intensity of the feature key point can include, but is not limited to, grayscale values, gradient magnitudes, and Gabor filter response values.
[0057] Feature consistency can be a measure of the similarity between the analyzed feature vector and the standard feature vector in terms of the fluctuation of intensity differences at key points. It can be used to assess the consistency of the target and the standard template in terms of intensity response patterns, thereby improving matching reliability. In an exemplary embodiment, feature consistency can be obtained by calculating the standard deviation of the intensity difference, the Pearson correlation coefficient, or the cosine similarity. The feature consistency of the analyzed feature vector is determined based on the fluctuation of the differences between the analyzed feature vector and the standard feature vector regarding the intensity of each feature key point. This can be achieved by calculating the dispersion of the intensity difference sequence or a similarity measure. Furthermore, this operation can be implemented by calculating the standard deviation of the intensity difference and taking its reciprocal, or by calculating the Pearson correlation coefficient as the consistency, thereby assessing the stability of the intensity response pattern and improving matching reliability.
[0058] The interference determination module is used to construct the superimposed interference feature value of each feature vector to be analyzed based on the background interference factor and feature consistency of each feature vector to be analyzed. Based on the superimposed interference feature value of all feature components to be analyzed, the effective safety helmet feature vector in the multi-scale image to be analyzed is extracted, and the target feature after background interference is eliminated is obtained through feature fusion.
[0059] The interference determination module can be a decision unit that integrates spatial structure interference information and feature intensity consistency information to filter effective features. It can adaptively identify and retain feature vectors that are less affected by interference and have strong discriminative power, eliminating the negative impact of low-quality features on the final detection. In this embodiment, the interference determination module can weightedly combine background interference factors and feature consistency to form superimposed interference feature values, and set a threshold or ranking mechanism to filter effective features. Furthermore, the interference determination module can simultaneously receive background interference factors from the difference modeling module and feature consistency from the consistency evaluation module, outputting effective safety helmet feature vectors for subsequent fusion.
[0060] Superimposed interference feature values can be a comprehensive interference evaluation index formed by fusing background interference factors and feature consistency. This index can be used to uniformly characterize the interference level and discriminative ability of feature vectors, and is used for feature selection. In a specific embodiment, superimposed interference feature values can be obtained by fusing two-dimensional indicators through weighted summation, product normalization, or a combination of logistic regression.
[0061] The effective safety helmet feature vector can be a set of low-interference, high-discriminative feature vectors retained after being filtered by the interference determination module. This set can be used as high-quality input for subsequent feature fusion, improving the purity of the target representation. For example, the effective safety helmet feature vector can be obtained by sorting based on superimposed interference feature values or by selecting the top K vectors or vectors that meet certain conditions after threshold truncation.
[0062] Feature fusion can be the process of integrating multiple valid safety helmet feature vectors into a unified target representation. This can be used to generate robust, compact, and interference-resistant target features for use by the detection module. In an exemplary embodiment, feature fusion can be achieved through vector concatenation followed by dimensionality reduction, weighted average fusion, or attention-based weighted aggregation. Furthermore, feature fusion can include, but is not limited to, channel-level feature concatenation, spatial-level feature weighted averaging, and cross-scale attention fusion. The target feature after background interference removal can be the final target representation obtained after feature fusion, which has suppressed the influence of background interference. This can be used as input to the target detection module to support high-precision classification and localization. For example, the target feature after background interference removal can include, but is not limited to, fusion intensity-structure feature vectors, multi-scale aggregated feature tensors, and interference suppression embedding vectors.
[0063] Based on the background interference factor and feature consistency of each feature vector to be analyzed, superimposed interference feature values are constructed for each feature vector. This can be achieved by fusing the two-dimensional indicators into a unified evaluation score. Furthermore, this operation can be implemented through product normalization: (1 - interference factor) × feature consistency, logistic regression weighted combination, etc., thus achieving a unified representation of multi-dimensional interference information and supporting feature selection. Effective safety helmet feature vectors are extracted from the multi-scale images to be analyzed based on the superimposed interference feature values of all feature components to be analyzed. This can be achieved by sorting by superimposed interference feature values or by threshold filtering to retain high-quality feature vectors. Furthermore, this operation can be achieved by selecting the top N highest-scoring feature vectors and retaining all vectors with superimposed values greater than a preset threshold, thus automatically filtering features that are severely interfered with or have weak discriminative power, improving the quality of subsequent processing. The target features after background interference elimination are obtained through feature fusion, which can be achieved by integrating the effective safety helmet feature vectors into a unified representation. Furthermore, this operation can be achieved by vector concatenation followed by dimensionality reduction through a fully connected layer, or by using attention-weighted weighted averaging fusion, thus generating robust and compact target descriptors and suppressing residual interference.
[0064] The target detection module is used to detect safety helmets in the scene image under test by using a classifier combined with target features after background interference removal, and to obtain the existence and category results of the safety helmets.
[0065] The target detection module is a detection unit that performs classification and localization tasks based on optimized target features. It can be used to accurately determine the existence and category of the safety helmet and output the final detection result. In this embodiment, the target detection module can use a pre-trained classifier (such as SVM or neural network) to perform inference on the target features after background interference removal. For example, the target detection module may include, but is not limited to, a support vector machine detector, a fully connected neural network classifier, or a lightweight convolutional integral classifier.
[0066] The classifier can be a machine learning model used to determine the category of input features and can be used to output the presence and specific category of the safety helmet (such as color, type). In a specific embodiment, the classifier can include, but is not limited to, a linear SVM classifier, a multi-layer perceptron, a lightweight convolutional neural network head, etc. By using the classifier to combine the target features after background interference elimination to detect the safety helmet in the image of the待测场景 (to-be-tested scenario), it can be to input the optimized target features into the classifier for inference. Further, this operation can be achieved by using SVM for binary classification (with / without a safety helmet), using a multi-layer perceptron to output multi-category probabilities, etc., so as to achieve high-precision determination of the presence and category of the safety helmet.
[0067] To obtain the presence and category results of the safety helmet, it can be to analyze the output of the classifier and generate a structured detection result. Further, this operation can be achieved by outputting a boolean presence flag and category label, generating a detection box with confidence and category, etc., so as to provide a final output that can be used for warning, recording, or linkage control.
[0068] For example, in the scenario of monitoring high-altitude operations at a construction site, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment can be: deploying a high-definition camera around the tower crane or high-rise scaffolding to collect real-time images of workers' operations; the system first performs multi-scale decomposition on the video frames and extracts the features of small-sized safety helmets at a long distance and large-sized safety helmets at a short distance respectively; the difference modeling module analyzes the position offset of key points such as the brim and the top of the helmet. If the head of the worker is partially blocked by the steel frame and the key point offset is severe, the background interference factor increases; the consistency evaluation module synchronously checks whether the intensity response of the yellow identification area of the safety helmet is consistent with the standard template. If the intensity fluctuates abnormally due to strong light reflection, the feature consistency decreases; the interference determination module synthesizes the two indicators, automatically eliminates low-quality features that are blocked by the steel frame and have severe reflection, and retains reliable features in the clearly visible area; after fusion, the target detection module accurately determines that the worker is wearing a yellow safety helmet, avoiding the traditional YOLO model misjudging the safety helmet as a yellow toolbox due to occlusion.
[0069] This embodiment utilizes an image layering module to enable the system to simultaneously perceive both distant small targets and near large targets through multi-scale decomposition. A difference modeling module constructs a background interference factor based on the differences in coordinates of key point locations and the average level, effectively quantifying structural shifts caused by occlusion, deformation, or background clutter. A consistency evaluation module determines feature consistency by analyzing fluctuations in the intensity differences of key feature points, ensuring matching reliability from the response intensity dimension. An interference determination module fuses the background interference factor and feature consistency to form a superimposed interference feature value, thereby selecting effective safety helmet feature vectors and generating target features after background interference elimination through feature fusion. The target detection module uses this optimized feature to drive a classifier to determine the existence and category of the safety helmet, achieving a highly robust detection effect for multi-scale, interfered safety helmet targets in complex industrial scenarios without relying on end-to-end deep learning.
[0070] In one embodiment, the extraction of each feature component to be analyzed and each standard feature component includes:
[0071] Construct the feature vector to be analyzed and the standard feature vector, and use the feature components of the feature vector to be analyzed after multi-scale feature decomposition as the feature components to be analyzed.
[0072] Multi-scale feature decomposition is performed on the standard feature vector, and the discrete coefficients of each feature component of the standard feature vector are calculated. The average of the discrete coefficients of all feature components of the standard feature vector is used as the judgment threshold. Feature components with discrete coefficients greater than or equal to the judgment threshold are all recorded as standard feature components.
[0073] Constructing the feature vector to be analyzed and the standard feature vector can be done by extracting vector representations that characterize the overall properties of the safety helmet from the original image or preprocessing results. Furthermore, the construction of the feature vector to be analyzed and the standard feature vector can be achieved by aggregating local descriptors (such as SIFT and LBP) to generate a global vector, or by extracting intermediate layer features using a shallow convolutional network and flattening them into vectors, thus providing a unified input format for subsequent multi-scale feature decomposition.
[0074] The feature components obtained by multi-scale eigenvalue decomposition of the feature vector to be analyzed can be used as the feature components to be analyzed. This can be achieved by applying a multi-scale decomposition method to the feature vector and directly treating its output components as the feature components to be analyzed. Furthermore, this operation can be achieved by using wavelet transform to decompose it into low-frequency and high-frequency coefficients, or by using a multi-resolution analysis framework to generate a scale subspace projection, thus preserving all scale information for subsequent comparison with the refined standard feature components. Multi-scale eigenvalue decomposition of the standard feature vector can be performed by inputting the standard feature vector into the same multi-scale decomposition process as the side to be analyzed, obtaining the feature component set at the corresponding scale. Furthermore, this operation can be achieved by synchronously using the same wavelet basis functions as the side to be analyzed, or by using a Gaussian-Laplace pyramid with the same number of layers, thus ensuring that the standard and the features to be tested are aligned in the same scale space, supporting effective comparison.
[0075] Calculating the coefficient of variation (COP) of each feature component of the standard eigenvector can be achieved by independently calculating the ratio of its standard deviation to its mean for each feature component. Furthermore, this operation can be implemented by directly calculating the statistic for the vector components or by flattening the tensor components before calculating the COP. This allows for the assessment of the inherent variability of each component within the standard template and the identification of highly discriminative components. The COP can be a dimensionless statistical indicator used to measure the degree of dispersion of the numerical distribution of each feature component after multi-scale decomposition of the standard eigenvector. It can be used to quantify the structural significance or discriminative potential of each feature component within the standard template; a high COP indicates that the component has stronger discriminative power. In this embodiment, the COP can be obtained by calculating the ratio of the standard deviation of a feature component to its mean, reflecting relative volatility. For example, the COP may include, but is not limited to, one or more of the following: coefficient of variation, normalized standard deviation, and relative dispersion.
[0076] The average of the discrete coefficients of all feature components in the standard feature vector is used as the judgment threshold. This can be achieved by summing and averaging the discrete coefficients of all components to form a unified screening benchmark. Furthermore, this operation can be implemented using an arithmetic mean or a weighted average (weighted according to scale importance), thus enabling an adaptive threshold mechanism that does not require manual setting and improving the system's generalization ability. The judgment threshold can be an adaptive threshold used to screen effective standard feature components, determined by the average of the discrete coefficients of all standard feature components. It can be used to dynamically set a baseline for retaining highly discriminative feature components, avoiding over- or under-screening caused by using a fixed threshold. In an exemplary embodiment, the judgment threshold can be obtained by calculating the arithmetic mean of the discrete coefficients of all feature components obtained from multi-scale decomposition of the standard feature vector. In a specific embodiment, the judgment threshold can include, but is not limited to, a global average threshold, a scale-independent threshold, and a template adaptive threshold. Feature components with discrete coefficients greater than or equal to the judgment threshold are recorded as standard feature components. This can be achieved by comparing the discrete coefficient of each standard feature component with the judgment threshold and retaining those that meet the condition. Furthermore, this operation can be achieved through hard threshold truncation (retaining only those ≥ the threshold) or soft screening (weighting the retained components by the discrete coefficient), thereby automatically eliminating low-discriminative or redundant standard feature components and improving the quality of subsequent matching.
[0077] For example, in a factory workshop monitoring scenario under strong light variations, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment can be as follows: In a workshop environment with strong metallic reflection, a standard safety helmet image is decomposed into multiple feature components at multiple scales; the system calculates the discrete coefficient of each component and finds that the component at the edge of the brim has a high discrete coefficient due to its clear structure, while the component in the diffuse reflection area at the top has a low discrete coefficient due to its sensitivity to light; the average discrete coefficient of all components is used as a threshold, and only components with high discrete coefficients, such as the brim and markings, are retained as standard feature components; when an interfering object resembling a yellow helmet but without a brim structure appears in the image to be analyzed, because it cannot be matched on the high discriminative standard component, the difference modeling module will detect a serious shift in the key point position, an increase in the background interference factor, and it will ultimately be filtered by the interference judgment module to avoid false alarms.
[0078] This embodiment achieves input consistency by uniformly constructing feature vectors, fully preserving multi-scale information of the analysis side to maintain detection sensitivity, introducing a discrete coefficient metric on the standard side to quantify the discriminative ability of each component, using a globally average discrete coefficient to adaptively set the judgment threshold to achieve template refinement, and filtering high-discriminative standard feature components based on the threshold to eliminate redundant or easily disturbed components. This allows the standard template to focus on structurally stable and discriminative feature dimensions, solving the problem of redundant or easily disturbed features in the standard template in traditional methods. This improves the accuracy of difference modeling and consistency evaluation, and ultimately enhances the robustness of the system in recognizing real safety helmet targets in complex industrial scenarios, especially improving the technical effect of preventing false detections and missed detections caused by background clutter, changes in illumination, or non-rigid deformation.
[0079] In one embodiment, the construction of the feature vector to be analyzed and the standard feature vector further includes:
[0080] All feature data in the multi-scale image to be analyzed are arranged in ascending order of scale to form the feature vector to be analyzed, and all feature data in the standard multi-scale image are arranged in ascending order of scale to form the standard feature vector. The feature data includes the coordinates of feature key points, edge intensity values and color channel data at each scale.
[0081] The feature data can be a set of raw perceptual information extracted from multi-scale images to construct feature vectors. It includes multimodal data encompassing spatial, structural, and appearance dimensions, and can serve as the basic unit for constructing the feature vectors to be analyzed and those used as standards, supporting unified representation across scales and multiple attributes. In this embodiment, the feature data can be organized according to the scale index generated by the image layering module when generating multi-scale images. For example, the feature data can include, but is not limited to, one or more of the following: feature keypoint coordinates, edge intensity values, and color channel data at each scale.
[0082] Ascending scale arrangement can be a way to organize feature data from different scale levels in order of resolution from low to high (or scale factor from small to large). It can be used to preserve the hierarchical topological relationship between scales, allowing feature vectors to implicitly encode multi-scale evolution paths, which is convenient for modeling positional shifts and intensity change trends. In an exemplary embodiment, ascending scale arrangement can sort the feature data according to the scale index (such as pyramid level number) when the image layering module generates multi-scale images.
[0083] The feature vector to be analyzed is formed by arranging all feature data in the multi-scale image in ascending order of scale. This can be achieved by extracting the coordinates of key points, edge intensity values, and color channel data from each scale layer of the multi-scale image, and then concatenating them into a one-dimensional vector according to the scale hierarchy from smallest to largest. Furthermore, this operation can be achieved by first extracting three types of features by scale grouping, and then flattening and concatenating them layer by layer; or by sorting each type of feature by scale separately and then cross-concatenating them to maintain attribute continuity. This results in a structured and ordered multi-scale fusion representation that retains scale evolution information and supports subsequent dynamic modeling of positional shifts and intensity fluctuations.
[0084] Arranging all feature data from a standard multi-scale image in ascending scale order to form a standard feature vector can be achieved using the exact same process as the analysis target: extracting and organizing three types of feature data from the standard multi-scale image in ascending scale order to form the standard feature vector. In a specific embodiment, this operation can be achieved by simultaneously using the same scale index and feature extraction operator; and by maintaining the same field order and dimensional layout as the analysis target vector during vector construction. This ensures structural alignment between the standard template and the target in the feature space, improving semantic consistency and comparability in the comparison process.
[0085] For example, in a scenario of panoramic monitoring of a construction site with mixed near and far views, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment can be: the monitoring screen simultaneously includes workers on a distant tower crane (with small-sized safety helmets) and ground workers nearby (with large-sized safety helmets). The system extracts the coordinates of key points on the top of the helmet, Canny edge response intensity, and RGB color values from the multi-scale image to be analyzed at three scale levels, and arranges them into feature vectors to be analyzed in ascending order of scale (small → medium → large); the standard feature vector is constructed according to the same rule. The difference modeling module finds that: at the small scale level, the key point coordinates shift significantly but the edge intensity is stable; at the large scale level, the coordinate shift is small but the color fluctuates significantly due to the influence of shadows. Since the three types of features in the vector are arranged in an orderly manner according to scale, the module can accurately identify the pattern of "the same structure behaves differently at different scales" rather than misjudging it as multiple interference objects; the consistency evaluation module integrates the three types of data to judge the overall matching degree, avoiding misjudging a yellow safety helmet as a yellow helmet based solely on color.
[0086] This embodiment provides a safety helmet detection system based on multi-scale image generation and analysis. It constructs a feature vector by arranging all feature data from the multi-scale image to be analyzed in ascending order of scale, and a standard feature vector by arranging all feature data from the standard multi-scale image in ascending order of scale. By explicitly preserving the order relationship between multi-scale levels, the subsequent difference modeling module can track the positional shift and intensity evolution trend of the same semantic structure under scale changes. It also integrates spatial, structural, and appearance-based heterogeneous features to enhance the integrity of the representation. Simultaneously, it ensures that the standard and test feature vectors use the exact same construction logic to ensure strict alignment in the feature space. This significantly improves the reliability of consistent evaluation, effectively solves the false detection and false negative problems caused by the lack of refined multi-scale modeling in traditional methods, and enhances the system's robustness to non-rigid deformation, scale diversity, and background interference.
[0087] In one embodiment, the formula for calculating the coefficient of variation of the key point position difference between the feature component to be analyzed and the standard feature component is as follows:
[0088] ;
[0089] In the formula, Let be the coefficient of variation of the keypoint position difference between the i-th eigenvector to be analyzed and the j-th standard eigenvector, norm() be the exponential normalization function, and M be the number of elements in the keypoint position difference vector between the i-th eigenvector to be analyzed and the j-th standard eigenvector. and These are the k-th and (k-1)-th elements in the key point position difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector, respectively.
[0090] The coefficient of variation of key point location differences can be a structural disturbance quantification index calculated based on the rate of change of adjacent elements in the key point location difference vector. It can reflect the severity of local geometric distortion in the spatial structure between the target and the standard template, and distinguish between reasonable deformation and abnormal interference. In this embodiment, the coefficient of variation of key point location differences can be obtained by calculating the sum of squares of the differences between adjacent elements in the key point location difference vector, and then processing it using an exponential normalization function. Furthermore, the coefficient of variation of key point location differences can be used as one of the core outputs of the difference modeling module to construct a background interference factor.
[0091] The exponential normalization function can be a monotonically decreasing function that nonlinearly compresses the input value and maps it to the interval [0, 1]. It can be used to suppress the influence of extreme values in the rate of change of key point position differences, enhance the sensitivity to moderate disturbances, and improve the stability of disturbance assessment. In an exemplary embodiment, the exponential normalization function can normalize the original rate of change using an exponential form such as exp(-x) or 1 / (1+exp(x)). Exemplarily, the exponential normalization function can include, but is not limited to, one or more of the following: negative exponential decay function, sigmoid normalization function, and Softmax normalization variant. The key point position difference vector can be a one-dimensional vector composed of the coordinate deviations of the corresponding key points between the feature vector to be analyzed and the standard feature vector arranged in order. It can be used to carry the spatial offset information of multiple key points and serve as the basic data structure for structural disturbance analysis. In a specific embodiment, the key point position difference vector can be formed by calculating the Euclidean distance or coordinate component difference for each pair of corresponding key points and then concatenating them in a preset order. Further, the key point position difference vector can include, but is not limited to, the X-direction offset vector, the Y-direction offset vector, and the polar coordinate radial offset vector.
[0092] The k-th element can be the k-th coordinate deviation value in the keypoint position difference vector, arranged sequentially. It can be used to characterize the spatial offset of the k-th keypoint and to calculate the local rate of change. For example, the k-th element can be one or more of the following: hat top offset, hat brim left corner offset, safety sign center offset, etc. The (k-1)-th element can be the coordinate deviation value preceding the k-th element in the keypoint position difference vector. It can be used together with the k-th element to form a local gradient calculation unit to measure the offset change of adjacent keypoints. In an exemplary embodiment, the (k-1)-th element can be, but is not limited to, the offset of the keypoint preceding the hat top, the hat brim right corner offset, the back of the head contour point offset, etc.
[0093] The number of elements M can be the number of keypoints contained in the keypoint position difference vector, which can be used to determine the number of summation terms in the calculation of the variation coefficient, ensuring comparability under different keypoint configurations. In this embodiment, the number of elements M can be directly given by the total number of keypoints determined in the keypoint detection stage. Calculating the variation coefficient of the keypoint position difference between the i-th feature vector to be analyzed and the j-th standard feature vector can be based on the sum of squared differences between adjacent elements in the keypoint position difference vector, after exponential normalization, divided by the total number of elements M. Further, this operation can be achieved by first calculating the sum of squared L2 norms of all adjacent differences and then substituting it into exp(-sum / M) to obtain the variation coefficient, or by normalizing each adjacent difference separately and then calculating the mean, finally applying exponential compression. This allows for the quantification of the local severity of structural disturbances and improves sensitivity to non-rigid deformation and occlusion.
[0094] Applying an exponential normalization function to the keypoint location difference vector can be achieved by inputting the original structural change rate into an exponential nonlinear function for compression mapping. In a specific embodiment, this operation can be achieved by directly compressing the original change rate using exp(-x), or by using 1 / (1+exp(α·x-β)) with parameter adjustment for Sigmoid normalization. This can reduce the dominant role of extreme offsets in the overall evaluation and enhance the model's ability to discriminate moderate disturbances.
[0095] Obtaining the difference sequence of adjacent elements within the keypoint position difference vector can be achieved by traversing the keypoint position difference vector and calculating the difference between the k-th and (k-1)-th elements sequentially. For example, this operation can be approximated by calculating the first-order forward difference or the central difference, thereby extracting the gradient of local structural changes and providing basic data for dynamic perturbation modeling. Quantifying the degree of structural perturbation based on the rate of change of adjacent elements in the keypoint position difference vector can be achieved by using the energy or variance of the adjacent difference sequence as a measure of structural instability. In an exemplary embodiment, this operation can be achieved by calculating the mean square value of the difference sequence as the perturbation energy, or by statistically analyzing the proportion of absolute differences exceeding a threshold as the severe perturbation ratio, thereby enabling an upgrade from static average offset to dynamic local distortion modeling paradigms and improving the accuracy of perturbation identification.
[0096] For example, in the scenario of detecting safety helmets at a distance under strong perspective distortion, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment can be as follows: In the monitoring screen of the top of a large steel structure factory, the shape of the safety helmet of a worker in the distance is significantly stretched due to the perspective effect, and the key points of the brim are asymmetrically offset. After the system extracts the difference vector of the key point position, it finds that the offset values of adjacent key points (such as the top of the helmet → left brim, left brim → right brim) jump sharply, and the variance of the difference sequence is high. After exponential normalization, the change coefficient still maintains a high value, indicating that there is abnormal structural disturbance. Although the average position deviation is small (because the overall is still in a reasonable range), the high change coefficient increases the background interference factor, and this feature component is downweighted by the interference judgment module. At the same time, the feature component with moderate resolution, slight deformation, low change coefficient, and high consistency at another scale is selected as an effective feature. Finally, the fusion result accurately identifies the wearing of a safety helmet and avoids misjudging it as no helmet or interference due to perspective distortion.
[0097] This embodiment calculates the variation coefficient of key point position differences between the feature components to be analyzed and the standard feature components. It quantifies structural disturbances based on the rate of change of adjacent elements in the key point position difference vector, applies exponential normalization to the rate of change to suppress the influence of extreme values, constructs a local gradient sequence using the offset difference between adjacent key points, and solves the disturbance index by normalizing the total number of key points. This achieves dynamic modeling of local disturbances in spatial structures by quantifying the rate of change of adjacent elements in the key point position difference vector and nonlinearly compressing it using an exponential normalization function. This mechanism surpasses the traditional static evaluation method that relies solely on average position deviation, and can sensitively capture non-rigid geometric distortions caused by occlusion, perspective distortion, or background clutter. Combined with the original process of the difference modeling module, this coefficient provides a highly discriminative dynamic dimension for background interference factors, enabling the system to effectively distinguish between reasonable scale scaling and abnormal structural disturbances. The resulting multi-stage collaborative framework significantly improves the robustness of detecting small-scale, partially occluded, or deformed safety helmets in complex industrial scenarios, mitigating the technical effects of false positives and false negatives caused by the coarse structural modeling of traditional methods.
[0098] In one embodiment, the construction of the keypoint location difference vector further includes:
[0099] For each feature vector to be analyzed and each standard feature vector, the coordinates corresponding to the positions of all feature key points in the feature vector to be analyzed and the standard feature vector are arranged according to the key point number to form a vector, and are denoted as the key point coordinate vector of the feature vector to be analyzed and the key point coordinate vector of the standard feature vector.
[0100] The difference vector between the feature vector to be analyzed and the standard feature vector with respect to the key point coordinate vector is denoted as the key point position difference vector between the feature vector to be analyzed and the standard feature vector.
[0101] The coordinates corresponding to the keypoint locations can be the position values of semantically significant salient structural points detected in the feature components within the image coordinate system. These coordinates can serve as the basic units for constructing keypoint coordinate vectors and carry local geometric information of the target. In an exemplary embodiment, the coordinates corresponding to the keypoint locations can be obtained through corner detectors, heatmap regression, or keypoint annotation tools.
[0102] The keypoint sequence number can be a predefined, unique index number used to identify the semantic identity of key points on various types of safety helmets. It can be used to ensure that key points in the target and standard feature vectors are arranged in the same semantic order and maintain structural topological consistency. For example, the keypoint sequence number can include, but is not limited to, one or more of the following: the keypoint sequence number on the top of the helmet, the keypoint sequence number on the left brim, and the keypoint sequence number at the center of the safety sign. In a specific embodiment, the keypoint sequence number can be manually defined based on the safety helmet's geometry during the system initialization phase or obtained in a fixed order through clustering learning. The keypoint coordinate vector can be an ordered vector formed by arranging the coordinates corresponding to the positions of all feature keypoints in a feature vector according to a preset keypoint sequence number. It can be used to achieve spatial semantic alignment of key points and ensure that the target and the standard template are geometrically comparable. Furthermore, the keypoint coordinate vector can be formed by concatenating two-dimensional or one-dimensional coordinates into a vector according to a fixed keypoint topological order. For example, the keypoint coordinate vector can include, but is not limited to, a two-dimensional coordinate concatenation vector, a polar coordinate representation vector, and a normalized image coordinate vector.
[0103] The keypoint coordinate vector of the feature vector to be analyzed can be a vector composed of the coordinates of all keypoints in the feature vector to be analyzed, arranged according to the keypoint index. It can be used to structurally represent the spatial keypoint layout of the target under test and to provide a point-by-point comparison with a standard template. In one specific embodiment, the keypoint coordinate vector of the feature vector to be analyzed can be obtained by extracting the keypoint coordinates from the feature components to be analyzed, rearranging them according to a preset index, and then concatenating them. The keypoint coordinate vector of the standard feature vector can be a reference vector composed of the coordinates of all keypoints in the standard feature vector arranged according to the same keypoint index. It can be used to provide a structural benchmark for an ideal safety helmet and to perform an ordered geometric comparison with the target under test. In an exemplary embodiment, the keypoint coordinate vector of the standard feature vector can be obtained by extracting keypoints from a standard multi-scale image and arranging them according to the same indexing rule as the target end.
[0104] The coordinates of all key points in the feature vector to be analyzed are arranged according to their key point numbers to form a vector, denoted as the key point coordinate vector of the feature vector to be analyzed. This can be achieved by reordering the detected key point coordinates according to a preset key point number and concatenating them into a single vector. Furthermore, this operation can be achieved by sequentially concatenating the (x, y) coordinates of each key point into a one-dimensional vector of the form [x1, y1, x2, y2, ..., xn, yn], or by converting the key points into relative coordinates relative to the target center and then concatenating them in order. This allows for the establishment of a structured geometric representation of the target and ensures semantic alignment with the standard template.
[0105] The coordinates corresponding to all keypoint positions in the standard feature vector are arranged according to the keypoint index to form a vector, denoted as the keypoint coordinate vector of the standard feature vector. This can be achieved by using the same indexing rule as the target end to vectorize the standard keypoint coordinates. Furthermore, this operation can be implemented by rearranging the coordinates using a keypoint index mapping table shared with the target end, or by directly reading pre-stored standardized keypoint coordinate vectors from a standard template library. This allows for the generation of a structurally consistent reference benchmark and ensures the semantic correctness of the difference vector calculation.
[0106] The difference vector of keypoint coordinate vectors can be obtained by subtracting the corresponding elements of the keypoint coordinate vectors to be analyzed from those of the standard keypoint coordinate vectors. This difference vector directly reflects the spatial offset of each semantic keypoint and forms the mathematical basis for the keypoint position difference vector. In one specific embodiment, the difference vector of keypoint coordinate vectors can be obtained by performing vector subtraction. Calculating the difference vector between the feature vector to be analyzed and the standard feature vector with respect to the keypoint coordinate vectors can be achieved by performing element-wise subtraction on two keypoint coordinate vectors of the same dimension. Furthermore, this operation can be achieved by directly calculating the Euclidean coordinate difference or by calculating the difference in a normalized coordinate system to eliminate scale effects, thereby obtaining the accurate spatial offset of each semantic keypoint and providing raw data for structural perturbation analysis.
[0107] The difference vector is denoted as the keypoint position difference vector between the feature vector to be analyzed and the standard feature vector. This can be achieved by formally naming the calculated difference vector as the keypoint position difference vector and passing it to subsequent modules. Furthermore, this operation can be implemented by storing the difference vector as a structured field for the difference modeling module to call, or by attaching a scale level label and then sending it to the change coefficient calculation unit. This allows for the transformation from an unordered set of keypoints to a structured difference representation and supports high-precision interference modeling.
[0108] The keypoint location difference vector can be obtained by subtracting the corresponding elements of the keypoint coordinate vectors to be analyzed from those of the standard feature vector. It can be used to structurally represent the point-by-point deviations of two sets of keypoints in spatial location and provide ordered, computable input for difference modeling. In one specific embodiment, the keypoint location difference vector can be obtained by first constructing the keypoint coordinate vectors to be analyzed and the standard, and then subtracting them element-by-element to form the difference vector.
[0109] For example, in the scenario of safety helmet detection in densely populated work areas, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment can be used in narrow construction passages where multiple workers' heads obscure each other, with some safety helmets only showing the top and one side of the brim. The system detects three key points at a certain scale: the top of the helmet, the left brim, and the center of the marker. It constructs the coordinate vector of the key points to be analyzed according to the preset sequence [1: top of the helmet, 2: left brim, 3: center of the marker]. The standard template has five complete points under the same sequence number, but the system only uses the first three points for alignment (the remaining points are marked as invalid or interpolated). When calculating the difference vector, only the valid key points participate in the calculation. Because the left brim is offset by a large amount due to occlusion while the top of the helmet is offset by a small amount, the difference vector shows a non-uniform distribution. The subsequent calculation of the change coefficient captures the drastic shift in the offset between adjacent key points (top of the helmet → left brim), which is judged as local structural damage rather than overall translation. This feature component is given a high background interference factor and is reduced in weight during the interference judgment stage. At another scale, due to the corrected viewpoint and the complete key points, its difference vector is smooth and is selected as a valid feature, ultimately achieving accurate detection.
[0110] This embodiment provides a safety helmet detection system based on multi-scale image generation and analysis. By arranging the coordinates of all key points in the feature vector to be analyzed and the standard feature vector according to the key point number to form their respective key point coordinate vectors, and calculating the difference vector between the two as the key point position difference vector, the system achieves the technical effect of distinguishing between overall geometric transformation and local structural damage under conditions of occlusion, deformation, or partial visibility, and improving the stability of feature matching and detection robustness.
[0111] In one embodiment, the formula for calculating the background interference factor of the feature component to be analyzed is:
[0112] ;
[0113] In the formula, Let N be the background interference factor for the i-th eigenvector to be analyzed, and N be the number of standard eigenvectors. Let be the mean of the elements in the key point position difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector.
[0114] The keypoint position difference vector can be a vector composed of the coordinate differences of corresponding keypoints in the i-th feature vector to be analyzed and the j-th standard feature vector. It can be used to characterize the local offset pattern of the two features in the spatial structure, providing basic data for background interference factors. In this embodiment, the keypoint position difference vector can be formed by calculating the difference between the horizontal and vertical coordinates (e.g., Δx, Δy) of each pair of corresponding keypoints and arranging them in order of the keypoints. Furthermore, the keypoint position difference vector can include, but is not limited to, one or more of the following: horizontal offset component vector, vertical offset component vector, and Euclidean distance deviation vector.
[0115] The mean value of elements within the keypoint position difference vector can be the arithmetic mean of all elements in the keypoint position difference vector. This can be used to compress multidimensional positional deviations into a single scalar, reflecting the overall spatial offset between the pair of features. In an exemplary embodiment, the mean value of elements within the keypoint position difference vector is obtained by summing the components of each dimension of the difference vector and dividing by the total number of elements. The number N of standard feature vectors can be the total number of feature vectors corresponding to the preset standard helmet templates in the system. This can be used to determine the diversity of templates participating in the comparison when calculating background interference factors, improving adaptability to changes in posture and angle. For example, the number N of standard feature vectors is determined by the number of independent feature vectors generated after multi-scale decomposition and keypoint extraction from the standard helmet image library.
[0116] Obtaining the keypoint position difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector can be achieved by calculating the difference point by point in the coordinates of the two sets of corresponding keypoints and forming a vector. Furthermore, this operation can be achieved by calculating the coordinate difference one-to-one according to the keypoint index, or by performing optimal keypoint matching using the Hungarian algorithm before calculating the difference, thereby establishing a quantitative deviation representation of the spatial structure between the target and a specific standard template.
[0117] Calculating the mean of elements within the keypoint location difference vector can be achieved by taking the arithmetic mean of all elements in the difference vector. In a specific embodiment, this operation can be achieved by directly averaging all Δx and Δy, or by first calculating the Euclidean distance of each keypoint and then averaging them, thereby compressing high-dimensional location deviation information into a comparable scalar index. Averaging the mean of elements corresponding to all standard feature vectors to obtain the background interference factor can be achieved by averaging the mean of elements calculated for the i-th feature vector to be analyzed and all N standard feature vectors. Furthermore, this operation can be achieved by simple arithmetic averaging or by template confidence-weighted averaging, thereby integrating the results of multi-template comparisons to obtain a more generalizable structural interference metric. Calculating the background interference factor for the i-th feature vector to be analyzed can be achieved by performing the above three steps, ultimately outputting a scalar interference index, enabling a quantitative assessment of the degree of spatial structural interference to a single feature vector, supporting subsequent feature selection.
[0118] Taking multi-angle safety helmet wearing recognition as an example, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment can be used in construction site entrance monitoring where workers pass by the camera at different angles, causing the safety helmets to exhibit postures such as forward tilting or side tilting. The system pre-stores standard feature vectors (N=5) corresponding to 5 standard viewpoints (front view, left 45°, right 45°, top view, and bottom view). When a side-tilted safety helmet is detected, its feature vector to be analyzed is compared with the 5 standard vectors to calculate the key point position difference vector (such as the offset of the left and right corner points of the brim); the mean of each difference vector is calculated to obtain 5 intermediate values; then these 5 values are averaged to obtain the final background interference factor. If the factor is low, it indicates that although there is a change in posture, the overall structure is still consistent with a certain type of standard template, and it is judged as a valid target; if the factor is too high (such as the key points being severely offset due to the helmet being obscured by steel bars), it is filtered by the interference judgment module.
[0119] This embodiment provides a safety helmet detection system based on multi-scale image generation and analysis. By calculating the element mean of the key point position difference vector between the feature vector to be analyzed and all standard feature vectors, and further averaging all the means, the system can accurately characterize local structural disturbances through the key point position difference vector, achieve dimensionality compression through its element mean, and enhance posture robustness and suppress the influence of individual abnormal matching through multi-template mean aggregation. This can effectively quantify the structural distortion caused by occlusion, deformation or background clutter, provide a reliable basis for the interference judgment module, and thus improve the anti-interference capability and discrimination accuracy of the entire detection system in complex industrial scenarios.
[0120] In one embodiment, the formula for calculating the feature consistency of the feature vector to be analyzed is:
[0121] ;
[0122] In the formula, Let be the feature consistency of the i-th feature vector to be analyzed, and exp() be an exponential function with the natural constant as the base. Let be the information entropy of the feature intensity difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector. Here, the intensity values of all feature keypoint positions in the i-th feature vector to be analyzed and the j-th standard feature vector are sorted according to the keypoint index to form the keypoint intensity vector of the i-th feature vector to be analyzed and the keypoint intensity vector of the j-th standard feature vector. The difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector with respect to the keypoint intensity vector is denoted as the feature intensity difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector.
[0123] The keypoint intensity vector can be an ordered vector formed by arranging the intensity values of all feature keypoint positions in a feature vector according to a unified keypoint index. It can be used to achieve comparability between the tested and standard features under semantic alignment, eliminating matching errors caused by inconsistent keypoint extraction order. In this embodiment, the keypoint intensity vector can rearrange the original intensity values according to a predefined or synchronously generated keypoint index order during detection. For example, the keypoint intensity vector can include, but is not limited to, one or more of the following: a brim-top-identifier ternary intensity vector, a circular sampling intensity sequence, and a gridded region intensity vector. The feature intensity difference vector can be the difference vector obtained by subtracting the corresponding elements of the keypoint intensity vectors of the feature vector to be analyzed from those of the standard feature vector. It can be used to characterize the local deviation pattern of the intensity response at each keypoint, serving as the basis for information entropy calculation. Furthermore, the feature intensity difference vector can be obtained by subtracting two intensity vectors aligned with the same keypoint index element-wise. In an exemplary embodiment, the feature intensity difference vector can include, but is not limited to, a positive offset difference vector, a negative offset difference vector, and a zero-center fluctuation difference vector.
[0124] Information entropy can be a statistical measure of the uncertainty in the numerical distribution of feature intensity difference vectors. It can quantify the degree of disorder in intensity differences; a lower entropy value indicates a more concentrated intensity difference and a more stable pattern, reflecting a higher consistency between the target and the standard template. In a specific embodiment, information entropy can be calculated by treating the difference vector as a discrete probability distribution (normalized or histogram-based) and then calculating the Shannon entropy. For example, information entropy can use Shannon information entropy, Rényi entropy, Tsallis entropy, etc. The exponential function can be an exponential operation function with the natural constant e as its base, used to map information entropy to a non-negative consistency score. It can be used to convert the negative value of information entropy into a monotonically decreasing consistency degree in the interval (0, 1), achieving non-linear enhancement of high consistency features. Furthermore, the exponential function can be used to calculate the information entropy... Calculate exp( This ensures that the smaller the entropy, the larger the output value. In an exemplary embodiment, the exponential function may include, but is not limited to, the natural exponential function, the negative exponential decay map, the monotonically decreasing activation function, etc.
[0125] Feature consistency can be a quantitative indicator that measures the consistency between the analyzed feature and the standard template in terms of keypoint intensity response, obtained by mapping the information entropy of the feature intensity difference vector through an exponential function. It can be used to reflect the uncertainty of the intensity difference distribution using information entropy, and to enhance high-consistency features and suppress low-consistency features through exponential mapping, thereby improving matching robustness. In this embodiment, feature consistency can be calculated by taking the negative exponent after calculating the information entropy of the feature intensity difference vector. Furthermore, feature consistency can rely on the information entropy of the feature intensity difference vector as input and output to the interference determination module to participate in the construction of superimposed interference feature values.
[0126] The intensity values of all keypoint positions in the i-th feature vector to be analyzed are sorted according to the keypoint index to form the keypoint intensity vector of the i-th feature vector to be analyzed. This can be achieved by rearranging the original unordered intensity values according to a unified keypoint index order. Furthermore, this operation can be implemented by using a predefined keypoint topology (such as top of hat → left brim → right brim) with fixed sorting rules, or by synchronously outputting intensity values with IDs during the keypoint detection stage and directly constructing vectors by ID index. This allows for the establishment of semantically aligned vector representations, ensuring dimension-by-dimensional comparison with the standard template at the same keypoint positions.
[0127] The intensity values of all keypoints in the j-th standard feature vector are sorted according to their keypoint indices to form the keypoint intensity vector of the j-th standard feature vector. This can be achieved by aligning the standard intensity values using the same sorting rules as the feature vector to be analyzed. Furthermore, this operation can be implemented by using the same topological indexing rules as the image to be tested, or by directly reading sorted reference intensity vectors from a standard template library, thus ensuring a strict dimensional and semantic correspondence between the standard and the vector to be tested, supporting effective difference calculation. The difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector with respect to the keypoint intensity vector is calculated, denoted as the feature intensity difference vector. This can be achieved by performing element-wise subtraction on two keypoint intensity vectors of the same dimension. Furthermore, this operation can be performed through direct vector subtraction or weighted difference calculation, thereby obtaining the complete distribution of local intensity response deviations, serving as the basis for subsequent uncertainty analysis.
[0128] Calculating the information entropy of a feature intensity difference vector can be achieved by transforming the difference vector into a probability distribution and then applying the Shannon entropy formula. Further, this operation can be performed by constructing a histogram of the difference values (binded and normalized to a probability distribution) and then calculating the entropy, or by treating the difference vector as a continuous signal, fitting the distribution using kernel density estimation, and then integrating to calculate the entropy. This allows for the quantification of the disorder of intensity differences; low entropy indicates a stable and highly consistent response pattern. The feature consistency of the i-th feature vector to be analyzed can be obtained by taking the negative exponent of the information entropy using an exponential function. Calculate exp( This serves as the final consistency score.
[0129] For example, in the scenario of detecting safety helmets under strong light variations, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment can be as follows: In a construction site under direct afternoon sunlight, a worker's safety helmet exhibits strong glare in certain areas, causing models such as YOLO to miss detection due to color distortion. This system first extracts its multi-scale features, detecting five key points (top of the helmet, left and right brims, and center of the front and rear markings) at a certain scale. Intensity vectors of these key points are constructed according to preset numbers [1: top of the helmet, 2: left brims, 3: right brims, 4: front marking, 5: rear marking], respectively, [240, 180, 190, 255, 170]. The vector corresponding to a standard yellow safety helmet is [200, 185, 185, 200, 180]. The difference vector is calculated to obtain [40, −5, 5, 55, −10], whose absolute values are concentrated around ±5 but contain two large deviations (40, 55). After histogram statistics, the information entropy is relatively high (e.g., ...). =1.8); with exp(−1.8)≈0.165, the feature consistency is low. If another worker is in the shadow area, the difference vector is [−10, −8, −7, −12, −9], which is concentrated, H=0.6, and the consistency exp(−0.6)≈0.55 is significantly higher. Based on this, the interference judgment module prioritizes retaining the latter's features to avoid misclassifying reflective areas as non-helmet targets.
[0130] This embodiment provides a safety helmet detection system based on multi-scale image generation and analysis. It forms a keypoint intensity vector by sorting the intensity values of feature keypoints according to a unified keypoint index, calculates the difference between the two to obtain a feature intensity difference vector, uses its information entropy to measure the uncertainty of the intensity response deviation, and applies the exponential function exp( By mapping entropy to feature consistency in the interval [0, 1], establishing semantically aligned keypoint intensity vectors to ensure comparability, using difference vectors to characterize local intensity deviations, quantifying response stability with information entropy, and achieving nonlinear enhancement of consistency scores through exponential mapping, we can achieve a refined evaluation of feature matching quality from the perspective of local intensity response stability. This overcomes the shortcomings of traditional methods that ignore intensity distribution characteristics and can effectively distinguish between real safety helmets and interfering objects under interference such as sudden changes in illumination and local occlusion. This provides a more discriminative consistency basis for the interference judgment module, thereby improving the robustness and accuracy of the entire multi-stage detection framework.
[0131] In one embodiment, the superimposed interference feature value of each feature vector to be analyzed is the ratio of the background interference factor to the feature consistency of each feature vector to be analyzed.
[0132] The ratio can be the result of a numerical division operation between the background interference factor and the feature consistency degree, and can be used as a specific calculation form for the superimposed interference feature value to quantify the comprehensive interference degree of the feature vector. In a specific embodiment, the ratio can be calculated by dividing the background interference factor as the numerator and the feature consistency degree as the denominator; when the feature consistency degree approaches 0, a smoothing constant can be introduced to prevent division by zero. Furthermore, the ratio can include, but is not limited to, directly calculating the background interference factor divided by (feature consistency degree plus smoothing constant) or first normalizing both to the [0, 1] interval before performing the ratio operation, or one or more of the following: For each feature vector to be analyzed, its corresponding background interference factor is divided by the feature consistency degree to obtain a single scalar as its superimposed interference feature value; or the background interference factor and feature consistency degree are first normalized to the [0, 1] interval, and then the ratio operation is performed to enhance numerical stability, thereby achieving a normalized evaluation of feature quality, so that features with low interference and high consistency obtain low ratios, which facilitates subsequent screening by threshold or sorting.
[0133] For example, in a construction site entrance monitoring scenario where strong sunlight and partial occlusion coexist, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment could be as follows: Under direct midday sunlight, a worker's safety helmet is partially obscured by a safety door frame. The system extracts feature vectors at multiple scales: one feature vector has low feature consistency (large intensity response fluctuations) due to glare from the helmet top, but small key point position offsets and low background interference factors; another feature vector comes from an unobstructed area, has high feature consistency, but slight jitter in key point positioning due to distant shooting, and a medium background interference factor. According to the ratio mechanism, the former ratio = 0.2 / 0.4 = 0.5, and the latter ratio = 0.5 / 0.9 ≈ 0.56, both below a set threshold (e.g., 1.0), and is retained; while a feature vector simultaneously affected by strong glare and severe occlusion (interference factor = 0.8, consistency = 0.2, ratio = 4.0) is automatically discarded. Finally, the retained features are fused to accurately determine the presence of the safety helmet.
[0134] This embodiment provides a safety helmet detection system based on multi-scale image generation and analysis. The system uses the superimposed interference feature value of each analyzed feature vector as the ratio of the background interference factor to the feature consistency of that vector. By dividing the corresponding background interference factor of each analyzed feature vector by the feature consistency, a single scalar is obtained as its superimposed interference feature value. This allows the interference judgment module to measure the reliability of features using a unified scale: a low ratio corresponds to high-quality features with stable structure and consistent response, while a high ratio reflects low-quality features that are severely interfered with or semantically ambiguous. This ratio mechanism replaces the unclear fusion method in the original scheme, realizing a more sensitive and discriminative feature selection strategy. This retains effective safety helmet features while suppressing mismatches caused by occlusion, lighting changes, or background clutter. Without changing the front-end image layering and back-end detection structure, it improves the accuracy and robustness of multi-scale safety helmet detection in complex industrial scenarios.
[0135] In one embodiment, the extraction of the effective safety helmet feature vector further includes:
[0136] Threshold segmentation is performed on the superimposed interference feature values of all feature vectors to be analyzed to obtain the segmentation threshold. Feature vectors to be analyzed whose superimposed interference feature values are lower than the segmentation threshold are recorded as valid safety helmet feature vectors.
[0137] The segmentation threshold can be an adaptive critical value used for binarizing and filtering superimposed interference feature values. It can be used to dynamically distinguish between high-interference and low-interference feature vectors, enabling automatic extraction of effective safety helmet feature vectors. In this embodiment, the segmentation threshold can be automatically determined by statistical analysis of the superimposed interference feature values of all feature vectors to be analyzed (e.g., Otsu method, histogram valley detection, or percentile estimation). For example, the segmentation threshold can include, but is not limited to, one or more of the following: Otsu adaptive threshold, histogram bimodal valley value, and fixed percentile threshold.
[0138] Thresholding is performed on the superimposed interference feature values of all feature vectors to be analyzed to obtain a segmentation threshold. This threshold can be calculated based on the distribution characteristics of the superimposed interference feature values to determine an optimal segmentation point. Furthermore, this operation can be achieved by using the Otsu algorithm to maximize the inter-class variance to determine the threshold, or by locating the valley bottom of the local minimum of the superimposed interference feature value histogram as the threshold. This allows for scene-adaptive quantization of the interference level, avoiding the failure problem of using a fixed threshold in different environments.
[0139] Feature vectors whose superimposed interference feature values are lower than the segmentation threshold are recorded as valid safety helmet feature vectors. This can be achieved by comparing the superimposed interference feature value of each feature vector with the segmentation threshold and retaining those below the threshold. In a specific embodiment, this operation can be implemented by comparing vectors one by one and then filtering with a Boolean mask, or by vectorized batch judgment and index extraction, so that only feature vectors with low interference and high consistency in structure and strength are retained for subsequent fusion and detection.
[0140] For example, in a steel structure construction site scenario where strong sunlight and partial occlusion coexist, the safety helmet detection system based on multi-scale image generation and analysis in this embodiment can be as follows: In a complex scene where direct sunlight causes partial reflection of the safety helmet, and the worker's head is partially occluded by a beam, the system calculates the superimposed interference feature values of multiple feature vectors to be analyzed at various scales. Due to the large intensity fluctuations in the reflective area and the drastic shift of key points in the occluded area, the superimposed interference feature values are generally high. The system uses the Otsu method to perform threshold segmentation on these values, automatically obtaining a high but reasonable segmentation threshold. Only feature vectors with superimposed interference feature values lower than this threshold (such as stable features from the unoccluded side view) are retained, thereby effectively eliminating noise features introduced by strong light and occlusion, ensuring that subsequent fusion and detection are based on high-quality input.
[0141] This embodiment provides a safety helmet detection system based on multi-scale image generation and analysis. It obtains a segmentation threshold by thresholding the superimposed interference feature values of all feature vectors to be analyzed, and records the feature vectors whose superimposed interference feature values are lower than the segmentation threshold as valid safety helmet feature vectors. The system adaptively determines the segmentation threshold based on the distribution characteristics of the superimposed interference feature values to achieve scene-adaptive quantification of the interference level. Furthermore, it retains low-interference feature vectors by comparing the superimposed interference feature values with the segmentation threshold to ensure the consistency of the structure and intensity of the input features. This system achieves the technical effects of dynamically adjusting feature selection criteria on the basis of the original framework, effectively suppressing the negative impacts of background clutter, non-rigid deformation, and occlusion, and significantly improving the purity of subsequent feature fusion and the accuracy of target detection. This enhances the robustness and engineering applicability of the overall solution in complex and dynamic operating scenarios.
[0142] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A hard hat detection system based on multi-scale image generation and analysis, characterized by, The system comprises an image layering module, a difference modeling module, a consistency evaluation module, an interference judgment module and a target detection module; The image layering module is configured to obtain multi-scale images of the to-be-tested scene image and the standard safety helmet image respectively, and record the multi-scale images as to-be-analyzed multi-scale images and standard multi-scale images; The to-be-analyzed multi-scale images and the standard multi-scale images are subjected to multi-scale feature decomposition respectively, and each to-be-analyzed feature component and each standard feature component are extracted; The difference modeling module is configured to analyze the difference between the to-be-analyzed feature components and the standard feature components in terms of the corresponding coordinates of the feature key point positions, construct a change coefficient of the key point position difference between the to-be-analyzed feature components and the standard feature components, and obtain a background interference factor of the to-be-analyzed feature components in combination with the average level of the difference between the to-be-analyzed feature components and the standard feature components in terms of the corresponding coordinates of the feature key point positions; The consistency evaluation module is configured to determine the feature consistency degree of the to-be-analyzed feature vector according to the fluctuation of the difference between the to-be-analyzed feature vector and the standard feature vector in terms of the intensity of each feature key point; The interference judgment module is configured to construct a superimposed interference feature value of each to-be-analyzed feature vector according to the background interference factor and the feature consistency degree of each to-be-analyzed feature vector, extract an effective safety helmet feature vector in the to-be-analyzed multi-scale image based on the superimposed interference feature values of all the to-be-analyzed feature components, and obtain a target feature after background interference elimination through feature fusion; The target detection module is configured to use a classifier to combine the target feature after background interference elimination to detect the safety helmet in the to-be-tested scene image, and obtain the existence and category of the safety helmet.
2. The safety helmet detection system based on multi-scale image generation and analysis as claimed in claim 1, wherein, The extraction of each to-be-analyzed feature component and each standard feature component comprises: constructing a to-be-analyzed feature vector and a standard feature vector, and taking each feature component after multi-scale feature decomposition of the to-be-analyzed feature vector as a to-be-analyzed feature component; performing multi-scale feature decomposition on the standard feature vector, calculating the dispersion coefficients of each feature component of the standard feature vector, taking the average of the dispersion coefficients of all the feature components of the standard feature vector as a judgment threshold, and taking each feature component with a dispersion coefficient greater than or equal to the judgment threshold as a standard feature component.
3. The safety hat detection system based on multi-scale image generation and analysis as claimed in claim 2, wherein, The construction of the to-be-analyzed feature vector and the standard feature vector further comprises: arranging all the feature data in the to-be-analyzed multi-scale image in ascending order of scale to form a to-be-analyzed feature vector, and arranging all the feature data in the standard multi-scale image in ascending order of scale to form a standard feature vector; the feature data comprises feature key point coordinates, edge intensity values and color channel data at each scale.
4. The safety hat detection system based on multi-scale image generation and analysis as claimed in claim 1, wherein, The change coefficient of the key point position difference between the to-be-analyzed feature component and the standard feature component corresponds to the following calculation formula: ; In the formula, is the variation coefficient of the key point position difference between the ith analyzed feature vector and the jth standard feature vector, norm() is an exponential normalization function, M is the number of elements in the key point position difference vector between the ith analyzed feature vector and the jth standard feature vector, and are the kth and k-1th elements in the key point position difference vector between the ith analyzed feature vector and the jth standard feature vector, respectively.
5. The safety hat detection system based on multi-scale image generation and analysis as claimed in claim 4, wherein, The construction of the key point position difference vector further comprises: for each to-be-analyzed feature vector and each standard feature vector, arranging the coordinates corresponding to all the feature key points in the to-be-analyzed feature vector and the standard feature vector in a vector according to the key point serial number, and recording the vector as a key point coordinate vector of the to-be-analyzed feature vector and a key point coordinate vector of the standard feature vector; The difference vector between the to-be-analyzed feature vector and the standard feature vector with respect to the key point coordinate vector is denoted as a key point position difference vector between the to-be-analyzed feature vector and the standard feature vector.
6. The safety hat detection system based on multi-scale image generation and analysis as claimed in claim 1, wherein, The calculation formula of the background interference factor of the to-be-analyzed feature component is: ; wherein is the background interference factor of the i-th feature vector to be analyzed, N is the number of standard feature vectors, is the average of the elements within the difference vector of the key point positions between the i-th feature vector to be analyzed and the j-th standard feature vector.
7. The safety hat detection system based on multi-scale image generation and analysis as claimed in claim 1, wherein, The calculation formula of the feature consistency degree of the to-be-analyzed feature vector is: ; In the formula, is the feature consistency degree of the i-th feature vector to be analyzed, exp() is an exponential function with a natural constant as the base, is the information entropy of the feature intensity difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector, wherein the intensity values of all feature key point positions in the i-th feature vector to be analyzed and the j-th standard feature vector are respectively sorted according to the key point serial numbers to form a key point intensity vector of the i-th feature vector to be analyzed and a key point intensity vector of the j-th standard feature vector, and the difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector with respect to the key point intensity vector is denoted as the feature intensity difference vector between the i-th feature vector to be analyzed and the j-th standard feature vector.
8. The safety hat detection system based on multi-scale image generation and analysis as claimed in claim 1, wherein, The superimposed interference feature value of each to-be-analyzed feature vector is a ratio of the background interference factor to the feature consistency degree of the to-be-analyzed feature vector.
9. The safety hat detection system based on multi-scale image generation and analysis as claimed in claim 1, wherein, The extraction of the effective safety helmet feature vector further includes: The superimposed interference feature values of all to-be-analyzed feature vectors are threshold segmented to obtain a segmentation threshold, and the to-be-analyzed feature vectors with superimposed interference feature values lower than the segmentation threshold are recorded as effective safety helmet feature vectors.