Reflective vest wearing state identification method and system based on deep learning
Through a deep learning-based method, the area image of the reflective vest is extracted and multi-scale feature extraction and region distribution analysis is performed, which solves the problem of low recognition accuracy in the prior art, and realizes high-accuracy wearable state recognition under different lighting conditions.
Patent Information
- Application Number
- CN202510470212.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art is susceptible to light conditions when identifying the wearable state of the reflective vest, and it is difficult to adapt to dynamic scene changes, resulting in low recognition accuracy.
Using a deep learning-based method, the wearable state is identified by acquiring real-time monitoring images, extracting the reflective vest area image, and performing multi-scale feature extraction and regional distribution analysis.
It significantly improves the robustness and accuracy of the wearable state recognition of reflective vests, and can accurately capture the dynamic changes of reflective materials under different lighting conditions, reducing the risk of misjudgment.
Smart Images

Figure CN119992602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular to a method and system for identifying the wearing status of a reflective vest based on deep learning. Background Art
[0002] In the fields of industrial security, traffic management, etc., workers wearing reflective vests is an important measure to ensure personal safety. In the prior art, the automatic recognition method for the wearing status of reflective vests usually relies on single-dimensional feature detection, such as color threshold segmentation or reflection intensity analysis based on reflective materials. For example, by setting a fixed color interval or reflection intensity threshold, the reflective area is segmented from the monitoring image, and the wearing status is judged based on the area or shape rules of the area; or the edge features of the reflective stripes are extracted using traditional image processing algorithms, and the wearing status is classified in combination with static template matching. However, the above methods have significant limitations in practical applications: first, single-dimensional feature detection is easily interfered by complex lighting conditions, such as strong light irradiation causing overexposure of the reflective area, shadow coverage causing a sudden drop in reflection intensity, etc., which will cause feature extraction failure or misjudgment; secondly, static thresholds or template matching are difficult to adapt to dynamic scene changes. For example, when the human body posture is tilted, the reflective vest is partially blocked by tools, or the clothing is wrinkled and deformed, the traditional method cannot accurately distinguish between normal wear and abnormal state, resulting in missed detection or false alarm. Therefore, there is an urgent need for a method that can improve the recognition accuracy. Summary of the invention
[0003] The purpose of the present invention is to provide a method and system for identifying the wearing status of a reflective vest based on deep learning. The embodiment of the present invention is implemented as follows: In a first aspect, an embodiment of the present invention provides a method for identifying a reflective vest wearing state based on deep learning, the method comprising: Acquire a real-time monitoring image of a target scene, wherein the real-time monitoring image contains at least one human object to be detected; Extracting a reflective vest region image corresponding to the human object to be detected from the real-time monitoring image, and performing multi-scale feature extraction on the reflective vest region image to generate initial wearing state features; Based on a pre-trained wearing state recognition model, the initial wearing state feature is analyzed for regional distribution to generate a target wearing state feature; wherein the pre-trained wearing state recognition model is obtained by fusing multimodal training data, and the multimodal training data includes sample images of reflective vests under different lighting conditions; According to the matching degree between the target wearing state feature and the preset wearing state threshold, the reflective vest wearing state recognition result of the human object to be detected is determined; wherein the recognition result includes normal wearing, not wearing or partially blocked state.
[0004] In a second aspect, the present invention provides a computer system, comprising: one or more processors; Memory; One or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method as described above is implemented.
[0005] Beneficial effects of the present invention: The reflective vest wearing state recognition method based on deep learning provided by the embodiment of the present invention can accurately capture the dynamic change characteristics of reflective materials under different lighting conditions by extracting the reflective vest area from the real-time monitoring image and performing multi-scale feature fusion, combined with the spatial distribution analysis ability of the pre-trained model, thereby significantly improving the robustness and accuracy of wearing state recognition. Specifically, the multi-scale feature extraction module comprehensively covers the structural characteristics and reflective characteristics of the reflective vest by fusing low-resolution texture features, medium-resolution edge features and high-resolution detail features, effectively overcoming the problem of feature loss caused by uneven lighting or local occlusion; the pre-trained wearing state recognition model adaptively learns the spatial distribution law of reflective vests in different lighting scenes based on multimodal training data, and enhances the ability to distinguish the distribution changes of reflective areas in complex environments; through the matching threshold judgment mechanism, combined with multi-dimensional statistics such as reflective material coverage area, stripe continuity and posture matching, efficient classification of wearing states is achieved to avoid the risk of single feature misjudgment. In addition, the dynamic correction mechanism can dynamically eliminate temporary occlusion interference through comprehensive analysis of multi-angle verification images and historical time series data, further improving the real-time and reliability of state recognition. In summary, the method provided by the embodiment of the present invention can provide accurate wearing status monitoring and real-time warning capabilities for industrial security scenarios, reduce the risk of human missed detection, and ensure the safety of operators.
[0006] Other features will be described in part in the following description. Those skilled in the art will discover these features in part upon inspection of the following content and drawings, or may learn these features through production or use. The features of the present application may be implemented and obtained by practicing or using various aspects of the methods, tools, and combinations listed in the detailed examples described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for describing the embodiments of the present invention are briefly introduced below.
[0008] Figure 1This is a flowchart of a reflective vest wearing status recognition method based on deep learning provided by an embodiment of the present invention.
[0009] Figure 2 It is a schematic diagram of the functional module architecture of a reflective vest wearing status recognition device based on deep learning provided in an embodiment of the present invention.
[0010] Figure 3 It is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0011] The following describes the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. The terms used in the implementation method part of the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.
[0012] The executor of the reflective vest wearing status recognition method based on deep learning in the embodiment of the present invention is a computer system, including but not limited to electronic equipment systems such as servers, personal computers, laptops, tablet computers, and smart phones.
[0013] The reflective vest wearing state recognition method based on deep learning provided by the embodiment of the present invention is as follows Figure 1 As shown, the following steps are included: Step S100: Acquire a real-time monitoring image of a target scene, wherein the real-time monitoring image contains at least one human object to be detected.
[0014] Exemplarily, the target scene may refer to a physical environment that needs to be monitored for safety, such as a construction site, a mining operation area, a traffic intersection, or a nighttime construction site. The real-time monitoring image refers to a sequence of digital images captured continuously at fixed time intervals by an optical camera deployed in the scene. The human object to be detected refers to a set of pixel areas with human morphological characteristics in the image. Specifically, the real-time monitoring image can be collected by a camera with high dynamic range imaging capabilities to ensure that the outline of the human object and the details of the attached wear can be fully recorded under strong light reflection or low illumination conditions. For example, in a construction site scene, the camera can collect RGB images with a resolution of 1920×1080 at a rate of 30 frames per second, in which there are three workers who are pouring concrete in one frame of the image, and their torso and limbs constitute three independent human objects to be detected. It should be noted that the real-time monitoring image meets the preset clarity threshold, that is, the pixel coverage of the key anatomical points of the human object (such as shoulders and waist) in the image must exceed the minimum detection size to avoid subsequent recognition failure due to excessive distance. On this basis, real-time monitoring images are transmitted to the central processing unit via wired or wireless networks to complete the format standardization and noise filtering preprocessing of image data, providing standardized input for subsequent analysis.
[0015] Step S200: extracting a reflective vest area image corresponding to the human object to be detected from the real-time monitoring image, and performing multi-scale feature extraction on the reflective vest area image to generate initial wearing state features.
[0016] Exemplarily, the reflective vest area image refers to a local image subset cropped based on preset geometric constraints after locating the torso area of the human object to be detected through the human key point detection algorithm, and its bounding box must completely cover the effective reflective material distribution range of the reflective vest. Multi-scale feature extraction refers to the use of a hierarchical convolution kernel structure in a convolutional neural network to extract texture, shape and reflection intensity features at different spatial resolution levels of the reflective vest area image, and the initial wearing state feature refers to the multi-dimensional tensor data formed by splicing feature maps of each scale. Specifically, an instance segmentation model based on Mask R-CNN can be used to perform pixel-level segmentation on human objects, and the precise boundary coordinates of the reflective vest in the torso are obtained through regression calculation. For example, the upper body area of the human object is divided into a rectangular area of 128×256 pixels as the reflective vest area image. Subsequently, a parallel feature extraction branch with three convolution kernel sizes of 3×3, 5×5, and 7×7 was constructed to capture micro-texture details (such as jagged edges of reflective strips), meso-structural features (such as spacing of reflective strips), and macro-contour morphology (such as overall symmetry of the vest) from the image of the reflective vest area. The multi-scale output is mapped to a unified dimensional space through a cross-layer feature fusion mechanism to form an initial wearing state feature vector with 512 channels. For example, in a night scene, the image of the reflective vest area contains bright reflective stripes caused by car lights. Multi-scale feature extraction can effectively distinguish the brightness gradient difference between normal reflective and overexposed areas, avoiding the misjudgment of strong light interference as not being worn.
[0017] Step S300: Based on a pre-trained wearing state recognition model, the initial wearing state feature is analyzed for regional distribution to generate a target wearing state feature; wherein the pre-trained wearing state recognition model is obtained by fusing multimodal training data, and the multimodal training data includes sample images of reflective vests under different lighting conditions.
[0018] Exemplarily, the pre-trained wearing state recognition model refers to a deep neural network that completes parameter optimization on a labeled data set containing multiple environmental conditions, regional distribution analysis refers to analyzing the statistical distribution law of the initial wearing state features in the reflective vest area through a spatial attention mechanism, and the target wearing state features refer to the high-order feature representation after semantic enhancement. The multimodal training data specifically covers sample images under complex conditions such as natural lighting, artificial fill lighting, backlighting, shadow occlusion, and rain and fog interference, ensuring that the model has cross-scenario generalization capabilities. In specific implementation, the wearing state recognition model can adopt a Transformer architecture with a self-attention module, encode the initial wearing state features into serialized inputs according to spatial positions, and calculate the association weights between different feature regions through a multi-head attention mechanism. For example, under backlight conditions, the model will increase the attention to the edge contour features of the vest and suppress the loss of details in the central area caused by backlight; in rainy and foggy weather samples, the model automatically corrects color distortion through cross-modal contrast learning and accurately identifies reflective strips partially covered by water stains. During the training process, light intensity normalization and adversarial sample generation techniques are used to enable the model to extract light invariance features from multimodal data. After regional distribution analysis, the target wearing state features will include discriminative indicators such as reflective material coverage, wearing symmetry index, and occlusion area confidence. For example, a 2048-dimensional dense vector is generated to represent the complete wearing degree of the vest.
[0019] Step S400: Determine the reflective vest wearing state recognition result of the human object to be detected according to the matching degree between the target wearing state feature and the preset wearing state threshold; wherein the recognition result includes normal wearing, not wearing or partially blocked state.
[0020] Exemplarily, the preset wearing state threshold refers to the classification decision boundary determined by statistical learning, the matching degree refers to the distance measurement value between the target wearing state feature and the center of each state cluster in the embedding space, and the recognition result is determined by maximum likelihood estimation or support vector machine classifier. Specifically, the normal wearing state corresponds to the feature space area where the reflective vest completely covers the torso and is not blocked by external objects, the unworn state corresponds to the feature distribution where the area of the reflective material area is lower than the minimum threshold, and the partially blocked state is between the two and there are local feature anomalies. When implementing the classification, the cosine similarity is used to calculate the matching score between the target wearing state feature and the three types of reference templates. When the matching degree of the normal wearing category exceeds 0.95, it is directly judged as a compliant state; if the matching degree is in the range of 0.75 to 0.95 and a local feature mutation is detected, a partial occlusion judgment is triggered; when the matching degree is lower than 0.4, it is judged as not worn. For example, if the similarity between the wear state feature of a target and the normal wear template is 0.92, but the similarity with the partially occluded template is 0.85, the system will further analyze the occlusion-sensitive dimension in the feature space to confirm whether the tool bag shoulder strap covers the reflective strip, and finally output the partial occlusion state. This method adapts to the recognition sensitivity requirements of different scenarios through a dynamic threshold adjustment mechanism to ensure that classification stability can be maintained under strong environmental interference.
[0021] As an implementation manner, in step S200, extracting the reflective vest area image corresponding to the human object to be detected from the real-time monitoring image may include the following steps: Step S210: performing human body contour segmentation on the real-time monitoring image to obtain a human body contour boundary box of the human object to be detected.
[0022] Human contour segmentation refers to the process of separating the pixel area of the human body from the real-time monitoring image through the semantic segmentation algorithm. The human contour bounding box refers to the set of geometric parameters that annotate the image position of the human object in the form of rectangular coordinates. Specifically, the instance segmentation model based on Mask R-CNN can be used to classify the input image at the pixel level, and the bounding box coordinates and the corresponding binary mask of each human object to be detected can be obtained through regression calculation. For example, in the construction site scene, there are three workers in the real-time monitoring image. The model outputs three independent human contour bounding boxes respectively, each of which contains the upper left corner coordinates (x1, y1) and the lower right corner coordinates (x2, y2), and its value is normalized to between 0 and 1 according to the image resolution. It should be noted that human contour segmentation needs to exclude non-human interference, such as hanging reflective warning signs or local reflective areas of mobile mechanical equipment, to ensure that the bounding box only covers the complete human form. In the implementation process, the model enhances the generalization ability through a variety of posture samples (such as bending, raising hands, squatting) contained in the pre-training data set to avoid contour breakage or bounding box offset due to changes in human posture. For example, when a worker bends over to operate equipment, the model can still accurately identify the contours of his back and legs, generating a bounding box that closely fits the actual position of the human body, providing reliable input for subsequent torso area positioning.
[0023] Step S220: Based on the human body contour boundary box, locate the torso region of the human body object to be detected, and perform color space conversion on the torso region to generate a first candidate region.
[0024] The trunk area refers to the anatomical range of the human body from the shoulder to the waist. The positioning process uses a proportional mapping algorithm to define a fixed-scale rectangular sub-area within the human body contour bounding box. The color space conversion refers to converting the original RGB image into an HSV (hue, saturation, brightness) color space that is more suitable for reflection intensity analysis. Specifically, according to the statistical law of human morphology, the trunk area occupies the height range of 60% to 80% of the middle part of the bounding box in the vertical direction and is the same width as the bounding box in the horizontal direction. For example, for a human body contour bounding box with a size of 320×480 pixels, the trunk area is defined as a rectangular area from the ordinate y=96 pixels to y=384 pixels, and the width remains unchanged at 320 pixels. Subsequently, the area is converted from the RGB space to the HSV space, and the reflective properties of the reflective material are enhanced using the value channel. For example, under backlight conditions, the silver-gray stripes of the reflective vest may appear dark in the RGB space due to underexposure, while the value channel of the HSV space can effectively distinguish the brightness difference between the reflective material and ordinary clothing through linear stretching, generating a first candidate area with higher contrast. It should be noted that the bicubic interpolation algorithm needs to be used during the color space conversion process to maintain the image resolution and avoid the loss of high-frequency details due to color quantization.
[0025] Step S230: performing high-reflection area detection on the first candidate area, and screening out a set of pixels with reflection intensity higher than a preset threshold.
[0026] Exemplarily, high-reflection area detection refers to a pixel-level threshold segmentation operation based on the brightness channel. The reflection intensity is quantified by the value of the brightness channel in the HSV space, and the preset threshold is calibrated by experiments according to the optical properties of the reflective material. Specifically, the brightness threshold V_threshold=200 (value range 0-255) is set, and all pixels with V≥200 in the first candidate area are classified as high-reflection pixels to form a binary mask image. For example, in a night scene, the glass bead material on the worker's reflective vest produces strong reflection due to the illumination of the car lights, and its brightness value can reach more than 240, while the brightness value of ordinary work clothes is usually less than 180. The pixel set of the reflective stripes can be accurately extracted through threshold segmentation. It should be noted that the preset threshold needs to be dynamically adjusted according to the ambient light. For example, under strong light conditions at noon, the brightness value of the reflective material may reach 255 due to overexposure. At this time, it is necessary to combine the saturation channel for auxiliary judgment to avoid misjudging white clothing as a reflective area. During implementation, an adaptive threshold algorithm can be used to dynamically calculate the local optimal segmentation threshold based on the peak position of the brightness histogram of the first candidate region to ensure detection robustness under different lighting conditions.
[0027] Step S240: determining the boundary coordinates of the reflective vest area image according to the connectivity distribution of the pixel set, and cropping the image area corresponding to the boundary coordinates from the real-time monitoring image.
[0028] Connectivity distribution refers to the spatial connection relationship between adjacent pixels in a binary image, and boundary coordinates refer to the coordinates of the minimum rectangle vertices covering all connected pixels. For example, a morphological closing operation can be performed on a set of highly reflective pixels to fill holes, and then the parameters of the circumscribed rectangle of the largest connected area can be extracted by an edge tracking algorithm. For example, in rainy and foggy weather, the stripes of a reflective vest may be partially covered by water droplets, resulting in multiple discrete connected areas in the pixel set. By calculating the area of each area and retaining the largest connected area, the interference of water droplet reflection can be eliminated and the boundary of the reflective vest area can be accurately delineated. During implementation, the calculation of boundary coordinates needs to take into account the curvature of the human body trunk, and the elastic bounding box algorithm is used to allow the rectangular box to bend in the vertical direction with the shape of the vest, thereby tightly wrapping the distribution area of the reflective material. For example, the actual distribution of a reflective vest is an inverted trapezoid, the left and right boundaries of the bounding box remain vertical, and the upper and lower boundaries are dynamically adjusted according to the extreme points of the connected area, and finally a reflective vest area image with a size of 256×128 pixels is generated for subsequent feature extraction.
[0029] As an implementation manner, the step S240, determining the boundary coordinates of the reflective vest area image according to the connectivity distribution of the pixel set, includes: Step S241: performing eight-neighborhood connected region detection on the pixel set to generate a connected region set including at least one connected region.
[0030] Eight-neighborhood connected region detection refers to checking the connection status of eight adjacent pixels in the upper, lower, left, right and diagonal directions with a pixel as the center. The connected region set refers to a list of independent regions consisting of spatially continuous pixel groups. For example, in a scene where the reflective vest stripes are broken, the pixel set may contain three vertically arranged connected regions, each corresponding to a reflective strip. In specific implementation, a two-pass scanning algorithm can be used. For example, the first pass traverses the image pixels and marks temporary labels, and the second pass merges equivalent labels to generate the final connected region. For example, a certain reflective vest area is blocked by a button, causing the middle reflective strip to be split into two parts, the algorithm identifies it as two connected regions and assigns labels ID=1 and ID=2 respectively, ensuring that subsequent processing can analyze each sub-region independently.
[0031] Step S242: extracting geometric attributes of each connected region in the connected region set, wherein the geometric attributes include region area, aspect ratio and minimum circumscribed rectangle angle.
[0032] Geometric attribute extraction refers to the calculation of the shape parameters of the connected area through mathematical morphological operations. The area of the region refers to the total number of pixels contained in the connected area, the aspect ratio refers to the ratio of the width to the height of the minimum bounding rectangle, and the minimum bounding rectangle angle refers to the angle between the main axis of the rectangle and the horizontal axis of the image. For example, a connected area contains 1200 pixels, its minimum bounding rectangle width is 60 pixels, the height is 20 pixels, the aspect ratio is 3:1, and the main axis angle is 1.5 degrees (close to the horizontal direction). It should be noted that the calculation of the aspect ratio needs to exclude noise interference. For example, when the connected area has jagged edges due to uneven lighting, the minimum area bounding rectangle algorithm based on the convex hull is used to improve the measurement accuracy and ensure that the geometric attributes truly reflect the structural characteristics of the reflective stripes.
[0033] Step S243: Based on the preset reflective vest shape rule, the candidate connected regions in the connected region set are screened, which satisfy the aspect ratio within the preset range, the region area is greater than the minimum area threshold, and the angle of the circumscribed rectangle is aligned with the axis of the human body trunk.
[0034] The shape rules of reflective vests refer to the aspect ratio range, minimum reflective area requirements and directional consistency constraints defined according to the standard reflective vest design specifications. For example, the preset aspect ratio interval is [2.5, 5.0], the minimum area threshold is 800 pixels, and the deviation between the circumscribed rectangle angle and the trunk axis does not exceed ±5 degrees. Specifically, during the screening process, a connected area with an aspect ratio of 4.2, an area of 950 pixels, and a circumscribed rectangle angle of 2 degrees (the human trunk axis is usually in the vertical direction) is determined to be a candidate connected area; and another area with an aspect ratio of 1.8 and an area of 600 pixels is eliminated because it does not meet the conditions. It should be noted that the human trunk axis is approximated by the perpendicular bisector of the human body contour bounding box. When the angle between the circumscribed rectangle angle and the perpendicular bisector exceeds the threshold, it indicates that the reflective strip may be tilted or twisted, which does not conform to the normal wearing form.
[0035] Step S244: performing region merging processing on the candidate connected regions, merging adjacent candidate connected regions whose spatial distance is less than a preset merging threshold into an extended connected region.
[0036] Exemplarily, the region merging process can be to merge adjacent regions into a single region by calculating the shortest Euclidean distance between candidate connected regions, and the preset merging threshold is set according to the standard spacing of reflective stripes. For example, the merging threshold is set to 15 pixels. When the minimum edge spacing of two candidate connected regions is less than this value, they are determined to belong to different sections of the same reflective strip and merged. In specific implementation, a method based on region growing is adopted: the centroid of each candidate connected region is used as a seed point, and it expands outward until it touches the merging threshold boundary of other regions. For example, there are two broken reflective strips caused by wrinkles in the middle of the reflective vest. The centroid spacing is 12 pixels, which is less than the merging threshold, so they are merged into an extended connected region, thereby restoring the complete reflective strip shape.
[0037] Step S245: performing convex hull detection on the extended connected area to determine the vertex coordinates of the minimum convex polygon covering all merged pixels.
[0038] Convex hull detection refers to calculating the convex shell of the extended connected area. The minimum convex polygon vertex coordinates refer to the set of polygon vertices that can contain all pixels and whose internal angles are all less than 180 degrees. Specifically, the Graham scanning algorithm is used to first find the pixel point with the smallest ordinate as the reference point, sort the remaining points by polar angle, and construct the convex hull in sequence. For example, an extended connected area contains scattered pixels, and the coordinates of its convex hull vertices are {(x1,y1), (x2,y2), (x3,y3), (x4,y4)}, forming a quadrilateral surrounding all reflective pixels. It should be noted that convex hull detection can eliminate the influence of depressions or burrs inside the area, generate a smooth edge contour, and provide a geometric basis for subsequent boundary coordinate calculations.
[0039] Step S246: Calculate the maximum left and right boundary coordinates in the horizontal direction and the maximum upper and lower boundary coordinates in the vertical direction according to the extreme point positions of the vertex coordinates of the minimum convex polygon, and generate the boundary coordinates of the reflective vest area image.
[0040] The extreme points are the maximum and minimum coordinate values of the convex polygon vertices in the horizontal and vertical directions. The boundary coordinates are defined by these extreme points as the four vertices of the rectangular area. Specifically, the left boundary in the horizontal direction is the minimum x value of all vertices, and the right boundary is the maximum x value; the boundary in the vertical direction is the minimum y value, and the bottom boundary is the maximum y value. For example, in the convex hull vertex set, x_min=100, x_max=300, y_min=50, y_max=200, then the boundary coordinates of the reflective vest area image are the rectangular box (100,50,300,200). During the implementation process, the boundary coordinates need to be expanded, expanding outward by 2-3 pixels to ensure that the edge of the reflective material is completely wrapped to avoid losing key features during cropping. For example, the size of the final cropped image area is (x_max - x_min + 6) × (y_max - y_min + 6), that is, 206×156 pixels, retaining the transition area between the reflective stripes and the surrounding background for multi-scale feature extraction.
[0041] As an implementation manner, in step S200, multi-scale feature extraction is performed on the reflective vest area image to generate initial wearing state features, including: Step S250: Input the reflective vest area image into a multi-branch feature extraction network, the multi-branch feature extraction network includes a first branch network, a second branch network and a third branch network; wherein the first branch network is used to extract low-resolution texture features of the reflective vest area image, the second branch network is used to extract medium-resolution edge features, and the third branch network is used to extract high-resolution detail features.
[0042] Exemplarily, the multi-branch feature extraction network is a plurality of convolutional neural network structures deployed in parallel, and each branch processes the input image at a different sampling rate to capture cross-scale features. The first branch network downsamples the input to 1 / 4 of the original size through a convolutional layer with a step size of 2, and extracts macro texture features such as the distribution density of reflective materials; the second branch network maintains the original resolution, enhances the edge response through the Sobel operator, and captures the boundary between the reflective strip and the background; the third branch network uses a hole convolution to expand the receptive field, while maintaining high resolution. While extracting micro details such as the arrangement pattern of reflective particles. For example, inputting a 256×128 pixel reflective vest area image, the first branch outputs a 64×32×64 feature map, the second branch outputs a 256×128×128 feature map, and the third branch outputs a 256×128×256 feature map, which respectively represent visual information at different levels of abstraction.
[0043] Step S260: performing cross-level feature fusion on the low-resolution texture features, medium-resolution edge features and high-resolution detail features to generate a fused feature map.
[0044] Cross-level feature fusion refers to aligning multi-scale features to a unified spatial dimension through upsampling, downsampling and channel splicing operations, and integrating complementary information. Specifically, the low-resolution texture features of the first branch are upsampled to the original size and added channel by channel with the medium-resolution edge features of the second branch; the result is then spliced with the high-resolution detail features of the third branch along the channel axis to form a 512-channel fused feature map. For example, after the low-resolution features are upsampled to 256×128 by bilinear interpolation, they are added to the 128 channels of the medium-resolution features to generate 192-channel intermediate features; they are then spliced with the 256 channels of the high-resolution features to finally obtain a 448-channel fused feature map. It should be noted that a 1×1 convolution kernel is used for channel dimension reduction during the fusion process to avoid dimensionality explosion and enhance feature correlation. In addition, the semantic alignment and information complementarity of multi-scale features can be achieved through the Feature Pyramid Network (FPN) structure. Specifically, the low-resolution texture features are upsampled and added element by element with the medium-resolution edge features, and the result is upsampled again and fused with the high-resolution detail features. For example, the 1 / 16 scale feature map of the first branch is upsampled to 1 / 8 scale through transposed convolution and added to the feature map of the same scale of the second branch; it is then upsampled to the original scale again and channel-joined with the features of the third branch. During the fusion process, 1×1 convolution is introduced to adjust the number of channels to reduce feature redundancy after fusion. The fused feature map will simultaneously contain the global layout information, local edge structure, and fine-grained texture details of the reflective vest, providing a multi-level basis for wearing status analysis.
[0045] Step S270: performing channel attention weighting on the fused feature map, enhancing the feature channel weights related to the reflective material, and compressing the spatial dimension to generate the initial wearing state feature.
[0046] Channel attention weighting dynamically adjusts the importance of each channel through the SENet (Squeeze-and-Excitation Network) mechanism. The reflective material related channels refer to the feature channels that are identified as having a high contribution to the classification of the wearing state during the training process. Specifically, the fused feature map is globally averaged and pooled to generate a channel description vector. The dependency between channels is learned through the fully connected layer and the weight coefficient is output. The weight is multiplied by the original feature map channel by channel. For example, after the 448-channel fused features are weighted by attention, the weights of the 112th, 215th, and 398th channels related to the reflective intensity of the reflective strip are increased to 1.5 times, while the weights of channels unrelated to the background are reduced to less than 0.3. Subsequently, the 256×128×448 feature map is compressed into a 448-dimensional vector by global maximum pooling, which is used as the initial wearing state feature to input the subsequent classification model. This feature vector can encode the wearing integrity, occlusion degree, and reflective characteristics of the reflective vest through end-to-end training, providing a discriminative basis for the final state recognition.
[0047] As an implementation manner, the training process of the pre-trained wearing state recognition model includes the following steps: Step S10: Acquire multiple groups of training samples, each group of training samples includes a reflective vest sample image, a corresponding wearing state label, and a lighting condition label.
[0048] The reflective vest sample images are raw image data containing the wearing status of reflective vests collected by different camera equipment in various scenes. The wearing status label refers to the classification result of the wearing status (normal wearing, not wearing or partially blocked) verified by manual annotation or sensor verification, and the lighting condition label refers to the classification of the ambient light intensity level (low light, high light or dynamic change conditions) obtained by photometer measurement or image metadata analysis. Specifically, the training sample set covers scenes such as construction sites, traffic intersections, and night work areas. For example, it contains 5,000 sample images with a resolution of 1920×1080, including 3,000 samples of normal wearing status (including 800 images of overexposed reflective stripes under strong light at noon, 1,200 images of low light at dusk, and 1,000 images of dynamic lighting in rain and fog), 1,000 samples of not wearing status (workers only wear ordinary work clothes), and 1,000 samples of partially blocked status (reflective vests are partially covered by tool bags or safety ropes). It should be noted that the sample images need to undergo geometric correction and white balance preprocessing to eliminate the interference of lens distortion and color temperature differences on model training. For example, the checkerboard calibration method is used to correct the barrel distortion caused by the wide-angle lens to ensure the consistency of the spatial proportions of the image in the reflective vest area.
[0049] Step S20: performing adaptive lighting enhancement on the reflective vest sample image to generate an enhanced sample image; wherein the adaptive lighting enhancement includes dynamically adjusting brightness compensation parameters according to the lighting condition label.
[0050] Adaptive illumination enhancement is a preprocessing method that adjusts image brightness, contrast, and noise distribution based on illumination condition labels. Brightness compensation parameters refer to adjustable coefficients that control the gamma correction curve, histogram equalization intensity, and noise injection amount. For example, for samples with low illumination labels, the restricted contrast adaptive histogram equalization (CLAHE) algorithm is used to enhance local contrast, while adding random noise that conforms to the Poisson distribution to simulate the granular effect of low illumination sensors; for samples with high illumination labels, the brightness value of the overexposed area is compressed by the highlight suppression function, and the global brightness mean is adjusted to a preset range (such as RGB mean 180-200) by linear scaling. In specific implementation, the dynamic adjustment mechanism is implemented through a lookup table, and a mapping relationship between illumination intensity levels and processing parameters is pre-established. For example, low illumination levels correspond to CLAHE grid size 8×8 and contrast limit 2.0, and high illumination levels correspond to brightness scaling factor 0.7 and saturation gain 1.2. This method ensures that the enhanced sample image retains the reflective characteristics of the reflective material while enhancing the robustness of the model to illumination interference.
[0051] As an implementation manner, the step S20 of adaptively enhancing the reflective vest sample image to generate an enhanced sample image includes: Step S21: determining the illumination intensity level of the current sample according to the illumination condition label.
[0052] The light intensity level is a discrete light intensity interval divided according to the photometric data. For example, low light corresponds to 0-200 lux (such as night scenes), high light corresponds to 2000-10000 lux (such as direct sunlight at noon), and the dynamic change condition refers to the light intensity fluctuation exceeding 500 lux in a single sample (such as intermittent shadows caused by cloud movement). Specifically, by parsing the aperture, shutter speed and ISO value in the image EXIF metadata, combined with the actual illumination value recorded by the ambient light sensor, the light level of the sample is comprehensively determined. For example, if the EXIF parameters of a sample are f / 2.8, 1 / 60s, ISO 1600, and the ambient light is recorded as 150lux, it is judged as a low light level; another sample parameter is f / 8, 1 / 1000s, ISO 100, and the ambient light is 8000 lux, it is classified as a high light level. This classification provides a control basis for subsequent adaptive enhancement.
[0053] Step S22: If the light intensity level is a low light condition, local contrast enhancement is performed on the reflective vest sample image, and random noise is superimposed to simulate low light interference.
[0054] Local contrast enhancement is an operation that improves the visibility of dark details in local areas of an image. Random noise simulation refers to adding signal distortion that conforms to the characteristics of low-light imaging. Specifically, the CLAHE algorithm is used to divide the image into local areas of 32×32 pixels. Each area is histogram-equalized and the contrast increase is limited to no more than 3.0 to prevent excessive noise amplification. Subsequently, a Gaussian noise matrix with a mean of 0 and a variance of 0.01 is generated and superimposed on the enhanced image at the pixel level. For example, the RGB value of the reflective stripes in the original low-light sample is (30,35,40), which is increased to (80,85,90) by CLAHE, and fine-tuned to (82,83,89) after superimposing noise, retaining the true noise characteristics while enhancing the reflective area. It should be noted that the noise injection intensity needs to be positively correlated with the ISO value. For example, the ISO 1600 sample adds noise with a variance of 0.02, and the ISO 3200 sample increases the variance to 0.03 to simulate the grain effect at different sensitivities.
[0055] Step S23: If the light intensity level is a high light condition, suppressing overexposed areas of the reflective vest sample image and reducing the global brightness to a preset range.
[0056] Overexposed area suppression is to restore highlight details through threshold detection and brightness compression. The preset brightness range refers to the RGB mean value interval (such as 180-200) set according to the reflective characteristics of the reflective material. Specifically, detect pixels in the image where any value of the RGB channel exceeds 240, and apply an S-curve adjustment to them: linearly compress the input value I to the interval [200, 220]. For example, the formula is I out =200+20*(I-240) / (255-240). Then, linear scaling is used to adjust the RGB mean of the entire image to 190±5. For example, the original value of the overexposed area of a high-light sample is (255,250,245), which is (220,215,210) after compression. The global brightness mean is reduced from 230 to 195, allowing the texture details of the reflective stripes (such as the direction of the stitching) to appear. It should be noted that the brightness adjustment needs to be performed in the HSV space to avoid color cast, and only the lightness channel (V value) is modified before converting back to the RGB space.
[0057] Step S24: If the illumination intensity level is a dynamically changing condition, multiple frames of the reflective vest sample image are fused in time series to generate an enhanced sample image with balanced illumination.
[0058] Multi-frame temporal fusion is to align and weighted average multiple frames of images collected continuously in the same scene to smooth the impact of illumination fluctuations. Specifically, 5 frames of images before and after the dynamic illumination sample (time window of 100ms) are selected, the inter-frame displacement is estimated by the optical flow method and affine transformation alignment is performed, and then the median of the RGB channel is taken for each pixel to generate a fused image. For example, a dynamic illumination sample sequence contains brightness fluctuations caused by cloud occlusion (the brightness difference between frames is up to 300 lux). After median fusion, the brightness standard deviation of the reflective vest area is reduced from 45 to 12, eliminating the impact of instantaneous shadows or overexposure on single-frame images. This method effectively improves the stability of reflective material detection under dynamic conditions. For example, the continuity of reflective stripes in partially occluded areas in the fused image is enhanced to avoid misclassification due to abnormal illumination of a single frame.
[0059] Step S30: input the enhanced sample image into the initial recognition model, and output the predicted wearing state features.
[0060] The initial recognition model is a deep convolutional neural network architecture that has not completed parameter optimization. The predicted wearing state features refer to the multi-dimensional feature vector before the last fully connected layer of the model. Exemplarily, the model can use the ResNet-50 backbone network, remove the top-level classifier and retain the 2048-dimensional feature vector as the output. For example, the input size is 256×256 pixels. After the spatial features are extracted step by step by the convolution layer and the residual block, a 1×1×2048 feature tensor is finally generated. It should be noted that the model randomly initializes the weight parameters at the beginning of training, and maps the enhanced sample image to the high-dimensional feature space through forward propagation. There is a large difference between its feature distribution and the actual wearing state label, which needs to be gradually optimized through subsequent loss calculation and back propagation.
[0061] Step S40: calculating the classification loss between the predicted wearing state feature and the wearing state label, and updating the parameters of the initial recognition model based on gradient back propagation until the classification loss converges.
[0062] For example, the classification loss is the difference between the model prediction probability distribution and the true label distribution measured by the cross entropy function, and the gradient back propagation refers to updating the parameters according to the partial derivatives of the loss function to the weight parameters of each layer to minimize the loss. Specifically, the 2048-dimensional predicted wearing state features are input into the fully connected classification layer to generate the logical output values of the three types of wearing states. After being converted into probability distributions by the Softmax function, the cross entropy loss between them and the one-hot encoded labels is calculated. For example, the true label of a sample is normal wear (encoded [1,0,0]), and the model output probability is [0.7, 0.2, 0.1], then the cross entropy loss is -log(0.7)=0.356. The Adam optimizer is used to update the parameters with an initial learning rate of 0.001. When the loss of the training set decreases by less than 0.1% for 10 consecutive epochs, it is considered to have converged. This process gradually aligns the model feature space with the wearing state semantics through iterative optimization. For example, the feature vectors of the normal wear sample form a tight cluster in the embedding space, and maintain a clear distance from the feature clusters of the non-wearing sample.
[0063] Step S50: performing end-to-end joint training on the converged initial recognition model and the multi-branch feature extraction network, fixing the parameters of the multi-branch feature extraction network, and optimizing the regional distribution analysis layer parameters of the initial recognition model.
[0064] End-to-end joint training is to connect the multi-branch feature extraction network and the initial recognition model into a unified calculation graph for overall optimization. Fixed parameters refer to prohibiting the weight update of the specified network layer during the back-propagation process. Specifically, the multi-branch feature extraction network has completed parameter convergence through the reflective vest regional image feature extraction task in the early training. During the joint training, its convolution kernel weights are frozen, and only the self-attention parameters of the regional distribution parsing layer (i.e., the Transformer encoder layer) in the initial recognition model are allowed to be adjusted. For example, in each batch of training, the initial wearing state features output by the multi-branch network are directly input into the initial recognition model, and the projection weights of its query and key-value matrices are optimized by calculating the wearing state classification loss and back-propagating to the multi-head attention module of the regional distribution parsing layer. This strategy ensures that the model retains the underlying feature extraction capability while enhancing the discriminability of high-level semantic parsing, for example, making the regional distribution parsing layer pay more attention to the spatial distribution pattern of the reflective material coverage rather than the basic texture details.
[0065] As an implementation manner, the step S300, based on the pre-trained wearing state recognition model, performs regional distribution analysis on the initial wearing state feature to generate the target wearing state feature, includes: Step S310: Divide the initial wearing state feature into a plurality of spatial grid units, and calculate the similarity between the feature vector of each spatial grid unit and a preset reflective feature template.
[0066] The spatial grid unit is a set of rectangular sub-areas that evenly divide the initial wearing state feature map along the height and width directions. The preset reflective feature template refers to a set of reference feature vectors extracted from the standard reflective vest sample. Specifically, the size of the initial wearing state feature map is 32×64×512 (height×width×number of channels), which is divided into 32 spatial grid units of 8×8 pixels according to a 4×8 grid, and each unit corresponds to a 512-dimensional feature vector. The preset reflective feature template contains typical features of reflective stripes in a normal wearing state, such as the mean vector extracted from 1,000 standard samples. When calculating the similarity, the cosine similarity is used to measure the directional consistency between the feature vector of each grid unit and the template vector. For example, the cosine similarity between the feature vector of a certain grid unit and the template is 0.92, indicating that the reflective characteristics of the area are highly consistent with the standard; while the similarity of another grid unit is only 0.35, indicating that there may be interference from occlusion or non-reflective materials. It should be noted that the grid division density needs to be adapted to the resolution of the reflective vest area image to ensure that each grid covers an area of approximately 5cm×5cm in actual physical size to match the minimum detectable width of the reflective stripes.
[0067] Step S320: generating a spatial weight matrix according to the distribution of the similarities, wherein the spatial weight matrix is used to identify the distribution probability of the effective reflective area in the reflective vest area image.
[0068] The spatial weight matrix is a two-dimensional matrix with the same dimension as the spatial grid unit, and its element value is the weighted result of the similarity of the corresponding grid unit. The probability of effective reflective area distribution refers to the confidence that each grid unit belongs to a compliant reflective material. Specifically, the similarities of the 32 grid units calculated in step S310 are arranged into a 4×8 matrix according to the original spatial position, and spatial smoothing is performed through Gaussian filtering to eliminate the influence of local outliers. For example, the similarity of a central grid is 0.88, and its adjacent grids are 0.75, 0.82, and 0.90. After Gaussian filtering with σ=1.5, the center value is adjusted to 0.85, which is more in line with the continuous distribution characteristics of actual reflective stripes. Further, the smoothed matrix is normalized so that the weight value range is mapped to the 0-1 interval, and the grids with weights ≥0.7 are determined to be effective reflective areas. This method can accurately identify the core reflective area of the reflective vest. For example, in a partially blocked scene, the weight of the right grid covered by the toolkit drops below 0.3, while the unblocked area on the left maintains a high weight of more than 0.8.
[0069] Step S330: multiplying the spatial weight matrix by the initial wearing state feature element by element to obtain a weighted intermediate feature.
[0070] Element-by-element multiplication refers to multiplying each element value of the spatial weight matrix by all channel eigenvalues of the corresponding spatial grid unit to suppress the response intensity of the low-weight area. In specific implementation, the 4×8 spatial weight matrix is upsampled to 32×64 size to align it with the spatial dimension of the initial wear state feature map, and then broadcast in the channel dimension for matrix dot multiplication. For example, the weight of a grid in the lower right corner is 0.2, and the eigenvalues of the corresponding 8×8 area on 512 channels are all multiplied by 0.2, which greatly reduces the feature contribution of the area blocked by the safety rope. This operation strengthens the discriminative features of the effective reflective area. For example, in the normal wearing state, the high-weight grid (0.9-0.95) arranged vertically in the center has more significant edge gradient features after weighting, which is helpful for the subsequent analysis of the structural integrity of the reflective stripes.
[0071] Step S340: performing non-maximum suppression processing on the intermediate features, removing redundant feature regions with an overlap rate higher than a preset threshold, and generating the target wearing state features.
[0072] The non-maximum suppression process is to retain the local maximum response area in the feature space, while suppressing the low confidence area that highly overlaps with it. The preset overlap rate threshold is set to 0.5 based on the minimum spacing of the reflective stripes. Specifically, the sliding window method is used to traverse the intermediate feature map, and only the grid with the maximum weight value is retained in each 8×8 window. If the overlapping area of adjacent windows accounts for more than 50%, the grid with lower weight is eliminated. For example, the weights of two horizontally adjacent grids are 0.92 and 0.87 respectively, and the overlap rate of their coverage areas reaches 60%, then the latter is suppressed and the former is retained. After this processing, the target wearing state feature will eliminate the multi-peak interference caused by wrinkles or reflections of the reflective material, and generate a compact and non-redundant reflective area representation. For example, the 5 discrete high-weight areas initially detected are merged into 2 core areas after suppression, which accurately reflects the wearing state of the main part of the reflective vest.
[0073] As an implementation manner, the step S400, determining the reflective vest wearing state recognition result of the human object to be detected according to the matching degree between the target wearing state feature and the preset wearing state threshold, includes: Step S410: extracting a first statistic related to the coverage area of the reflective material, a second statistic related to the continuity of the reflective stripes, and a third statistic related to the matching degree of the human body posture from the target wearing state characteristics.
[0074] Exemplarily, the first statistic is the percentage of the effective reflective area to the total area of the target wearing state feature, calculated as the ratio of the number of high-weighted grids to the total number of grids. For example, if 24 of the 32 grids have a weight ≥ 0.7, the coverage area is 75%. The second statistic is obtained by calculating the spatial connectivity of the high-weighted grids, merging adjacent grids into continuous regions using a morphological expansion operation, and counting the number of grids in the largest connected region. For example, 20 of the 24 high-weighted grids form a single connected region with a continuity of 83.3%. The third statistic is obtained by comparing the distribution of the reflective area with the matching degree of the standard human posture template, which defines the standard position of the reflective stripes on the front of the torso (such as left-right symmetrical distribution, longitudinal coverage from the shoulder to the waist). Specifically, the average Euclidean distance between the center point of the actual reflective area and the template reference point is calculated. When the distance is ≤5 pixels, the matching degree is 1.0, and it decreases linearly with the distance. For example, the center of the reflective area of a sample deviates from the standard position by 8 pixels, and the matching degree drops to 0.6.
[0075] Step S420: input the first statistic, the second statistic and the third statistic into a preset multi-condition decision tree model to generate a comprehensive matching score.
[0076] The multi-conditional decision tree model is a hierarchical classifier trained based on historical data. The comprehensive matching score is the weighted evaluation value of the three types of statistics, with a value range of 0-1. Specifically, the first statistic is assigned a weight of 0.5, the second statistic is 0.3, and the third statistic is 0.2. The scores are gradually accumulated through the hierarchical judgment of the decision tree. For example, the first statistic of a sample of 0.75 triggers the "≥70%" branch plus 0.375 points, the second statistic of 0.8 triggers the "≥80%" branch plus 0.24 points, and the third statistic of 0.7 triggers the "≥60%" branch plus 0.14 points, with a total score of 0.755. It should be noted that the decision condition threshold is dynamically set through cluster analysis. For example, the cluster center of the first statistic of the normal wear class sample is 0.85, the partially occluded class is 0.55, and the unworn class is 0.15, and the decision boundary is divided accordingly.
[0077] Step S430: If the comprehensive matching score is higher than the first threshold, it is determined to be in a normal wearing state.
[0078] The first threshold is set to 0.75 based on the minimum comprehensive score of the normal wear samples. For example, a construction worker wears a reflective vest correctly, with a coverage area of 85%, a continuity of 90%, a posture matching of 0.95, and a comprehensive score of 0.85*0.5+0.9*0.3+0.95*0.2=0.86, which is considered normal if it exceeds the threshold. This threshold ensures that correct classification can be maintained even when there is a slight change in lighting or posture deviation (such as a decrease in the weight of the local grid due to arm lifting). For example, a sample with a score of 0.76 can still be judged as compliant even though the third statistic is 0.7 due to the shooting angle.
[0079] Step S440: If the comprehensive matching score is lower than the second threshold, it is determined to be in a not-worn state.
[0080] The second threshold is set to 0.3 based on the highest score of the samples not wearing. For example, a worker only wears ordinary work clothes, with reflective material coverage of 5%, continuity of 0%, and posture matching of 0.1. The comprehensive score is 0.05*0.5+0*0.3+0.1*0.2=0.045, which is far below the threshold and is judged as not wearing. This threshold needs to cover the situation where the reflective vest is completely missing or extremely damaged. For example, a sample with a coverage of 10% but a continuity of 0 and a score of 0.12 is still classified as not wearing.
[0081] Step S450: If the comprehensive matching score is between the first threshold and the second threshold, it is determined to be a partial occlusion state.
[0082] The partial occlusion state is when the reflective vest is partially covered by external objects but not completely ineffective. For example, the right half of a worker's reflective vest is blocked by a tool bag, with a coverage area of 55%, a continuity of 60% (the left half remains connected), and a posture matching degree of 0.5. The comprehensive score is 0.55*0.5+0.6*0.3+0.5*0.2=0.275+0.18+0.1=0.555, which is between 0.3-0.75, and is judged as partially occluded. It should be noted that this interval needs to exclude misjudgments near the boundary. For example, although the sample with a score of 0.74 is close to the normal threshold, it still maintains a partial occlusion judgment due to a continuity of 85% and a posture matching degree of 0.8, avoiding misjudgment of severe occlusion as normal.
[0083] As an implementation method, the construction process of the multi-condition decision tree model includes the following steps: Step S401: collecting historical identification data, wherein the historical identification data includes multiple groups of the first statistics, the second statistics, the third statistics and corresponding real wearing state labels.
[0084] Historical identification data is the wear status identification records accumulated from the deployed system, and needs to cover all lighting conditions and wear status combinations. For example, 10,000 records are collected, of which 6,000 are normally worn (the first statistic has a mean of 0.82 and a standard deviation of 0.08), 3,000 are partially blocked (mean of 0.52 and a standard deviation of 0.12), and 1,000 are not worn (mean of 0.18 and a standard deviation of 0.10). Each piece of data contains three floating-point values of statistics and a label confirmed by manual review. The data needs to be deduplicated and balanced, such as performing SMOTE oversampling on the minority class samples (not worn) to make the ratio of the number of samples of the three categories 1:1:1 to avoid model skew.
[0085] Step S402: performing cluster analysis on the historical recognition data to determine the distribution boundaries of different wearing status categories in the three-dimensional statistical space.
[0086] Cluster analysis uses the K-means algorithm to divide the data into three clusters, corresponding to three wearing states. Specifically, the three-dimensional space coordinate axes are the first statistic (X-axis), the second statistic (Y-axis), and the third statistic (Z-axis), and the sample similarity is measured by the Euclidean distance. For example, the center coordinates of the normal wear cluster are (0.85, 0.88, 0.90), the partially occluded cluster is (0.55, 0.60, 0.50), and the unworn cluster is (0.15, 0.10, 0.12). The distribution boundary is determined by calculating the decision hyperplane of the adjacent areas between clusters. For example, the boundary plane equation of the normal and partially occluded clusters is 0.6X+0.5Y+0.4Z=0.72. When the sample point is substituted into the left value ≥0.72, it is classified as normal, otherwise it is partially occluded.
[0087] Step S403: generating a plurality of decision rules according to the distribution boundary, each decision rule corresponding to a hyperplane segmentation condition.
[0088] The decision rule divides the three-dimensional space into multiple sub-regions, each of which is associated with a specific wearing state. For example, the first rule is "if the first statistic ≥ 0.7 and the second statistic ≥ 0.75 and the third statistic ≥ 0.8, then it is determined to be normal wear"; the second rule is "if 0.4 ≤ the first statistic < 0.7 and the second statistic ≥ 0.5 and the third statistic ≥ 0.4, then it is determined to be partially occluded". The hyperplane equation is obtained by linearly segmenting the inter-cluster samples through the support vector machine (SVM). For example, the segmentation hyperplane of the normal and partially occluded classes is trained by 1000 boundary samples, and the classification interval is maximized to ensure generalization ability.
[0089] Step S404: Optimize the order of the decision rules based on the Gini coefficient and construct hierarchical decision nodes until each leaf node contains only samples of a single wearing status category.
[0090] The Gini coefficient is used to measure the degree to which the decision rule reduces the data impurity. The optimization goal is to select the rule that can maximize the purity of the subsequent branches and execute it first. For example, the initial node contains all the data, and the Gini coefficient is 0.6 (three-class mixture). First, the "first statistic ≥ 0.7" rule is applied to divide the data into two subsets: the left subset Gini coefficient is 0.2 (normal class accounts for 90%), and the right subset is 0.5 (partial occlusion and non-wearing mixture). Continue to apply the "second statistic < 0.4" rule to the right subset, and the Gini coefficient drops to 0.1 (non-wearing class accounts for 95%). The final decision tree contains 5 layers of nodes, and the sample category purity of each leaf node is ≥ 95%. This structure ensures efficient classification. For example, normal wearing samples can reach the leaf node after an average of 2 judgments, while complex partially occluded samples require 4 judgments to exclude the possibility of non-wearing and normal.
[0091] As an implementation manner, the method further includes a dynamic correction process for the recognition result, which may specifically include the following steps: Step S500: When the recognition result is a partial occlusion state, a plurality of continuous frames of historical monitoring images of the human object to be detected are obtained.
[0092] Continuous multi-frame historical monitoring images refer to at least 10 frames of image sequences collected at fixed time intervals before the recognition result is triggered. The time window length is set to 5 seconds based on the human body movement frequency, and the frame rate is 2 frames per second to ensure time continuity. Specifically, in the construction site scene, a worker's reflective vest right shoulder is judged to be partially blocked due to temporary carrying of a tool bag. The system immediately retrieves 10 frames of images collected within the previous 5 seconds (time stamps are t-4.5s to t-0.5s), and each frame of the image is pre-processed by human body detection and reflective vest area alignment. It should be noted that historical images must meet the spatiotemporal alignment conditions, that is, the pixel offset caused by human body movement is eliminated by optical flow method or feature point matching. For example, the SIFT algorithm is used to extract the human body key points in each frame, and the reflective vest area of each frame is mapped to a unified coordinate system through affine transformation to ensure the spatial consistency of temporal feature extraction. If the alignment fails due to violent movement (such as translation of more than 50 pixels or rotation of more than 15 degrees), the frame is discarded and traced back to the image frame that meets the alignment conditions until the minimum frame number threshold is collected.
[0093] Step S600: extracting time series features from the historical monitoring images to generate a dynamic change trend of the wearing status of the reflective vest.
[0094] Temporal feature extraction is the process of quantifying the evolution of the area, position and morphological parameters of the partially occluded area from the multi-frame aligned images. The dynamic change trend is fitted by linear regression or sliding average algorithm to the change slope and fluctuation amplitude of the time series data. Specifically, the pixel area ratio of the partially occluded area is calculated frame by frame (such as 15% of the occlusion ratio in the t-4.5s frame, 18% in the t-3.5s frame, ..., 35% in the t-0.5s frame), and the moving trajectory of the centroid of the occluded area is recorded (such as the X coordinate increases from 120 pixels to 150 pixels, and the Y coordinate remains unchanged at 80 pixels). The area change curve is fitted by the least squares method, and the slope is +4% / s, which is judged as a continuous expansion trend; if the slope fluctuates within ±1% / s, it is considered to be a stable state. Further analysis of morphological parameters, such as the aspect ratio of the occluded area gradually changes from 2.1 (vertical stripes) to 1.3 (approximately circular), indicating that the toolkit has changed its occlusion morphology due to human body shaking. This dynamic trend is encoded into three sets of time series vectors: area change rate, center of mass displacement rate, and morphological distortion index, which are input into the preset trend classifier for state evolution pattern recognition.
[0095] As an implementation manner, the step S600, extracting time series features from the historical monitoring image to generate a dynamic change trend of the wearing state of the reflective vest, includes: Step S610: Track key points of the reflective vest area in multiple consecutive frames of historical monitoring images to determine the motion trajectory of the blocked area.
[0096] Key point tracking refers to tracking the displacement trajectory of specific spatial points in the reflective vest area in the time series image through a feature matching algorithm. The motion trajectory of the occluded area is composed of the dynamic coverage range change path of the occluded reflective material. Specifically, in the construction site monitoring scene, the right shoulder of a worker's reflective vest is gradually covered by a moving safety rope, and the system collects 10 consecutive frames of images at a rate of 5 frames per second (time span 2 seconds). First, the intersection of high-reflective stripes in the reflective vest area (such as the intersection of horizontal and vertical reflective stripes) and the center point of the reflective particle dense area are located in the first frame image as the initial key point set, and a total of 32 key points are marked. Subsequently, the KLT (Kanade-Lucas-Tomasi) optical flow algorithm is used to calculate the pixel displacement vector of the key point in each pair of adjacent frames. For example, between the first frame and the second frame, the key point numbered K15 moves from the coordinates (120,80) to (122,81), and the displacement vector is (+2,+1). By accumulating the displacement vectors of consecutive frames, the motion trajectory curve of the key points is constructed. For example, the trajectory of K15 in 10 frames shows a linear movement pattern from the lower left to the upper right. It should be noted that the key point screening needs to exclude instantaneous noise points caused by changes in the reflective properties of reflective materials. For example, when a key point jumps by more than 20 pixels due to overexposure in the third frame, it needs to be corrected or eliminated through the trajectory smoothing algorithm. This process can accurately depict the expansion direction and speed of the occluded area (such as the safety rope) in the time dimension, providing basic spatial motion data for subsequent trend analysis.
[0097] As an implementation manner, the step S610 of tracking key points of the reflective vest area in the continuous multiple frames of historical monitoring images to determine the motion trajectory of the blocked area includes: Step S611: Based on the reflective vest area image of the current frame, a set of key points related to the geometric distribution of reflective stripes is detected, where the key point set includes the intersection points of the reflective material edges and the center points of the high reflectivity areas.
[0098] The key points related to the geometric distribution of reflective stripes refer to the feature points with fixed spatial relationships in the standard design of reflective vests, such as the intersection of horizontal stripes and vertical stripes, the maximum curvature point of the end of the reflective stripes, and the centroid of the high-reflective particle aggregation area. Specifically, the Harris corner detection algorithm is used to calculate the eigenvalues of the pixel grayscale gradient matrix in the reflective vest area image of the current frame, and the pixels whose response values exceed the threshold of 0.05 are marked as candidate key points. For example, in the reflective vest area with a size of 256×128 pixels, 48 candidate points are detected, and 32 key points with uniform spatial distribution are retained after non-maximum suppression, including the intersection of the third horizontal and second vertical reflective stripes (coordinates (80,60)), the centroid of the reflective particle area in the lower right corner (coordinates (220,100)), etc. It should be noted that key point detection needs to adapt to different lighting conditions. For example, in low-light environments, an adaptive threshold adjustment strategy is used to dynamically reduce the Harris threshold to 0.03 to maintain detection sensitivity. At the same time, morphological closing operations are used to fill in the broken parts of the reflective strips to ensure the integrity of the geometric features.
[0099] Step S612: performing optical flow estimation on the key point set and the corresponding reflective vest area in the previous frame of historical monitoring image to obtain the displacement vector of each key point between adjacent frames.
[0100] Optical flow estimation refers to the calculation of the pixel-level displacement of key points between adjacent frames based on the assumption of constant brightness and small motion. The displacement vector contains horizontal and vertical components. Specifically, the pyramid Lucas-Kanade algorithm is used to construct a three-layer image pyramid (scaling factor 0.5), and iterative optical flow calculation is performed at each layer to improve the accuracy of large displacement tracking. For example, the coordinates of key point K09 in the previous frame are (150,90). After pyramid calculation, it is found that it moves to (155,88) in the current frame, and the displacement vector is (+5,-2). For the center point of the high reflectivity area, due to the reflective characteristics that may cause sudden brightness changes, it is necessary to introduce reverse optical flow verification: that is, reverse tracking from the current frame to the previous frame. If the bidirectional tracking error exceeds 1 pixel, it is determined to be a failure point and eliminated. For example, the forward optical flow displacement of a key point is (+6,+3), and the reverse tracking error is (+1,-1), which meets the error tolerance and the displacement vector is valid; if the error reaches (+4,-2), it is considered a mismatch and is excluded.
[0101] Step S613: Filter out matching key point pairs according to the displacement vectors, and construct initial motion trajectory segments based on the coordinate offsets of the matching key point pairs.
[0102] Matching key point pairs refers to point pairs that are successfully tracked between two frames and whose displacements conform to rigid motion constraints. The initial motion trajectory fragment consists of the displacement sequence of the same key point in consecutive frames. Specifically, the optical flow estimation results are screened by RANSAC (random sampling consistency), the global motion model (such as affine transformation) is fitted, and abnormal points that deviate from the model by more than 2 pixels are removed. For example, among the 50 candidate displacement vectors, 42 conform to the affine transformation model (translation + rotation), and the remaining 8 are removed due to abnormal displacement caused by local occlusion. Subsequently, the retained key points are linked in chronological order to form the initial trajectory fragment: the displacement sequence of key point K22 in frames 1-5 is [(0,0), (+2,+1), (+3,+2), (+5,+3), (+7,+4)], showing a uniform motion pattern in the lower right corner. This process can effectively suppress trajectory breaks caused by reflective flashes or temporary occlusions. For example, a key point is lost in frame 3 due to short-term overexposure, and its trajectory fragment is completed by interpolation algorithm.
[0103] Step S614: performing continuity check on the initial motion trajectory segments, removing abnormal trajectory points whose displacement directions do not conform to the human body motion posture constraints, and generating a corrected candidate motion trajectory.
[0104] Continuity verification refers to verifying whether the displacement direction of the trajectory conforms to the movement rules of the trunk, arms and other parts based on the human joint kinematic model. Specifically, a human posture constraint rule library is constructed: the horizontal movement speed of the shoulder is ≤5 pixels / frame, the vertical movement speed is ≤3 pixels / frame; the waist rotation angle change rate is ≤2° / frame. For example, a trajectory segment has a displacement of (+15, +0) between frames 4-5, which exceeds the horizontal speed threshold of the shoulder and is judged to be abnormal (possibly caused by reflective artifacts) and needs to be removed from the candidate trajectory. The corrected candidate motion trajectory must meet the motion smoothness conditions: the angle change of the displacement vectors of adjacent frames is ≤30°, and the acceleration is ≤5 pixels / frame². For example, the displacement of the corrected K22 trajectory in frame 3 is adjusted from (+3, +2) to (+4, +2), so that the angle between frames 2-3 is reduced from 18° to 12°, which meets the smoothness requirements.
[0105] Step S615: extracting the motion direction consistency parameter of the occluded area and the area change rate of the track coverage area according to the spatial distribution density of the candidate motion track in multiple consecutive frames.
[0106] The motion direction consistency parameter refers to the average angular variance of all candidate trajectory displacement vectors, reflecting the concentration of the overall moving direction of the occluded area; the rate of change of the area covered by the trajectory is obtained by calculating the time derivative of the convex hull area of the trajectory points. Specifically, in 10 frames, the displacement angle of 20 candidate trajectories has an average value of 85°, a variance of 5°, and a direction consistency parameter of 0.92 (1-variance / 90°); the convex hull area of the trajectory points increases from 200 pixels² in frame 1 to 800 pixels² in frame 10, and the area change rate is (600 / 500) / 2=0.6 / s. This parameter can distinguish different types of occlusion: the sliding of the tool bag leads to high directional consistency (>0.9) and continuous area growth, while the wind blowing the hem may cause a directional variance of >20° and area fluctuations.
[0107] Step S616: Based on the motion direction consistency parameter and the area change rate of the track coverage area, fit the motion trend curve of the occluded area, and determine the starting position, ending position and path shape of the motion track.
[0108] The motion trend curve refers to the spatiotemporal distribution function of the trajectory points fitted by polynomial regression or spline interpolation, and the path shape is characterized by the curvature and extension direction of the curve. Specifically, a cubic polynomial is used to fit the displacement data of 20 candidate trajectories, and the parametric equations x(t)=0.5t²+2t, y(t)=0.2t³-0.1t²+3t, R²=0.98 are obtained, indicating that the occlusion area extends along an approximate parabola. The starting position is the mean of the trajectory points of frame 1 (120±5, 80±3), and the ending position is (250±8, 110±5) of frame 10. The path shape is detected as a straight line (curvature <0.01) by Hough transform. This fitting result is used to predict future movement: if the trend curve shows a continuous extension to the core area of the reflective vest (such as the center of the chest), an early warning is triggered; if the path turns to the edge and the curvature increases, it may be a temporary occlusion.
[0109] Step S620: Calculate the area change rate and shape similarity of the occluded region between adjacent frames.
[0110] The area change rate of adjacent frames refers to the relative increase or decrease ratio of the pixel area of the occluded area per unit time. The shape similarity measures the morphological consistency of the occluded area contour between consecutive frames through the Hausdorff distance. Specifically, in the t-th frame and the t+1 frame, the binary mask of the occluded area determined by key point tracking is extracted respectively, and the pixel area difference is calculated as ΔA=|A_{t+1}-A_t|, and the change rate R=ΔA / ((A_t+A_{t+1}) / 2) / Δt, where Δt is the frame interval time of 0.2 seconds. For example, the area of a certain occluded area increases from 1500 pixels in the third frame to 1800 pixels in the fourth frame, and the change rate R=(300 / 1650) / 0.2≈0.909 / s, indicating that the occlusion is rapidly expanding. When calculating shape similarity, the mask contours of the two frames are discretized into a sequence of polygonal vertices, and the Hausdorff distance of the maximum and minimum vertex spacing in both directions is calculated. For example, the contour distance between the 5th and 6th frames is 8 pixels, and the normalized similarity is S=1-(8 / 256)=0.969 (image size 256×256). This indicator effectively distinguishes the stability of the occlusion form, such as the strip occlusion caused by the sliding of the safety rope and the triangular occlusion formed by the temporary wind blowing the corner of the clothes. The shape similarity of the former is higher than 0.9, while the latter may be lower than 0.7.
[0111] Step S630: constructing a time series feature vector according to the motion trajectory, area change rate and shape similarity.
[0112] The time series feature vector refers to the encoding of multidimensional time series data into a numerical vector of fixed dimension, which is used to characterize the pattern characteristics of occlusion evolution. Specifically, within a 2-second time window, the motion trajectory direction angle, area change rate and shape similarity are collected every 0.2 seconds, and a total of 10 sets of original data are generated. First, the motion trajectory direction angle is filtered by sliding window average (window size 3 frames) to eliminate instantaneous jitter noise. For example, the direction angle sequence of the 2nd to 4th frames [85°, 88°, 82°] is filtered and output as [85°, 85°, 83°]. Subsequently, the area change rate is arranged in chronological order as [0.2, 0.5, 0.9, 1.1, 0.8,...], and its variance value in the first-order difference sequence is calculated to reflect the fluctuation intensity of the change rate. The shape similarity sequence is extracted by fast Fourier transform to extract the main frequency component. For example, a periodic fluctuation of 0.5Hz is detected, indicating that the occlusion morphology is affected by regular external forces (such as equipment vibration). Finally, 12 statistics such as the azimuth mean, area change rate variance, and main frequency amplitude are aligned by timestamp to generate a 10×12-dimensional feature matrix, which is then normalized to the [0,1] interval to form a 120-dimensional time series feature vector. This vector can comprehensively describe the trend, volatility, and morphological laws of the occlusion movement. For example, the continuously expanding occlusion is characterized by a stable azimuth, a positive accumulation of area change rate, and a high degree of shape similarity.
[0113] As an implementation manner, the step S630, constructing a time series feature vector according to the motion trajectory, area change rate and shape similarity, includes: Step S631: Segmentally sample the direction change pattern of the motion trajectory, extract the direction angle offset sequence in each time window, and generate direction change pattern parameters.
[0114] The direction change pattern parameters refer to the mean, variance and autocorrelation characteristics of the trajectory direction angle in the sliding time window to quantify the regularity of the movement. Specifically, the window size is set to 3 frames (0.6 seconds), the step size is 1 frame, and the 10-frame direction angle sequence [82°, 85°, 83°, 87°, 84°, 88°, 85°, 86°, 84°, 87°] is segmented and sampled to obtain 8 window data. For example, the direction angle of windows 1-3 has a mean of 83.3°, a variance of 1.25, and an autocorrelation coefficient of 0.78; the mean of windows 2-4 is 85°, the variance is 2.0, and the autocorrelation coefficient is 0.65. The three statistics of each window are spliced to generate a 24-dimensional directional feature subvector to reflect the stability of the direction change. For example, the window autocorrelation coefficient of continuous linear motion is greater than 0.7, while the autocorrelation coefficient of random fluctuations is less than 0.3.
[0115] Step S632: Based on the time series data of the area change rate, the relative change gradient between adjacent time windows is calculated, and the area increase and decrease trends at different time scales are accumulated to generate the accumulated area change.
[0116] The cumulative change in area refers to the integrated area change rate at multiple time scales (such as 1 second for short time, 2 seconds for medium time, and 3 seconds for long time), which characterizes the persistence of occlusion expansion. Specifically, for the area change rate sequence [0.2, 0.5, 0.9, 1.1, 0.8, 1.2, 1.3, 1.4, 1.5, 1.6] / s, the 1-second scale cumulative amount is calculated: the first 5 frames (1 second) accumulate 0.2+0.5+0.9+1.1+0.8=3.5, and the next 5 frames accumulate 1.2+1.3+1.4+1.5+1.6=7.0, an increase of 100%. At the same time, the gradients of adjacent windows are calculated: the gradient of window 1-2 is (0.5-0.2) / 0.2=1.5, and the gradient of window 2-3 is (0.9-0.5) / 0.2=2.0, reflecting the increase in speed. This parameter distinguishes short-term fluctuations (the gradient alternates between positive and negative) from sustained growth (the gradient increases monotonically).
[0117] Step S633: performing frequency domain conversion on the continuous fluctuation of the shape similarity, extracting the energy distribution ratio of the low-frequency component and the high-frequency component, and generating a shape stability index.
[0118] Exemplarily, the shape stability index decomposes the shape similarity time series signal into a spectrum through fast Fourier transform, and calculates the ratio of 0-0.5Hz low-frequency energy to 0.5-2Hz high-frequency energy. For example, after FFT, the 10-frame shape similarity sequence [0.95, 0.94, 0.93, 0.92, 0.91, 0.90, 0.89, 0.88, 0.87, 0.86] has a low-frequency energy share of 85% and a high-frequency share of 15%. The index value is 85 / 15≈5.67, indicating that the morphology changes slowly and stably; if the sequence contains [0.7, 0.9, 0.6, 0.8, 0.5, ...], the high-frequency energy share is greater than 40%, and the index value is less than 1.5, indicating that the morphology fluctuates violently.
[0119] Step S634: align the direction change mode parameters, area cumulative change and shape stability index according to the timestamp, and perform normalization processing to generate a standardized time series data block.
[0120] Standardizing the time series data block means aligning the multi-dimensional heterogeneous data according to a unified time reference and eliminating the dimension difference through Z-score normalization. Specifically, at the timestamp t=2.0s, the directional feature subvector is [83.3°, 1.25, 0.78], the area accumulation is 3.5, and the shape stability index is 5.67, which are normalized to [0.75, 0.12, 0.81], 0.62, and 0.92 (based on the global maximum value), respectively. After standardization, the data is arranged in chronological order into a 10×5 matrix, and missing values (such as optical flow tracking failure in a certain frame) are filled by linear interpolation or the average of adjacent frames.
[0121] Step S635: splicing the direction change pattern parameters, area cumulative change amounts and shape stability indicators corresponding to the same timestamp in the standardized time series data block to form a multi-dimensional feature segment.
[0122] The multidimensional feature fragment is the vertical concatenation vector of all features at each timestamp. For example, at t=2.0s, the direction parameter [0.75, 0.12, 0.81], the area accumulation 0.62, and the shape index 0.92 are concatenated into a 5-dimensional vector [0.75, 0.12, 0.81, 0.62, 0.92]. The vectors of 10 consecutive timestamps are stacked vertically into a 10×5 matrix to form the basic feature fragment.
[0123] Step S636: Perform sliding window aggregation on the multi-dimensional feature segments of continuous timestamps, calculate the mean, variance and maximum value of each dimension in the window, and generate the time series feature vector.
[0124] Sliding window aggregation refers to traversing the basic feature fragments with a window size of 3 and a step size of 1, and calculating the statistics of each feature in each window. For example, the directional mean parameter of windows 1-3 is (0.75+0.68+0.72) / 3≈0.72, with a variance of 0.03; the maximum area accumulation is 0.65. The statistics of the three windows are concatenated into (0.72, 0.03, 0.65, ...) × 3 windows = 45 dimensions, and finally a 45×3 = 135-dimensional time series feature vector is generated. This vector enhances the model's ability to identify trend patterns by capturing the time domain statistical characteristics of the features, for example, the variance of the continuously increasing area accumulation approaches 0, while the variance of the fluctuating state increases significantly.
[0125] Step S640: input the time series feature vector into the long short-term memory network, predict the evolution path of the occluded area in the next several frames, and generate the dynamic change trend according to the evolution path.
[0126] Long short-term memory network refers to a recurrent neural network structure with a gating mechanism, which is used to model long-term dependencies in time series data. The evolution path prediction iteratively outputs the occlusion parameters of the future time step through autoregression. Specifically, a network model containing 3 layers of LSTM units is constructed. The input layer receives a 120-dimensional feature vector, the hidden layer dimension is 256, and the output layer generates the direction angle, area change rate and shape similarity prediction values for the next 5 frames (1 second). During training, 10,000 groups of time series samples in historical data are used, and the network parameters are optimized by the mean square error loss function. For example, when the feature vector of a safety rope that continues to slide is input, the network outputs the area change rate prediction for the next 5 frames as [1.2, 1.3, 1.4, 1.5, 1.6] / s, the direction angle is maintained at 82°±2°, and the shape similarity is maintained above 0.91. Based on this, a dynamic trend curve of linear expansion of the occlusion area along a fixed direction is generated. This trend will be used to determine whether the recognition result needs to be corrected: if the predicted area exceeds the 60% threshold within 3 seconds and the direction points to the core reflective area, the correction of the unworn state will be triggered; if the predicted area falls below 10% and the shape returns to stability, it will be corrected to the normal worn state.
[0127] Step S700: If the dynamic change trend indicates that the blocked area continues to expand, the recognition result is corrected to the not-worn state.
[0128] Exemplarily, continuous expansion means that the area of the occluded area increases monotonically within the time window and the average growth rate exceeds the preset threshold (such as ≥3% / s), accompanied by the movement of the center of mass toward the core area of the reflective vest (such as moving more than 20 pixels from the edge to the center). For example, a worker's reflective vest gradually slips off within 5 seconds due to improper wearing. The occluded area increases linearly from 10% to 60%, and the center of mass moves from (100,80) to (160,90). The growth rate reaches 10% / s and the moving direction is opposite to the normal coverage area of the vest. At this time, the system determines that the occlusion is irreversible and the overall reflection function fails, and corrects the partial occlusion state to the unworn state, triggering a secondary alarm signal. It should be noted that the correction logic needs to exclude temporary occlusion interference (such as waving action briefly covering the reflective strip), and the judgment robustness is improved by setting a minimum frame limit for the continuous growth rate (such as 4 consecutive frames of growth rate ≥2% / s) to avoid false corrections.
[0129] Step S800: If the dynamic change trend indicates that the blocked area returns to the normal range within the preset time, the recognition result is corrected to the normal wearing state.
[0130] The preset time is set to 8 seconds based on the reasonable duration of the human body's adjustment action. The normal range refers to the occlusion area ≤10% and the center of mass is located in the non-core area of the edge of the reflective vest (such as ≤15 pixels from the boundary). Specifically, a worker actively adjusted to the side waist position after carrying the tool bag for 3 seconds, causing the occlusion area to drop from the peak of 40% to 8% in the next 5 frames, and the center of mass retreated from (140,85) to (180,120) (the boundary coordinates of the lower right corner of the reflective vest are (200,130)). The system detects that the area regression rate reaches -6.4% / s and the displacement direction of the center of mass deviates from the core area. The dynamic trend classifier outputs a "return to normal" signal, triggering the correction of the recognition result. It should be noted that the recovery judgment must meet two conditions: first, the difference between the occlusion area of the final frame and the initial normal state does not exceed 5 percentage points; second, there is no fluctuation of secondary occlusion expansion during the recovery process (such as the area rebounding by more than 3 percentage points) to ensure the stability of the correction result. For example, a temporary obstruction may cause area fluctuations (35%→28%→40%→15%) due to wind blowing the safety rope. Due to the rebound phenomenon, the observation window needs to be extended to 10 seconds, and corrections can only be performed after the fluctuation disappears and stabilizes within the threshold.
[0131] As an optional implementation, in step S400, after determining the reflective vest wearing state recognition result of the human object to be detected according to the matching degree between the target wearing state feature and the preset wearing state threshold, the method may further include: Step S400A: When the recognition result is a partial occlusion state, extract the local occlusion features related to the occlusion area contour in the target wearing state features, and obtain the auxiliary verification image synchronously shot by the multi-angle acquisition device of the real-time monitoring image.
[0132] The local occlusion feature is a sub-feature vector related to the geometric shape and reflective characteristics of the occluded area segmented from the target wearing state feature. The auxiliary verification image refers to a set of multi-view images captured at the same time by optical acquisition devices deployed at different spatial positions. Specifically, in the construction site safety monitoring system, the main camera detects that there is an occlusion area of 35% on the right shoulder of a worker's reflective vest. The system immediately triggers three auxiliary cameras distributed on the left, top and back of the work area to synchronously collect images. The local occlusion features of the main camera include the boundary coordinates of the occlusion area (x1=180, y1=90, x2=220, y2=120), the average reflection intensity value of 85 (range 0-255) and the sequence of contour polygon vertices; the resolution and frame rate of the auxiliary verification image are consistent with the main camera to ensure the accuracy of spatiotemporal alignment. For example, the left camera captures the side image of the worker at a 45-degree angle, the top overlooking camera records the unfolded state of the reflective vest, and the rear camera supplements the perspective of the back area, forming four synchronous image streams. The timestamp error is controlled within ±10ms, providing a multi-dimensional data source for cross-perspective verification.
[0133] Step S400B: performing perspective alignment processing on the auxiliary verification image to generate a multi-angle verification image consistent with the human body posture in the real-time monitoring image, and extracting corresponding reflective vest area verification features from the multi-angle verification image.
[0134] The perspective alignment process maps multi-angle images to a unified human coordinate system through three-dimensional posture estimation and affine transformation. The reflective vest area verification feature refers to the standardized feature vector extracted from each perspective image. Specifically, the OpenPose algorithm is used to extract human skeleton key points (such as shoulders, hips, and knee joints) from the main camera image, build a three-dimensional posture model, and calculate the projection matrix under the perspective of the auxiliary camera. For example, the worker's main perspective is a standing posture, and the left arm is raised 30 degrees, resulting in wrinkles on the reflective vest. After the left camera image is projected, the reflective vest area is corrected to a plane unfolded state consistent with the main perspective. Subsequently, the same reflective vest area detection process is applied to the corrected auxiliary image: the reflective area boundary coordinates (x1=175, y1=85, x2=215, y2=115) are detected in the top view image, and the average reflection intensity is 82; the reflective area ratio in the back image is reduced to 25% due to backpack occlusion. After normalization, the verification features of each perspective generate a 512-dimensional feature vector with consistent dimensions for cross-perspective consistency verification.
[0135] Step S400C: performing cross-view feature matching on the local occlusion feature and the reflective vest area verification feature, and calculating the visibility confidence of the occlusion area at different viewing angles.
[0136] Cross-view feature matching refers to measuring the semantic consistency of local occlusion features and verification features of each view through cosine similarity, and visibility confidence reflects the probability of reproducibility of occluded areas under multiple views. Specifically, the similarity between the local occlusion feature of the main view and the verification feature of the left view is 0.78, with the top view is 0.85, and with the back view is 0.35. The visibility confidence CV is calculated as the weighted harmonic mean of the similarities of each view, and the weight is determined by the angle between the spatial position of the camera and the main view: the left view weight is 0.4 (angle 45 degrees), the top view is 0.3 (90 degrees), and the back view is 0.3 (180 degrees), then CV=(0.4×0.78+0.3×0.85+0.3×0.35) / (0.4+0.3+0.3)=0.66. It should be noted that when the similarity of a certain perspective is lower than 0.2, the perspective data is judged to be invalid and the weight is reallocated. For example, when the similarity of the rear perspective is 0.1 due to complete occlusion, its weight is transferred to the valid perspective to ensure the robustness of the confidence calculation.
[0137] Step S400D: If the visibility confidence is lower than a preset verification threshold, the blocked area is determined to be covered by a temporary interference object, and the recognition result is corrected to a normal wearing state.
[0138] The preset verification threshold is set to 0.7 based on the historical false alarm rate analysis. Temporary interference coverage refers to short-lived non-fixed occlusions (such as splashing water stains and short-lived flying insects). For example, when a worker is welding, the main camera detects that there is 15% occlusion in the reflective area on the chest, but the similarities of the verification features of the left and top perspectives are 0.85 and 0.88 respectively, and the similarity of the rear perspective is 0.25 due to the strong light interference of the arc. The visibility confidence CV after weight adjustment is 0.72, which is still lower than the threshold of 0.7. The system determines that it is a momentary light spot interference rather than a physical occlusion, corrects the recognition result to a normal wearing state, and suppresses false alarm triggering. The correction logic needs to be combined with time continuity verification: if the CV of three consecutive frames is lower than the threshold and the standard deviation of the occlusion area fluctuation is greater than 10%, the correction is performed; if only a single frame is lower than the threshold but the area is stable, the partial occlusion judgment is maintained.
[0139] Step S400E: If the visibility confidence is higher than a preset verification threshold, an obstruction type inference result is generated based on the shape consistency parameters of the obstruction area in the multi-angle verification image, and the inference result is searched for similarity with a preset obstruction database to determine whether the obstruction is an acceptable safety accessory.
[0140] The shape consistency parameter calculates the morphological matching degree of the multi-view occluded area contour through the Hausdorff distance. The occluder database contains typical contour templates of preset safety equipment such as safety buckles, tool kits, and walkie-talkies. For example, the Hausdorff distance of the occluded area between the main view and the left view is 5 pixels (similarity 0.95), and 8 pixels (0.92) with the top view. The shape consistency parameter SC = 1-(5+8) / (256+256) = 0.97. The contour vertex sequence is matched with the database. The similarity of the tool kit template is 0.89, the safety buckle is 0.65, and the walkie-talkie is 0.42. The occluder type is determined to be a tool kit. The list of acceptable equipment in the database includes necessary protective equipment such as safety buckles and breathing masks. If the matching result belongs to the list and the similarity is greater than 0.8, it is considered a compliant occlusion; otherwise, further inspection is triggered. For example, the tool kit is not included in the acceptable list, and it is still determined to be an unacceptable device even if the similarity is 0.89.
[0141] Step S400F: When the obstruction is determined to be an unacceptable safety accessory, alarm trigger data including the location and type information of the obstruction is generated, and the alarm trigger data is bound to the timestamp and spatial coordinates of the real-time monitoring image to generate a structured alarm record.
[0142] The alarm trigger data includes the coordinates of the occlusion bounding box (x1=180, y1=90, x2=220, y2=120), the type label "toolkit", the confidence score 0.89, and the multi-view verification image index. The timestamp is accurate to the millisecond level (such as 2023-09-15 14:23:45.789), and the spatial coordinates are converted into the construction site three-dimensional coordinate system (X=35.2m, Y=12.7m, Z=1.5m) through the laser rangefinder and the camera's internal and external parameter matrix. The structured alarm record is encapsulated in JSON format, which contains the above fields and the original image storage path. For example, you can refer to the following encapsulation example: json { "alert_id": "20230915142345789_001", "timestamp": "2023-09-15 14:23:45.789", "coordinates": {"X":35.2, "Y":12.7, "Z":1.5}, "obstruction_type": "Toolkit", "confidence": 0.89, "image_paths": [" / cam1 / frame_789.jpg", " / cam2 / frame_789.jpg"], "status": "pending" } The record is written into a distributed time series database in real time, supporting millisecond-level concurrent writing and multi-condition retrieval, and providing a standardized data interface for subsequent auditing and behavioral analysis.
[0143] Step S400G: Prioritize the structured alarm records, send warning signals of different levels to the target terminal according to preset alarm response rules, and store the alarm records in a historical database for subsequent behavior analysis model calls.
[0144] Priority sorting is dynamically calculated based on the risk level of the type of obstruction, spatial location sensitivity, and duration. The preset rule base defines: the risk level of tool bag obstruction is level 2 (medium). If it is located in the high-altitude working area (Z>5m), it will be upgraded to level 1 (high risk), triggering an audible and visual alarm and pushing it to the on-site safety officer's handheld terminal; if it is located in the ordinary area on the ground and the duration is less than 30 seconds, it will be marked as level 3 (low risk), only recording the log without triggering a real-time alarm. For example, in an alarm record, the tool bag obstruction is located in the scaffolding area 8 meters above the ground. The system immediately activates the warning light in the area to flash red, and sends an early warning message containing a location map to the safety officer's PDA within a radius of 50 meters. The historical database adopts a columnar storage structure, and a composite index is established by time partition (day / month / year) and spatial grid (10m×10m grid). The behavior analysis model can efficiently retrieve alarm records in specific time periods and areas, and explore the spatiotemporal distribution of illegal wearing patterns. For example, the incidence of tool bag obstruction events from 9 to 10 am on Mondays is 300% higher than that of other time periods. Based on this, the pre-job equipment inspection process is optimized.
[0145] Based on Figure 1 Based on the same principle as the method shown in , the embodiment of the present invention also provides a reflective vest wearing state recognition device 10 based on deep learning, such as Figure 2 As shown, the device 10 comprises: An image acquisition module 11 is used to acquire a real-time monitoring image of a target scene, wherein the real-time monitoring image contains at least one human object to be detected; A feature extraction module 12 is used to extract a reflective vest area image corresponding to the human object to be detected from the real-time monitoring image, and perform multi-scale feature extraction on the reflective vest area image to generate an initial wearing state feature; A feature analysis module 13 is used to perform regional distribution analysis on the initial wearing state feature based on a pre-trained wearing state recognition model to generate a target wearing state feature; wherein the pre-trained wearing state recognition model is obtained by fusing multimodal training data, and the multimodal training data includes sample images of reflective vests under different lighting conditions; The state recognition module 14 is used to determine the reflective vest wearing state recognition result of the human object to be detected according to the matching degree between the target wearing state feature and the preset wearing state threshold; wherein the recognition result includes normal wearing, not wearing or partially blocked state.
[0146] The above embodiment introduces the reflective vest wearing state recognition device 10 based on deep learning from the perspective of a virtual module. The following introduces a computer system from the perspective of a physical module, as shown below: An embodiment of the present invention provides a computer system, such as Figure 3 As shown, the computer system 100 includes: a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the computer system 100 may also include a transceiver 104. It should be noted that in actual applications, the transceiver 104 is not limited to one, and the structure of the computer system 100 does not constitute a limitation on the embodiments of the present invention.
[0147] Processor 101 may be a CPU, a general-purpose processor, a GPU, a DSP, an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 101 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0148] The bus 102 may include a path to transmit information between the above components. The bus 102 may be a PCI bus or an EISA bus, etc. The bus 102 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0149] The memory 103 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disk storage (including a compressed optical disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0150] The memory 103 is used to store application code for executing the solution of the present invention, and the execution is controlled by the processor 101. The processor 101 is used to execute the application code stored in the memory 103 to implement the content shown in any of the above method embodiments.
[0151] An embodiment of the present invention provides a computer system. The computer system in the embodiment of the present invention includes: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more programs are executed by the processor, the above method is implemented.
[0152] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a processor, the processor can execute corresponding contents in the aforementioned method embodiment.
[0153] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0154] The above descriptions are only partial embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A reflective vest wearing status recognition method based on deep learning, characterized in that: The method comprises: Acquire a real-time monitoring image of a target scene, wherein the real-time monitoring image contains at least one human object to be detected; Extracting a reflective vest region image corresponding to the human object to be detected from the real-time monitoring image, and performing multi-scale feature extraction on the reflective vest region image to generate initial wearing state features; Based on a pre-trained wearing state recognition model, the initial wearing state feature is analyzed for regional distribution to generate a target wearing state feature; wherein the pre-trained wearing state recognition model is obtained by fusing multimodal training data, and the multimodal training data includes sample images of reflective vests under different lighting conditions; According to the matching degree between the target wearing state feature and the preset wearing state threshold, the reflective vest wearing state recognition result of the human object to be detected is determined; wherein the recognition result includes normal wearing, not wearing or partially blocked state.
2. The method according to claim 1, characterized in that The extracting the reflective vest area image corresponding to the human object to be detected from the real-time monitoring image comprises: Performing human body contour segmentation on the real-time monitoring image to obtain a human body contour boundary frame of the human object to be detected; Based on the human body contour bounding box, locate the torso region of the human object to be detected, and perform color space conversion on the torso region to generate a first candidate region; Performing high-reflection area detection on the first candidate area, and screening out a pixel set whose reflection intensity is higher than a preset threshold; According to the connectivity distribution of the pixel set, the boundary coordinates of the reflective vest area image are determined, and the image area corresponding to the boundary coordinates is cropped from the real-time monitoring image.
3. The method according to claim 2, characterized in that The step of extracting multi-scale features from the reflective vest area image to generate initial wearing state features includes: Inputting the reflective vest area image into a multi-branch feature extraction network, the multi-branch feature extraction network includes a first branch network, a second branch network and a third branch network; wherein the first branch network is used to extract low-resolution texture features of the reflective vest area image, the second branch network is used to extract medium-resolution edge features, and the third branch network is used to extract high-resolution detail features; Performing cross-level feature fusion on the low-resolution texture features, medium-resolution edge features, and high-resolution detail features to generate a fused feature map; Channel attention weighting is performed on the fused feature map to enhance the feature channel weights related to the reflective material and compress the spatial dimension to generate the initial wearing state feature.
4. The method according to claim 3, characterized in that The training process of the pre-trained wearing state recognition model includes: Obtain multiple sets of training samples, each set of training samples includes a reflective vest sample image, a corresponding wearing state label, and a lighting condition label; Adaptively performing illumination enhancement on the reflective vest sample image to generate an enhanced sample image; wherein the adaptive illumination enhancement includes dynamically adjusting brightness compensation parameters according to the illumination condition label; Inputting the enhanced sample image into an initial recognition model and outputting predicted wearing state features; Calculating the classification loss between the predicted wearing state feature and the wearing state label, and updating the parameters of the initial recognition model based on gradient back propagation until the classification loss converges; The converged initial recognition model is end-to-end jointly trained with the multi-branch feature extraction network, the parameters of the multi-branch feature extraction network are fixed, and the regional distribution analysis layer parameters of the initial recognition model are optimized.
5. The method according to claim 4, characterized in that The adaptive illumination enhancement comprises the following steps: Determining the light intensity level of the current sample according to the light condition label; If the light intensity level is a low light condition, locally enhancing the contrast of the reflective vest sample image, and superimposing random noise to simulate low light interference; If the light intensity level is a high light condition, suppressing the overexposed area of the reflective vest sample image and reducing the global brightness to a preset range; If the illumination intensity level is a dynamically changing condition, multiple frames of the reflective vest sample image are fused in time series to generate an enhanced sample image with balanced illumination.
6. The method according to claim 1, characterized in that The pre-trained wearing state recognition model performs regional distribution analysis on the initial wearing state feature to generate a target wearing state feature, including: Dividing the initial wearing state feature into a plurality of spatial grid units, and calculating the similarity between a feature vector of each spatial grid unit and a preset reflective feature template; Generate a spatial weight matrix according to the distribution of the similarities, wherein the spatial weight matrix is used to identify the distribution probability of the effective reflective area in the reflective vest area image; Multiplying the spatial weight matrix by the initial wearing state feature element by element to obtain a weighted intermediate feature; Non-maximum suppression processing is performed on the intermediate features to remove redundant feature areas with an overlap rate higher than a preset threshold, thereby generating the target wearing state features.
7. The method according to claim 6, characterized in that The step of determining the reflective vest wearing state recognition result of the human object to be detected according to the matching degree between the target wearing state feature and the preset wearing state threshold comprises: Extracting a first statistic related to the coverage area of the reflective material, a second statistic related to the continuity of the reflective stripes, and a third statistic related to the matching degree of the human body posture from the target wearing state characteristics; Inputting the first statistic, the second statistic and the third statistic into a preset multi-condition decision tree model to generate a comprehensive matching score; If the comprehensive matching score is higher than the first threshold, it is determined to be in a normal wearing state; If the comprehensive matching score is lower than a second threshold, it is determined to be in a not-worn state; If the comprehensive matching score is between the first threshold and the second threshold, it is determined to be a partial occlusion state; The construction process of the multi-condition decision tree model includes: Collecting historical identification data, the historical identification data including multiple groups of the first statistics, the second statistics, the third statistics and corresponding real wearing state labels; Performing cluster analysis on the historical recognition data to determine distribution boundaries of different wearing status categories in a three-dimensional statistical space; generating a plurality of decision rules according to the distribution boundary, each decision rule corresponding to a hyperplane segmentation condition; The order of the decision rules is optimized based on the Gini coefficient, and hierarchical decision nodes are constructed until each leaf node contains samples of only a single wearing state category.
8. The method according to claim 1, characterized in that The method also includes a step of dynamically correcting the recognition result: When the recognition result is a partial occlusion state, obtaining a plurality of continuous frames of historical monitoring images of the human object to be detected; Extracting time series features from the historical monitoring images to generate a dynamic change trend of the wearing state of the reflective vest; If the dynamic change trend indicates that the blocked area continues to expand, the recognition result is corrected to the not-worn state; If the dynamic change trend indicates that the blocked area returns to the normal range within the preset time, the recognition result is corrected to the normal wearing state.
9. The method according to claim 8, characterized in that The extracting time series features of the historical monitoring images to generate the dynamic change trend of the wearing state of the reflective vest includes: Track the key points of the reflective vest area in multiple consecutive frames of historical surveillance images to determine the motion trajectory of the blocked area; Calculate the area change rate and shape similarity of the occluded area between adjacent frames; Constructing a time series feature vector according to the motion trajectory, area change rate and shape similarity; The time series feature vector is input into a long short-term memory network to predict the evolution path of the occluded area in several future frames, and the dynamic change trend is generated according to the evolution path.
10. A computer system, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Safety helmet and reflective vest detection method for multi-scene construction
CN114005089A
Reflective vest detection method under complex background
CN118351481A
High-altitude operation lifeline early warning method
CN119068412A
Image recognition and processing system based on deep learning
CN119741559A
Real-time object recognition using cascaded features, deep learning and multi-target tracking
US11055872B1
Cited By
Image data acquisition method and system for lens plastic part
CN120525734A
AI camera-based fishing information acquisition system of lamplight cover net fishing boat
CN120673331A
Method for identifying maintenance tools in shelter based on machine vision
CN120725946A
Target identification method and system based on deep learning
CN120876834A
Equipment state recognition system
CN121074340A