Deep Learning-Based Reflective Vest Wearing State Recognition Method and System

Through deep learning methods, multi-scale features of the reflective vest area are extracted and combined with pre-trained models, the recognition misjudgment problem caused by lighting and occlusion interference in the prior art is solved, and efficient wearable state recognition is achieved in complex environments, improving the accuracy and reliability of recognition.

CN119992602BActive Publication Date: 2025-07-18GUIZHOU NEW THINKING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510470212.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the prior art, the automatic identification method of the wearable state of the reflective vest is susceptible to interference with light conditions and changes in dynamic scenes, resulting in failure or misjudgment of feature extraction, making it difficult to accurately identify the wearable state in complex environments.

Method used

Using a deep learning-based method, the multi-scale features of the reflective vest area are extracted by acquiring real-time monitoring images, and the pre-trained wearable state recognition model is used to analyze the region distribution. Combined with multimodal training data, the dynamic correction mechanism eliminates occlusion interference and achieves efficient classification.

Benefits of technology

It significantly improves the robustness and accuracy of the wearable state recognition of reflective vests, and can accurately identify normal wearable, unwearable or partially obstructed states under different lighting conditions, reduces the risk of missed inspections, and ensures the safety of operators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992602B_ABST
    Figure CN119992602B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for recognizing the wearing state of a reflective vest based on deep learning. The method includes: obtaining a real-time monitoring image of a target scene, extracting an image of the reflective vest area corresponding to the human object to be detected from the real-time monitoring image, and performing multi-scale feature extraction on the reflective vest area image to generate initial wearing state features; based on a pre-trained wearing state recognition model, performing regional distribution analysis on the initial wearing state features to generate target wearing state features, and then determining the recognition result of the wearing state of the reflective vest of the human object to be detected according to the matching degree between the target wearing state features and a preset wearing state threshold; wherein the recognition result includes normal wearing, not wearing or partially blocked states. Through the present invention, the accuracy of recognizing the wearing state of the reflective vest can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and in particular, to a method and system for recognizing the wearing state of a reflective vest based on deep learning. Background Art

[0002] In the fields of industrial safety and security, traffic management, etc., it is an important measure to ensure personal safety for operators to wear reflective vests. In the prior art, for the automatic recognition method of the wearing state of a reflective vest, it usually relies on the feature detection of a single dimension, such as color threshold segmentation or reflection intensity analysis based on reflective materials. For example, by setting a fixed color interval or reflection intensity threshold, the reflective area is segmented from the surveillance image, and the wearing state is judged based on the area or shape rules of the region; or traditional image processing algorithms are used to extract the edge features of the reflective stripes, and the wearing state classification is carried out in combination with static template matching. However, the above methods have significant limitations in practical applications: First, the feature detection of a single dimension is easily interfered by complex lighting conditions. For example, strong light irradiation causes overexposure of the reflective area, and shadow coverage causes a sudden drop in reflection intensity, etc., which will cause feature extraction failure or misjudgment; Second, static thresholds or template matching are difficult to adapt to dynamic scene changes. For example, when the human body posture is inclined, part of the reflective vest is blocked by tools or the clothing folds are deformed, the traditional method cannot accurately distinguish normal wearing from abnormal states, resulting in missed detections or false alarms. Therefore, there is an urgent need for a method that can improve the recognition accuracy. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for recognizing the wearing state of a reflective vest based on deep learning. The embodiments of the present invention are implemented as follows:

[0004] In a first aspect, an embodiment of the present invention provides a method for recognizing the wearing state of a reflective vest based on deep learning, and the method includes:

[0005] Obtain a real-time surveillance image of a target scene, where the real-time surveillance image includes at least one human object to be detected;

[0006] Extract an image of the reflective vest area corresponding to the human object to be detected from the real-time surveillance image, and perform multi-scale feature extraction on the image of the reflective vest area to generate an initial wearing state feature;

[0007] Based on a pre-trained wearing state recognition model, perform regional distribution analysis on the initial wearing state feature to generate a target wearing state feature; wherein, the pre-trained wearing state recognition model is obtained by fusing multi-modal training data, and the multi-modal training data includes sample images of reflective vests under different lighting conditions;

[0008] Determine the recognition result of the reflective vest wearing state of the human object to be detected according to the matching degree between the target wearing state characteristics and the preset wearing state threshold; wherein, the recognition result includes normal wearing, not wearing or partially blocked state.

[0009] In a second aspect, the present invention provides a computer system, including:

[0010] One or more processors;

[0011] A memory;

[0012] One or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processor, the method described above is implemented.

[0013] Advantages of the present invention:

[0014] The method for recognizing the wearing state of a reflective vest based on deep learning provided by the embodiments of the present invention can accurately capture the dynamic change characteristics of the reflective material under different lighting conditions by extracting the reflective vest area from the real-time monitoring image and performing multi-scale feature fusion, combined with the spatial distribution analysis ability of the pre-trained model, thereby significantly improving the robustness and accuracy of the wearing state recognition. Specifically, the multi-scale feature extraction module comprehensively covers the structural and reflective characteristics of the reflective vest by fusing low-resolution texture features, medium-resolution edge features, and high-resolution detail features, effectively overcoming the problem of feature loss caused by uneven lighting or partial occlusion; the pre-trained wearing state recognition model adaptively learns the spatial distribution law of the reflective vest in different lighting scenarios based on multi-modal training data, enhancing the discrimination ability of the distribution change of the reflective area in complex environments; through the matching degree threshold determination mechanism, combined with multi-dimensional statistics such as the coverage area of the reflective material, stripe continuity, and pose matching degree, efficient classification of the wearing state is realized, avoiding the risk of misjudgment of a single feature. In addition, the dynamic correction mechanism can dynamically eliminate the interference of temporary occlusion through comprehensive analysis of multi-angle verification images and historical time-series data, further improving the real-time performance and reliability of state recognition. In summary, the method provided by the embodiments of the present invention can provide accurate wearing state monitoring and real-time warning capabilities for industrial security scenarios, reduce the risk of human missed inspections, and ensure the safety of operators.

[0015] In the following description, other features will be partly stated. When examining the following content and drawings, those skilled in the art will partly discover these features, or may learn these features through production or application. Through practicing or using various aspects of the methods, tools, and combinations listed in the detailed examples described later, the features in the current application can be implemented and obtained. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present invention.

[0017] Figure 1 It is a flowchart of a method for identifying the wearing state of a reflective vest based on deep learning provided by an embodiment of the present invention.

[0018] Figure 2 It is a schematic diagram of the functional module architecture of a device for identifying the wearing state of a reflective vest based on deep learning provided by an embodiment of the present invention.

[0019] Figure 3 It is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. Detailed implementation manners

[0020] The following describes the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. The terms used in the implementation manner part of the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.

[0021] The execution subject of the method for identifying the wearing state of a reflective vest based on deep learning in the embodiments of the present invention is a computer system, including but not limited to electronic device systems such as servers, personal computers, laptop computers, tablet computers, and smart phones.

[0022] The method for identifying the wearing state of a reflective vest based on deep learning provided by the embodiments of the present invention is as Figure 1 shown, and includes the following steps:

[0023] Step S100: Obtain a real-time monitoring image of a target scene, where the real-time monitoring image includes at least one human object to be detected.

[0024] Exemplarily, the target scenario may refer to a physical environment that requires security monitoring, such as a construction site, a mining operation area, a traffic intersection, or a night construction site. The real-time monitoring image refers to a sequence of digital images continuously captured at fixed time intervals by an optical camera device deployed in the scenario. The human object to be detected refers to a set of pixel regions in the image that have human morphological characteristics. Specifically, the real-time monitoring image can be collected by a camera with high dynamic range imaging capabilities to ensure that the contours of the human object and the details of the attached wearable items can still be completely recorded under strong light reflection or low illuminance conditions. For example, in the construction site scenario, the camera can collect RGB images with a resolution of 1920×1080 at a rate of 30 frames per second. There are three workers performing concrete pouring operations in one frame of the image, and their torso and limb regions form three independent human objects to be detected. It should be noted that the real-time monitoring image meets the preset clarity threshold, that is, the pixel coverage range of the key anatomical points of the human object (such as the shoulders and waist) in the image needs to exceed the minimum detection size to avoid subsequent recognition failure due to too far a distance. On this basis, the real-time monitoring image is transmitted to the central processing unit through a wired or wireless network to complete the format standardization and noise filtering preprocessing of the image data, providing a standardized input for subsequent analysis.

[0025] Step S200: Extract the reflective vest region image corresponding to the human object to be detected from the real-time monitoring image, and perform multi-scale feature extraction on the reflective vest region image to generate an initial wearing state feature.

[0026] Exemplarily, the reflective vest area image refers to a local image subset cropped based on preset geometric constraint conditions after positioning the torso area of the human object to be detected through a human key point detection algorithm, and its bounding box needs to completely cover the effective reflective material distribution range of the reflective vest. Multi-scale feature extraction refers to using the hierarchical convolutional kernel structure in a convolutional neural network to extract texture, shape, and reflection intensity features at different spatial resolution levels of the reflective vest area image. The initial wearing state feature refers to the multi-dimensional tensor data formed by splicing feature maps of each scale. Specifically, an instance segmentation model based on Mask R-CNN can be used to perform pixel-level segmentation on the human object, and the precise boundary coordinates of the reflective vest at the torso position can be obtained through regression calculation. For example, the upper body area of the human object can be divided into a rectangular area of 128×256 pixels as the reflective vest area image. Subsequently, parallel feature extraction branches with three convolutional kernel sizes of 3×3, 5×5, and 7×7 are constructed to capture microscopic texture details (such as the serrations on the edge of the reflective strip), mesoscopic structural features (such as the spacing between reflective strips), and macroscopic contour forms (such as the overall symmetry of the vest) from the reflective vest area image respectively. The multi-scale outputs are mapped to a unified dimensional space through a cross-layer feature fusion mechanism to form an initial wearing state feature vector with 512-dimensional channels. For example, in a night scene, the reflective vest area image contains bright reflective stripes generated by the illumination of vehicle lights, and multi-scale feature extraction can effectively distinguish the brightness gradient differences between normal reflection and overexposed areas, avoiding misjudging strong light interference as the non-wearing state.

[0027] Step S300: Based on the pre-trained wearing state recognition model, perform regional distribution analysis on the initial wearing state feature to generate a target wearing state feature; wherein, the pre-trained wearing state recognition model is obtained by fusing multi-modal training data, and the multi-modal training data includes reflective vest sample images under different lighting conditions.

[0028] Exemplarily, the pre-trained wearable state recognition model refers to a deep neural network that has completed parameter optimization on an annotated dataset containing various environmental conditions. The regional distribution analysis refers to analyzing the statistical distribution law of the initial wearable state features within the reflective vest area through a spatial attention mechanism. The target wearable state features refer to the high-order feature representations after semantic enhancement. The multi-modal training data specifically covers sample images under complex conditions such as natural light, artificial supplementary light, backlight, shadow occlusion, and rain and fog interference, ensuring that the model has cross-scene generalization ability. Specifically in implementation, the wearable state recognition model can adopt a Transformer architecture with a self-attention module, encoding the initial wearable state features into a serialized input according to spatial positions, and calculating the correlation weights between different feature regions through a multi-head attention mechanism. For example, under backlight conditions, the model will enhance the attention to the edge contour features of the vest and suppress the loss of details in the central region caused by backlight. In the samples of rainy and foggy weather, the model automatically corrects color distortion through cross-modal contrast learning and accurately identifies the reflective strips covered by water stains. During the training process, techniques such as light intensity normalization and adversarial sample generation are adopted to enable the model to extract light-invariant features from multi-modal data. After the regional distribution analysis, the target wearable state features will include discriminative indicators such as the coverage rate of reflective materials, the wearing symmetry index, and the confidence level of the occlusion area. For example, a 2048-dimensional dense vector is generated to represent the complete wearing degree of the vest.

[0029] Step S400: Determine the recognition result of the reflective vest wearing state of the human object to be detected according to the matching degree between the target wearable state features and the preset wearable state threshold; wherein, the recognition result includes the states of normal wearing, not wearing, or partial occlusion.

[0030] Exemplarily, the preset wearing state threshold refers to the classification decision boundary determined by statistical learning, the matching degree refers to the distance metric value of the target wearing state feature from the center of each state cluster in the embedding space, and the recognition result is determined by maximum likelihood estimation or a support vector machine classifier. Specifically, the normal wearing state corresponds to the feature space region where the reflective vest completely covers the torso without external object occlusion, the non-wearing state corresponds to the feature distribution where the area of the reflective material region is lower than the minimum threshold, and the partially occluded state is between the two and has local feature anomalies. When performing classification, the cosine similarity is used to calculate the matching scores of the target wearing state feature with the three types of reference templates. When the matching degree of the normal wearing category exceeds 0.95, it is directly determined as the compliant state; if the matching degree is in the range of 0.75 to 0.95 and local feature mutations are detected, the partial occlusion determination is triggered; when the matching degree is lower than 0.4, it is determined as not worn. For example, the similarity of a certain target wearing state feature with the normal wearing template is 0.92, but the similarity with the partially occluded template reaches 0.85. At this time, the system will further analyze the occlusion-sensitive dimensions in the feature space to confirm whether there is a situation where the tool bag strap covers the reflective strip, and finally output the partially occluded state. This method adapts to the recognition sensitivity requirements of different scenarios through a dynamic threshold adjustment mechanism to ensure the classification stability under strong environmental interference.

[0031] As an implementation manner, in the step S200, extracting the reflective vest region image corresponding to the human body object to be detected from the real-time monitoring image may include the following steps:

[0032] Step S210: Perform human body contour segmentation on the real-time monitoring image to obtain the human body contour bounding box of the human body object to be detected.

[0033] Human body contour segmentation refers to the process of separating the human body pixel region from the real-time monitoring image through a semantic segmentation algorithm. The human body contour bounding box refers to a set of geometric parameters that mark the position of the human body object in the image in the form of rectangular coordinates. Specifically, an instance segmentation model based on Mask R-CNN can be used to perform pixel-level classification on the input image, and the bounding box coordinates and corresponding binary masks of each human body object to be detected are obtained through regression calculation. For example, in a construction site scenario, there are three workers in the real-time monitoring image, and the model outputs three independent human body contour bounding boxes respectively. Each bounding box contains the upper left corner coordinates (x1, y1) and the lower right corner coordinates (x2, y2), and their values are normalized to between 0 and 1 according to the image resolution. It should be noted that non-human interference objects, such as hanging reflective warning signs or local reflective areas of moving mechanical equipment, need to be excluded in human body contour segmentation to ensure that the bounding box only covers the complete human body form. During the implementation process, the model enhances its generalization ability through various pose samples (such as bending, raising hands, squatting) included in the pre-trained dataset to avoid contour breakage or bounding box offset caused by human body pose changes. For example, when a worker bends down to operate the equipment, the model can still accurately identify the contours of his back and legs and generate a bounding box that closely fits the actual occupancy of the human body, providing a reliable input for subsequent trunk region positioning.

[0034] Step S220: Based on the human body contour bounding box, locate the trunk region of the human body object to be detected, and perform color space conversion on the trunk region to generate a first candidate region.

[0035] The trunk region refers to the anatomical range of the human body from the shoulders to the waist. The positioning process delimits a rectangular sub-region with a fixed ratio within the human body contour bounding box through a proportional mapping algorithm. Color space conversion refers to converting the original RGB image into the HSV (hue, saturation, value) color space, which is more suitable for reflection intensity analysis. Specifically, according to the statistical laws of human body morphology, the trunk region occupies 60% to 80% of the height range in the middle in the vertical direction of the bounding box and is as wide as the bounding box in the horizontal direction. For example, for a human body contour bounding box with a size of 320×480 pixels, the trunk region is defined as a rectangular region from the ordinate y = 96 pixels to y = 384 pixels, and the width remains 320 pixels unchanged. Subsequently, this region is converted from the RGB space to the HSV space, and the reflection characteristics of the reflective material are enhanced using the value channel. For example, under backlight conditions, the silver-gray stripes of the reflective vest may appear dark due to insufficient exposure in the RGB space, while the value channel of the HSV space can effectively distinguish the brightness difference between the reflective material and ordinary clothing through linear stretching, generating a first candidate region with higher contrast. It should be noted that the bicubic interpolation algorithm needs to be used during the color space conversion process to maintain the image resolution and avoid the loss of high-frequency details caused by color quantization.

[0036] Step S230: Detect high-reflection regions in the first candidate region, and screen out a set of pixels with a reflection intensity higher than a preset threshold.

[0037] Exemplarily, the high-reflection region detection refers to a pixel-level threshold segmentation operation based on the lightness channel. The reflection intensity is quantified by the value of the lightness channel in the HSV color space, and the preset threshold is calibrated through experiments according to the optical properties of the reflective material. Specifically, set the lightness threshold V_threshold = 200 (value range 0 - 255), and classify all pixels with V ≥ 200 in the first candidate region as high-reflection pixels to form a binary mask image. For example, in a night scene, the glass bead material on the worker's reflective vest reflects strongly due to the headlight illumination, and its lightness value can reach above 240, while the lightness value of ordinary work clothes is usually lower than 180. Through threshold segmentation, the pixel set of the reflective stripe can be accurately extracted. It should be noted that the preset threshold needs to be dynamically adjusted according to the ambient light. For example, under strong noon light conditions, the lightness value of the reflective material may reach 255 due to overexposure. At this time, the saturation channel needs to be combined for auxiliary judgment to avoid misjudging white clothes as reflective regions. During the implementation process, an adaptive threshold algorithm can be used to dynamically calculate the local optimal segmentation threshold based on the peak position of the lightness histogram of the first candidate region to ensure the detection robustness under different lighting conditions.

[0038] Step S240: Determine the boundary coordinates of the reflective vest region image according to the connectivity distribution of the pixel set, and crop the image region corresponding to the boundary coordinates from the real-time monitoring image.

[0039] The connectivity distribution refers to the spatial connection relationship of adjacent pixels in a binary image, and the boundary coordinates refer to the vertex coordinates of the smallest rectangle covering all connected pixels. Exemplarily, morphological closing operations can be performed on the high-reflection pixel set to fill holes, and then the circumscribed rectangle parameters of the largest connected region are extracted through an edge tracking algorithm. For example, in rainy and foggy weather, the stripes of the reflective vest may be partially covered by water droplets, resulting in multiple discrete connected regions in the pixel set. By calculating the area of each region and retaining the largest connected region, the interference of water droplet reflection can be excluded, and the boundary of the reflective vest region can be accurately delimited. During implementation, the calculation of the boundary coordinates needs to consider the curvature of the human torso. The elastic bounding box algorithm is used to allow the rectangle to bend vertically with the shape of the vest, so as to tightly wrap the distribution area of the reflective material. For example, the actual distribution of a certain reflective vest is trapezoidal, the left and right boundaries of the bounding box remain vertical, and the upper and lower boundaries are dynamically adjusted according to the extreme points of the connected region. Finally, a reflective vest region image with a size of 256×128 pixels is generated for subsequent feature extraction.

[0040] As an implementation manner, step S240 of determining the boundary coordinates of the reflective vest area image according to the connectivity distribution of the pixel set includes:

[0041] Step S241: Perform an eight-neighborhood connected region detection on the pixel set to generate a connected region set including at least one connected region.

[0042] The eight-neighborhood connected region detection refers to checking the connection status of the eight adjacent pixels including the upper, lower, left, right, and diagonal pixels centered on a pixel. The connected region set refers to a list of independent regions composed of spatially continuous pixel groups. For example, in the scenario where the reflective vest stripes are broken, the pixel set may include three vertically arranged connected regions, each corresponding to a reflective stripe. Specifically, in implementation, a two-pass scanning algorithm can be adopted. For example, in the first pass, the image pixels are traversed and temporary labels are marked, and in the second pass, equivalent labels are merged to generate the final connected regions. For example, in a certain reflective vest area, due to button occlusion, the middle reflective stripe is split into upper and lower parts. The algorithm identifies them as two connected regions and assigns label IDs = 1 and 2 respectively to ensure that each sub-region can be analyzed independently in subsequent processing.

[0043] Step S242: Extract geometric attributes of each connected region in the connected region set, where the geometric attributes include region area, aspect ratio, and minimum bounding rectangle angle.

[0044] The geometric attribute extraction refers to calculating the shape parameters of the connected region through mathematical morphology operations. The region area refers to the total number of pixels included in the connected region. The aspect ratio refers to the ratio of the width to the height of the minimum bounding rectangle. The minimum bounding rectangle angle refers to the angle between the main axis of the rectangle and the horizontal axis of the image. For example, a certain connected region contains 1200 pixels, the width of its minimum bounding rectangle is 60 pixels, the height is 20 pixels, the aspect ratio is 3:1, and the main axis angle is 1.5 degrees (close to the horizontal direction). It should be noted that when calculating the aspect ratio, noise interference needs to be excluded. For example, when the connected region has a jagged edge due to uneven illumination, the minimum area bounding rectangle algorithm based on the convex hull is used to improve the measurement accuracy to ensure that the geometric attributes truly reflect the structural characteristics of the reflective stripe.

[0045] Step S243: Based on a preset reflective vest shape rule, screen the candidate connected regions in the connected region set that satisfy the aspect ratio within a preset interval, the region area is greater than the minimum area threshold, and the bounding rectangle angle is aligned with the human torso axis.

[0046] The regular shape of the reflective vest refers to the aspect ratio range, minimum reflection area requirement, and direction consistency constraint defined according to the standard design specifications of reflective vests. For example, the preset aspect ratio range is [2.5, 5.0], the minimum area threshold is 800 pixels, and the deviation of the circumscribed rectangle angle from the torso axis does not exceed ±5 degrees. Specifically, during the screening process, if the aspect ratio of a connected region is 4.2, the area is 950 pixels, and the circumscribed rectangle angle is 2 degrees (the human torso axis is usually in the vertical direction), it is determined as a candidate connected region; while another region with an aspect ratio of 1.8 and an area of 600 pixels is excluded because it does not meet the conditions. It should be noted that the human torso axis is approximated by the perpendicular bisector of the human contour bounding box. When the angle between the circumscribed rectangle angle and this perpendicular bisector exceeds the threshold, it indicates that the reflective strip may be tilted or distorted, not conforming to the normal wearing form.

[0047] Step S244: Perform region merging processing on the candidate connected regions, and merge adjacent candidate connected regions with a spatial distance less than the preset merging threshold into an extended connected region.

[0048] Exemplarily, the region merging processing can be to fuse adjacent regions into a single region by calculating the shortest Euclidean distance between candidate connected regions. The preset merging threshold is set according to the standard spacing of the reflective stripes. For example, the merging threshold is set to 15 pixels. When the minimum edge distance between two candidate connected regions is less than this value, it is determined that they belong to different segments of the same reflective strip and are merged. Specifically in implementation, a region-growing-based method is adopted: taking the centroid of each candidate connected region as the seed point, and expanding outward until the merging threshold boundary of other regions is reached. For example, in the middle of the reflective vest, there are two broken reflective strips caused by wrinkles, and the centroid distance between them is 12 pixels, less than the merging threshold, so they are merged into an extended connected region, thus restoring the complete shape of the reflective strip.

[0049] Step S245: Perform convex hull detection on the extended connected region to determine the vertex coordinates of the minimum convex polygon covering all merged pixel points.

[0050] Convex hull detection refers to calculating the convex hull of the extended connected region, and the vertex coordinates of the minimum convex polygon refer to the set of polygon vertices that can contain all pixels and the interior angles are all less than 180 degrees. Specifically, the Graham scan algorithm is used. First, find the pixel point with the minimum ordinate as the reference point, sort the remaining points by polar angle, and then construct the convex hull in sequence. For example, an extended connected region contains scattered pixel points, and its convex hull vertex coordinates are {(x1,y1), (x2,y2), (x3,y3), (x4,y4)}, forming a quadrilateral to enclose all reflective pixels. It should be noted that convex hull detection can eliminate the influence of internal depressions or burrs in the region, generate a smooth edge contour, and provide a geometric basis for subsequent boundary coordinate calculation.

[0051] Step S246: Calculate the maximum left and right boundary coordinates in the horizontal direction and the maximum upper and lower boundary coordinates in the vertical direction according to the extreme point positions of the vertex coordinates of the minimum convex polygon, and generate the boundary coordinates of the reflective vest area image.

[0052] The extreme points are the maximum and minimum coordinate values of the convex polygon vertices in the horizontal and vertical directions. The boundary coordinates define the four vertices of a rectangular area by these extreme points. Specifically, the left boundary in the horizontal direction is the minimum x value among all vertices, and the right boundary is the maximum x value; the upper boundary in the vertical direction is the minimum y value, and the lower boundary is the maximum y value. For example, if x_min = 100, x_max = 300, y_min = 50, and y_max = 200 in the convex hull vertex set, then the boundary coordinates of the reflective vest area image are the rectangular box (100, 50, 300, 200). During the implementation process, dilation processing needs to be performed on the boundary coordinates, expanding outward by 2 - 3 pixels to ensure complete wrapping of the edge of the reflective material and avoid losing key features during cropping. For example, the size of the finally cropped image area is (x_max - x_min + 6) × (y_max - y_min + 6), that is, 206 × 156 pixels, retaining the transition area between the reflective stripes and the surrounding background for multi-scale feature extraction.

[0053] As an implementation manner, in the step S200, multi-scale feature extraction is performed on the reflective vest area image to generate initial wearing state features, including:

[0054] Step S250: Input the reflective vest area image into a multi-branch feature extraction network, and the multi-branch feature extraction network includes a first branch network, a second branch network, and a third branch network; wherein, the first branch network is used to extract the low-resolution texture features of the reflective vest area image, the second branch network is used to extract the medium-resolution edge features, and the third branch network is used to extract the high-resolution detail features.

[0055] Exemplarily, the multi-branch feature extraction network is multiple convolutional neural network structures deployed in parallel. Each branch processes the input image at a different sampling rate to capture cross-scale features. The first branch network downsamples the input to 1 / 4 of the original size through a convolutional layer with a stride of 2, and extracts macroscopic texture features such as the distribution density of reflective materials; the second branch network maintains the original resolution, enhances the edge response through the Sobel operator, and captures the boundary between the reflective strip and the background; the third branch network uses dilated convolution to expand the receptive field and extracts microscopic details such as the arrangement pattern of reflective particles while maintaining high resolution. For example, for an input image of the reflective vest area with 256×128 pixels, the first branch outputs a feature map of 64×32×64, the second branch outputs a feature map of 256×128×128, and the third branch outputs a feature map of 256×128×256, respectively representing visual information at different abstraction levels.

[0056] Step S260: Perform cross-level feature fusion on the low-resolution texture features, medium-resolution edge features, and high-resolution detail features to generate a fused feature map.

[0057] Cross-level feature fusion refers to aligning multi-scale features to a unified spatial dimension through upsampling, downsampling, and channel concatenation operations, and integrating complementary information. Specifically, the low-resolution texture features of the first branch are upsampled to the original size and added to the medium-resolution edge features of the second branch channel by channel; then the result is concatenated with the high-resolution detail features of the third branch along the channel axis to form a fused feature map with 512 channels. For example, after the low-resolution features are upsampled to 256×128 by bilinear interpolation, they are added to the 128 channels of the medium-resolution features to generate an intermediate feature with 192 channels; then it is concatenated with the 256 channels of the high-resolution features to finally obtain a fused feature map with 448 channels. It should be noted that a 1×1 convolutional kernel is used for channel dimensionality reduction during the fusion process to avoid dimensional explosion and enhance feature correlation. In addition, the semantic alignment and information complementation of multi-scale features can be achieved through the Feature Pyramid Network (FPN) structure. Specifically, the low-resolution texture features are upsampled and added to the medium-resolution edge features element by element, and the result is upsampled again and fused with the high-resolution detail features. For example, the 1 / 16-scale feature map of the first branch is upsampled to the 1 / 8 scale through transposed convolution and added to the same-scale feature map of the second branch; then it is upsampled again to the original image scale and concatenated with the features of the third branch. A 1×1 convolution is introduced during the fusion process to adjust the number of channels and reduce the feature redundancy after fusion. The fused feature map will simultaneously contain the global layout information, local edge structure, and fine-grained texture details of the reflective vest, providing a multi-level basis for the analysis of the wearing state.

[0058] Step S270: Perform channel attention weighting on the fused feature map to enhance the weights of the feature channels related to the reflective material and compress the spatial dimension to generate the initial wearable state feature.

[0059] Channel attention weighting dynamically adjusts the importance of each channel through the SENet (Squeeze-and-Excitation Network) mechanism. The channels related to the reflective material refer to the feature channels that are identified as having a high contribution to the classification of the wearable state during the training process. Specifically, global average pooling is performed on the fused feature map to generate a channel description vector. The channel dependencies are learned through a fully connected layer and the weight coefficients are output, and the weights are multiplied with the original feature map channel by channel. For example, after attention weighting of the 448-channel fused feature, the weights of channels 112, 215, and 398 related to the reflection intensity of the reflective strip are increased to 1.5 times, while the weights of the channels unrelated to the background are reduced to below 0.3. Subsequently, global max pooling is used to compress the 256×128×448 feature map into a 448-dimensional vector, which is used as the initial wearable state feature and input into the subsequent classification model. This feature vector can encode the wearing integrity, occlusion degree, and reflection characteristics of the reflective vest through end-to-end training, providing a discriminative basis for the final state recognition.

[0060] As an implementation, the training process of the pre-trained wearable state recognition model includes the following steps:

[0061] Step S10: Obtain multiple groups of training samples, each group of training samples including a reflective vest sample image, the corresponding wearable state label, and the illumination condition label.

[0062] The sample images of the reflective vests are the original image data containing the wearing states of the reflective vests collected by different camera devices in diverse scenarios. The wearing state labels refer to the classification results of the wearing states (properly worn, not worn, or partially blocked) manually labeled or verified by sensors. The lighting condition labels refer to the classification of the environmental lighting intensity levels (low lighting, high lighting, or dynamically changing conditions) obtained by photometric measurement or image metadata parsing. Specifically, the training sample set covers scenarios such as construction sites, traffic intersections, and night operation areas. For example, it contains 5000 sample images with a resolution of 1920×1080, among which there are 3000 sample images in the properly worn state (including 800 overexposed images of reflective stripes under strong noon sunlight, 1200 images under low twilight lighting, and 1000 images under dynamic lighting conditions of rain and fog), 1000 sample images in the not worn state (workers only wearing ordinary work clothes), and 1000 sample images in the partially blocked state (the reflective vests are partially covered by tool bags or safety ropes). It should be noted that the sample images need to be preprocessed by geometric correction and white balance to eliminate the interference of lens distortion and color temperature difference on model training. For example, the checkerboard calibration method is used to correct the barrel distortion caused by wide-angle lenses to ensure the spatial proportion consistency of the images in the reflective vest area.

[0063] Step S20: Perform adaptive illumination enhancement on the sample images of the reflective vests to generate enhanced sample images; wherein, the adaptive illumination enhancement includes dynamically adjusting the brightness compensation parameters according to the lighting condition labels.

[0064] Adaptive illumination enhancement is a preprocessing method for directionally adjusting the image brightness, contrast, and noise distribution based on lighting condition labels. The brightness compensation parameters refer to the adjustable coefficients that control the gamma correction curve, histogram equalization intensity, and noise injection amount. For example, for samples with low lighting labels, the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm is used to enhance the local contrast, and at the same time, random noise conforming to the Poisson distribution is added to simulate the grain effect of low-light sensors; for samples with high lighting labels, the brightness values of overexposed areas are compressed through a highlight suppression function, and the global brightness mean is adjusted to a preset range (such as RGB mean 180 - 200) by linear scaling. Specifically in implementation, the dynamic adjustment mechanism is realized through a look-up table, and the mapping relationship between the lighting intensity levels and processing parameters is established in advance. For example, the low lighting level corresponds to a CLAHE grid size of 8×8 and a contrast limit of 2.0, and the high lighting level corresponds to a brightness scaling factor of 0.7 and a saturation gain of 1.2. This method ensures that the enhanced sample images retain the reflection characteristics of the reflective materials while enhancing the robustness of the model to lighting interference.

[0065] As an implementation manner, step S20, performing adaptive illumination enhancement on the sample images of the reflective vests to generate enhanced sample images, includes:

[0066] Step S21: Determine the illumination intensity level of the current sample according to the illumination condition label.

[0067] The illumination intensity level is a discrete illumination intensity interval divided based on photometric measurement data. For example, low illumination corresponds to 0 - 200 lux (such as a night scene), high illumination corresponds to 2000 - 10000 lux (such as direct noon sunlight), and the dynamic change condition means that the illumination intensity fluctuates by more than 500 lux within a single sample (such as intermittent shadows caused by cloud movement). Specifically, by analyzing the aperture, shutter speed, and ISO value in the image EXIF metadata, combined with the actual illuminance value recorded by the ambient light sensor, the illumination level of the sample is comprehensively determined. For example, if the EXIF parameters of a sample are f / 2.8, 1 / 60s, ISO 1600, and the ambient light is recorded as 150 lux, it is determined to be a low illumination level; for another sample with parameters f / 8, 1 / 1000s, ISO 100 and ambient light of 8000 lux, it is classified as a high illumination level. This classification provides a control basis for subsequent adaptive enhancement.

[0068] Step S22: If the illumination intensity level is a low illumination condition, perform local contrast enhancement on the reflective vest sample image and superimpose random noise to simulate low illumination interference.

[0069] Local contrast enhancement is an operation to improve the visibility of dark details in a local area of the image, and random noise simulation refers to adding signal distortion that conforms to the characteristics of low illumination imaging. Specifically, the CLAHE algorithm is used to divide the image into local areas of 32×32 pixels, perform histogram equalization on each area, and limit the contrast increase to no more than 3.0 to prevent excessive noise amplification. Subsequently, a Gaussian noise matrix with a mean of 0 and a variance of 0.01 is generated and superimposed on the enhanced image at the pixel level. For example, the RGB values of the reflective stripe in the original low illumination sample are (30, 35, 40), which are increased to (80, 85, 90) by CLAHE, and fine-tuned to (82, 83, 89) after superimposing the noise, while retaining the true noise characteristics while enhancing the reflective area. It should be noted that the noise injection intensity needs to be positively correlated with the ISO value. For example, a sample with ISO 1600 adds noise with a variance of 0.02, and the variance of the ISO 3200 sample is increased to 0.03 to simulate the grain effect under different sensitivities.

[0070] Step S23: If the illumination intensity level is a high illumination condition, suppress the overexposed area of the reflective vest sample image and reduce the global brightness to a preset range.

[0071] Overexposure area suppression recovers the highlight details by threshold detection and brightness compression. The preset brightness range refers to the RGB mean value interval (such as 180 - 200) set according to the reflection characteristics of the reflective material. Specifically, pixels with any value in the RGB channels of the detected image exceeding 240 are detected, and an S-shaped curve adjustment is applied to them: the input value I is linearly compressed to the interval [200, 220], and the formula is, for example, I out = 200 + 20 * (I - 240) / (255 - 240). Subsequently, the RGB mean value of the entire image is adjusted to 190 ± 5 by linear scaling. For example, the original values of the overexposed area of a high-light sample are (255, 250, 245), and after compression, they are (220, 215, 210), and the global brightness mean value drops from 230 to 195, enabling the texture details of the reflective stripes (such as the suture direction) to be revealed. It should be noted that the brightness adjustment needs to be performed in the HSV space to avoid color deviation. Only the value of the lightness channel (V value) is modified and then converted back to the RGB space.

[0072] Step S24: If the light intensity level is a dynamically changing condition, then perform multi-frame temporal fusion on the reflective vest sample image to generate an enhanced sample image with balanced illumination.

[0073] Multi-frame temporal fusion is to align and perform weighted averaging on multiple consecutive frames of images collected under the same scene to smooth the influence of illumination fluctuations. Specifically, five frames of images before and after the dynamically illuminated sample (time window 100 ms) are selected, the inter-frame displacement is estimated by the optical flow method and an affine transformation is performed for alignment, and then the median value of the RGB channels of each pixel is taken to generate a fused image. For example, in a sequence of dynamically illuminated samples, there are brightness fluctuations caused by cloud occlusion (the inter-frame brightness difference reaches 300 lux). After median fusion, the brightness standard deviation of the reflective vest area drops from 45 to 12, eliminating the influence of instantaneous shadows or overexposure on a single-frame image. This method effectively improves the stability of reflective material detection under dynamic conditions. For example, the continuity of the reflective stripes in some occluded areas in the fused image is enhanced, avoiding misclassification caused by abnormal illumination in a single frame.

[0074] Step S30: Input the enhanced sample image into the initial recognition model to output the predicted wearing state features.

[0075] The initial recognition model is a deep convolutional neural network architecture with unoptimized parameters. The predicted wearing state features refer to the multi-dimensional feature vector before the last fully connected layer of the model. Exemplarily, the model can adopt the ResNet-50 backbone network. After removing the top classifier, a 2048-dimensional feature vector is retained as the output. For example, when an enhanced sample image with an input size of 256×256 pixels is input, after gradually extracting spatial features through convolutional layers and residual blocks, a feature tensor of 1×1×2048 is finally generated. It should be noted that the weights of the model are randomly initialized at the beginning of training. Through forward propagation, the enhanced sample image is mapped to a high-dimensional feature space, and there is a large difference between its feature distribution and the true wearing state label. It needs to be gradually optimized through subsequent loss calculation and backpropagation.

[0076] Step S40: Calculate the classification loss between the predicted wearing state features and the wearing state label, and update the parameters of the initial recognition model based on gradient backpropagation until the classification loss converges.

[0077] Exemplarily, the classification loss is the difference between the predicted probability distribution of the model and the true label distribution measured by the cross-entropy function. Gradient backpropagation means updating the parameters according to the partial derivatives of the loss function with respect to the weights of each layer to minimize the loss. Specifically, the 2048-dimensional predicted wearing state features are input into the fully connected classification layer to generate the logical output values of the three wearing states. After being converted into a probability distribution through the Softmax function, the cross-entropy loss with the one-hot encoded label is calculated. For example, if the true label of a sample is normal wearing (encoded [1,0,0]) and the model output probability is [0.7, 0.2, 0.1], then the cross-entropy loss is -log(0.7)=0.356. The Adam optimizer is used to update the parameters with an initial learning rate of 0.001. When the decrease in the training set loss is less than 0.1% for 10 consecutive epochs, it is determined to converge. Through this iterative optimization process, the feature space of the model gradually aligns with the semantics of the wearing state. For example, the feature vectors of normal wearing samples form a tight cluster in the embedding space, maintaining an obvious distance from the feature clusters of non-wearing samples.

[0078] Step S50: Conduct end-to-end joint training on the converged initial recognition model and the multi-branch feature extraction network, fix the parameters of the multi-branch feature extraction network, and optimize the parameters of the regional distribution analysis layer of the initial recognition model.

[0079] End-to-end joint training connects a multi-branch feature extraction network and an initial recognition model into a unified computational graph for overall optimization. Fixed parameters mean that during backpropagation, the weight updates of the specified network layers are prohibited. Specifically, the multi-branch feature extraction network has completed parameter convergence through the task of extracting image features in the reflective vest area during the early training. During joint training, its convolution kernel weights are frozen, and only the self-attention parameters of the region distribution parsing layer (i.e., the Transformer encoder layer) in the initial recognition model are allowed to be adjusted. For example, in each training batch, the initial wearing state features output by the multi-branch network are directly input into the initial recognition model. By calculating the wearing state classification loss and backpropagating it to the multi-head attention module of the region distribution parsing layer, the projection weights of its query and key-value matrices are optimized. This strategy ensures that while the model retains the underlying feature extraction ability, it enhances the discriminability of high-level semantic parsing. For example, it makes the region distribution parsing layer pay more attention to the spatial distribution pattern of the reflective material coverage rate rather than the basic texture details.

[0080] As an implementation, step S300, based on the pre-trained wearing state recognition model, performs region distribution parsing on the initial wearing state features to generate target wearing state features, including:

[0081] Step S310: Divide the initial wearing state features into multiple spatial grid units, and calculate the similarity between the feature vector of each spatial grid unit and a preset reflective feature template.

[0082] Spatial grid units are a set of rectangular sub-regions obtained by evenly dividing the initial wearing state feature map along the height and width directions. The preset reflective feature template refers to a set of reference feature vectors extracted from standard reflective vest samples. Specifically, the size of the initial wearing state feature map is 32×64×512 (height×width×number of channels), which is divided into 32 8×8 pixel spatial grid units according to a 4×8 grid, and each unit corresponds to a 512-dimensional feature vector. The preset reflective feature template contains the typical features of reflective stripes in the normal wearing state, such as the mean vector extracted from 1000 standard samples. When calculating the similarity, the cosine similarity is used to measure the direction consistency between the feature vector of each grid unit and the template vector. For example, the cosine similarity between the feature vector of a certain grid unit and the template is 0.92, indicating that the reflective characteristics of this region are highly consistent with the standard; while the similarity of another grid unit is only 0.35, suggesting that there may be occlusion or interference from non-reflective materials. It should be noted that the grid division density needs to be adapted to the resolution of the reflective vest area image to ensure that each grid covers an area with an actual physical size of approximately 5cm×5cm to match the minimum detectable width of the reflective stripes.

[0083] Step S320: Generate a spatial weight matrix according to the distribution of the similarity. The spatial weight matrix is used to identify the distribution probability of the effective reflective regions in the image of the reflective vest area.

[0084] The spatial weight matrix is a two-dimensional matrix with the same dimension as the spatial grid cells. The element values of it are the weighted results of the similarity of the corresponding grid cells. The distribution probability of the effective reflective regions refers to the confidence that each grid cell belongs to the compliant reflective material. Specifically, arrange the similarities of the 32 grid cells calculated in step S310 in the original spatial positions to form a 4×8 matrix, and perform spatial smoothing processing through Gaussian filtering to eliminate the influence of local outliers. For example, the similarity of a central grid is 0.88, and its adjacent grids are 0.75, 0.82, and 0.90. After Gaussian filtering with σ = 1.5, the central value is adjusted to 0.85, which is more in line with the continuous distribution characteristics of the actual reflective stripes. Further, perform normalization processing on the smoothed matrix to map the weight value range to the 0-1 interval, and the grids with weights ≥ 0.7 are determined as effective reflective regions. This method can accurately identify the core reflection regions of the reflective vest. For example, in some occlusion scenarios, the weight of the right grid covered by the tool kit drops below 0.3, while the unoccluded left region maintains a high weight above 0.8.

[0085] Step S330: Multiply the spatial weight matrix element by element with the initial wearing state feature to obtain a weighted intermediate feature.

[0086] Element-by-element multiplication means multiplying each element value of the spatial weight matrix by the feature values of all channels of the corresponding spatial grid cell to suppress the response intensity of the low-weight regions. In specific implementation, upsample the 4×8 spatial weight matrix to the 32×64 size to align its spatial dimension with the initial wearing state feature map, and then perform matrix dot multiplication by broadcasting in the channel dimension. For example, the weight of a grid in the lower right corner is 0.2, and the feature values of its corresponding 8×8 region on 512 channels are all multiplied by 0.2, which greatly reduces the feature contribution degree of the region blocked by the safety rope. This operation strengthens the discriminative features of the effective reflective regions. For example, in the normal wearing state, after weighting the high-weight grids (0.9 - 0.95) arranged longitudinally in the center, their edge gradient features are more prominent, which helps to analyze the structural integrity of the reflective stripes in the subsequent process.

[0087] Step S340: Perform non-maximum suppression processing on the intermediate feature to remove redundant feature regions with an overlap rate higher than a preset threshold, and generate the target wearing state feature.

[0088] Non-maximum suppression processing is to retain the local maximum response area in the feature space while suppressing the low-confidence areas that highly overlap with it. The preset overlap rate threshold is set to 0.5 according to the minimum spacing of the reflective stripes. Specifically, the sliding window method is used to traverse the intermediate feature map, and only the grid with the maximum weight value is retained within each 8×8 window. If the proportion of the overlapping area of adjacent windows exceeds 50%, the grid with the lower weight is removed. For example, if the weights of two horizontally adjacent grids are 0.92 and 0.87 respectively, and the overlap rate of their coverage areas reaches 60%, then the latter is suppressed and the former is retained. After this processing, the target wearing state features will eliminate the multi-peak interference caused by the wrinkles or reflections of the reflective material, and generate a compact and non-redundant representation of the reflective area. For example, the initially detected 5 discrete high-weight areas are merged into 2 core areas after suppression, accurately reflecting the wearing state of the main part of the reflective vest.

[0089] As an implementation manner, in step S400, according to the matching degree between the target wearing state feature and the preset wearing state threshold, determining the recognition result of the reflective vest wearing state of the human object to be detected includes:

[0090] Step S410: Extract a first statistic related to the coverage area of the reflective material, a second statistic related to the continuity of the reflective stripes, and a third statistic related to the matching degree of the human posture in the target wearing state feature.

[0091] Exemplarily, the first statistic is the percentage of the effective reflective area in the total area of the target wearing state feature, and the calculation method is the ratio of the number of high-weight grids to the total number of grids. For example, if 24 out of 32 grids have a weight ≥ 0.7, the coverage area is 75%. The second statistic is obtained by calculating the spatial connectivity of the high-weight grids. The morphological dilation operation is used to merge adjacent grids into a continuous area, and the proportion of the number of grids in the largest connected area is statistically calculated. For example, if 20 out of 24 high-weight grids form a single connected area, the continuity is 83.3%. The third statistic is obtained by comparing the matching degree between the distribution of the reflective area and the standard human posture template. The template defines the standard positions of the reflective stripes on the front of the torso (such as symmetric distribution on the left and right, longitudinally covering from the shoulders to the waist). Specifically, the average Euclidean distance between the center points of the actual reflective area and the reference points of the template is calculated. When the distance ≤ 5 pixels, the matching degree is 1.0, and it linearly decreases as the distance increases. For example, if the center of the reflective area of a sample deviates from the standard position by 8 pixels, the matching degree drops to 0.6.

[0092] Step S420: Input the first statistic, the second statistic, and the third statistic into a preset multi-condition decision tree model to generate a comprehensive matching score.

[0093] The multi - condition decision tree model is a hierarchical classifier trained based on historical data. The comprehensive matching score is a weighted evaluation value of three types of statistics, with a value range of 0 - 1. Specifically, the first statistic is given a weight of 0.5, the second statistic 0.3, and the third statistic 0.2. The scores are gradually accumulated through the hierarchical judgment of the decision tree. For example, for a certain sample, the first statistic of 0.75 triggers the "≥70%" branch and adds 0.375 points, the second statistic of 0.8 triggers the "≥80%" branch and adds 0.24 points, and the third statistic of 0.7 triggers the "≥60%" branch and adds 0.14 points, with a total score of 0.755. It should be noted that the decision - condition thresholds are dynamically set through cluster analysis. For example, the clustering center of the first statistic for normally - worn samples is 0.85, for partially - occluded samples is 0.55, and for non - worn samples is 0.15, based on which the decision boundaries are divided.

[0094] Step S430: If the comprehensive matching score is higher than the first threshold, it is determined to be in the normal - wearing state.

[0095] The first threshold is set to 0.75 according to the lowest comprehensive score of normally - worn samples. For example, a construction worker wears a reflective vest correctly, with a coverage area of 85%, continuity of 90%, and pose matching degree of 0.95. The comprehensive score is 0.85 * 0.5+0.9 * 0.3 + 0.95 * 0.2 = 0.86, which exceeds the threshold and is determined to be normal. This threshold ensures correct classification can still be maintained under slight light changes or pose offsets (such as when the arm is raised, causing the local grid weight to decrease). For example, a sample with a score of 0.76, although the third statistic is 0.7 due to the shooting angle, can still be determined to be compliant.

[0096] Step S440: If the comprehensive matching score is lower than the second threshold, it is determined to be in the non - wearing state.

[0097] The second threshold is set to 0.3 based on the highest score of non - worn samples. For example, a worker only wears ordinary work clothes, with a reflective material coverage area of 5%, continuity of 0%, and pose matching degree of 0.1. The comprehensive score is 0.05 * 0.5+0 * 0.3 + 0.1 * 0.2 = 0.045, which is far lower than the threshold and is determined to be non - worn. This threshold needs to cover the situation where the reflective vest is completely missing or extremely damaged. For example, a sample with a coverage area of 10% but continuity of 0 still has a score of 0.12 and is classified as non - worn.

[0098] Step S450: If the comprehensive matching score is between the first threshold and the second threshold, it is determined to be in the partially - occluded state.

[0099] The partial occlusion state refers to the situation where the reflective vest is partially covered by an external object but not completely ineffective. For example, the right half of a worker's reflective vest is blocked by a tool bag, with a coverage area of 55%, a continuity of 60% (the left half remains connected), and an attitude matching degree of 0.5. The comprehensive score is 0.55 * 0.5 + 0.6 * 0.3 + 0.5 * 0.2 = 0.275 + 0.18 + 0.1 = 0.555, which is between 0.3 and 0.75, and it is determined as partial occlusion. It should be noted that misjudgments near the boundary in this interval need to be excluded. For example, a sample with a score of 0.74, although close to the normal threshold, still maintains the partial occlusion determination due to a continuity of 85% and an attitude matching degree of 0.8, to avoid misjudging severe occlusion as normal.

[0100] As an implementation manner, the construction process of the multi - condition decision tree model includes the following steps:

[0101] Step S401: Collect historical recognition data, where the historical recognition data includes multiple groups of the first statistic, the second statistic, the third statistic, and the corresponding true wearing state labels.

[0102] The historical recognition data is the recognition records of wearing states accumulated from the deployed system, which needs to cover all combinations of lighting conditions and wearing states. For example, 10,000 records are collected, including 6,000 with normal wearing (the mean of the first statistic is 0.82 and the standard deviation is 0.08), 3,000 with partial occlusion (the mean is 0.52 and the standard deviation is 0.12), and 1,000 with no wearing (the mean is 0.18 and the standard deviation is 0.10). Each piece of data contains the floating - point values of the three statistics and the labels confirmed by manual review. The data needs to be de - duplicated and balanced. For example, the minority class samples (no wearing) are oversampled by SMOTE so that the ratio of the number of samples in the three classes is 1:1:1, to avoid model skewness.

[0103] Step S402: Conduct cluster analysis on the historical recognition data to determine the distribution boundaries of different wearing state categories in the three - dimensional statistic space.

[0104] The cluster analysis uses the K - means algorithm to divide the data into three clusters corresponding to the three wearing states. Specifically, the three - dimensional space coordinate axes are the first statistic (X - axis), the second statistic (Y - axis), and the third statistic (Z - axis), and the sample similarity is measured by the Euclidean distance. For example, the center coordinates of the normal - wearing cluster are (0.85, 0.88, 0.90), the partial - occlusion cluster is (0.55, 0.60, 0.50), and the no - wearing cluster is (0.15, 0.10, 0.12). The distribution boundaries are determined by calculating the decision hyperplane of the adjacent regions between clusters. For example, the boundary plane equation between the normal and partial - occlusion clusters is 0.6X + 0.5Y + 0.4Z = 0.72. When the sample point is substituted into the left - hand side value and ≥0.72, it is classified into the normal class, otherwise it is partial occlusion.

[0105] Step S403: Generate a plurality of decision rules according to the distribution boundary, and each decision rule corresponds to a hyperplane segmentation condition.

[0106] The decision rules divide the three-dimensional space into a plurality of sub-regions, and each sub-region is associated with a specific wearing state. For example, the first rule is "if the first statistic ≥ 0.7 and the second statistic ≥ 0.75 and the third statistic ≥ 0.8, then it is determined to be normally worn"; the second rule is "if 0.4 ≤ the first statistic < 0.7 and the second statistic ≥ 0.5 and the third statistic ≥ 0.4, then it is determined to be partially occluded". The hyperplane equation is obtained by linearly segmenting the inter-cluster samples through a support vector machine (SVM). For example, the segmentation hyperplane between the normal and partially occluded classes is obtained by training with 1000 boundary samples, and the classification interval is maximized to ensure the generalization ability.

[0107] Step S404: Optimize the order arrangement of the decision rules based on the Gini coefficient, and construct a hierarchical decision node until each leaf node contains only samples of a single wearing state category.

[0108] The Gini coefficient is used to measure the degree of reduction of data impurity by the decision rule, and the optimization goal is to select the rule that can maximize the improvement of the subsequent branch purity to be executed first. For example, the initial node contains all data, and the Gini coefficient is 0.6 (a mixture of three categories). First, apply the rule "the first statistic ≥ 0.7", and divide the data into two subsets: the Gini coefficient of the left subset is 0.2 (the normal class accounts for 90%), and the Gini coefficient of the right subset is 0.5 (a mixture of partially occluded and not worn). Continue to apply the rule "the second statistic < 0.4" to the right subset, and the Gini coefficient drops to 0.1 (the not worn class accounts for 95%). The finally formed decision tree contains 5 layers of nodes, and the sample category purity of each leaf node ≥ 95%. This structure ensures efficient classification. For example, normally worn samples can reach the leaf node after an average of 2 judgments, while complex partially occluded samples require 4 judgments to exclude the possibilities of not worn and normal.

[0109] As an implementation manner, the method further includes a dynamic correction process for the recognition result, which may specifically include the following steps:

[0110] Step S500: When the recognition result is in a partially occluded state, obtain a continuous multi-frame historical monitoring image of the human object to be detected.

[0111] Continuous multi-frame historical monitoring images refer to at least 10 frames of image sequences collected at fixed time intervals before the recognition result is triggered. The time window length is set to 5 seconds based on the human body movement frequency, and the frame rate is 2 frames per second to ensure time continuity. Specifically, in the construction site scene, a worker's reflective vest right shoulder is judged to be partially blocked due to temporary carrying of a tool bag. The system immediately retrieves 10 frames of images collected within the previous 5 seconds (time stamps are t-4.5s to t-0.5s), and each frame of the image is pre-processed by human body detection and reflective vest area alignment. It should be noted that historical images must meet the spatiotemporal alignment conditions, that is, the pixel offset caused by human body movement is eliminated by optical flow method or feature point matching. For example, the SIFT algorithm is used to extract the human body key points in each frame, and the reflective vest area of each frame is mapped to a unified coordinate system through affine transformation to ensure the spatial consistency of temporal feature extraction. If the alignment fails due to violent movement (such as translation of more than 50 pixels or rotation of more than 15 degrees), the frame is discarded and traced back to the image frame that meets the alignment conditions until the minimum frame number threshold is collected.

[0112] Step S600: extracting time series features from the historical monitoring images to generate a dynamic change trend of the wearing status of the reflective vest.

[0113] Temporal feature extraction is the process of quantifying the evolution of the area, position and morphological parameters of the partially occluded area from the multi-frame aligned images. The dynamic change trend is fitted by linear regression or sliding average algorithm to the change slope and fluctuation amplitude of the time series data. Specifically, the pixel area ratio of the partially occluded area is calculated frame by frame (such as 15% of the occlusion ratio in the t-4.5s frame, 18% in the t-3.5s frame, ..., 35% in the t-0.5s frame), and the moving trajectory of the centroid of the occluded area is recorded (such as the X coordinate increases from 120 pixels to 150 pixels, and the Y coordinate remains unchanged at 80 pixels). The area change curve is fitted by the least squares method, and the slope is +4% / s, which is judged as a continuous expansion trend; if the slope fluctuates within ±1% / s, it is considered to be a stable state. Further analysis of morphological parameters, such as the aspect ratio of the occluded area gradually changes from 2.1 (vertical stripes) to 1.3 (approximately circular), indicating that the toolkit has changed its occlusion morphology due to human body shaking. This dynamic trend is encoded into three sets of time series vectors: area change rate, center of mass displacement rate, and morphological distortion index, which are input into the preset trend classifier for state evolution pattern recognition.

[0114] As an implementation manner, the step S600, extracting time series features from the historical monitoring image to generate a dynamic change trend of the wearing state of the reflective vest, includes:

[0115] Step S610: Track key points of the reflective vest area in multiple consecutive frames of historical monitoring images to determine the motion trajectory of the blocked area.

[0116] Key point tracking refers to tracking the displacement trajectory of specific spatial points in the reflective vest area in a time series of images through a feature matching algorithm. The motion trajectory of the occluded area is composed of the dynamic coverage change path of the occluded reflective material. Specifically, in the construction site monitoring scenario, the right shoulder of a worker's reflective vest is gradually covered by a moving safety rope. The system collects 10 consecutive frames of images at a rate of 5 frames per second (time span of 2 seconds). First, in the first frame of the image, the intersection points of the high-reflection stripes in the reflective vest area (such as the intersection of horizontal and vertical reflective strips) and the center points of the dense areas of reflective particles are located as the initial key point set, and a total of 32 key points are marked. Subsequently, the KLT (Kanade-Lucas-Tomasi) optical flow algorithm is used to calculate the pixel displacement vectors of the key points in each pair of adjacent frames. For example, between the first frame and the second frame, the key point numbered K15 moves from the coordinates (120, 80) to (122, 81), and the displacement vector is (+2, +1). By accumulating the displacement vectors of consecutive frames, the motion trajectory curve of the key points is constructed. For example, the trajectory of K15 within 10 frames shows a linear movement pattern from the lower left to the upper right. It should be noted that when screening key points, instantaneous noise points caused by changes in the reflective characteristics of the reflective material need to be excluded. For example, when the coordinates of a key point jump by more than 20 pixels due to overexposure in the third frame, it needs to be corrected or removed through a trajectory smoothing algorithm. This process can accurately depict the expansion direction and speed of the occluded area (such as the safety rope) in the time dimension, providing basic spatial motion data for subsequent trend analysis.

[0117] As an implementation manner, step S610 of performing key point tracking on the reflective vest area in a plurality of consecutive historical monitoring images to determine the motion trajectory of the occluded area includes:

[0118] Step S611: Based on the image of the reflective vest area of the current frame, detect a key point set related to the geometric distribution of the reflective stripes, where the key point set includes the intersection points of the edges of the reflective material and the center points of the high-reflectivity areas.

[0119] The key points related to the geometric distribution of reflective stripes refer to the characteristic points with fixed spatial relationships in the standard design of reflective vests, such as the intersection points of horizontal and vertical stripes, the points with the maximum curvature at the ends of reflective strips, and the centroid of the high-reflection particle aggregation area. Specifically, the Harris corner detection algorithm is used to calculate the eigenvalues of the pixel gray-scale gradient matrix in the current frame of the reflective vest area image, and the pixel points with a response value exceeding the threshold of 0.05 are marked as candidate key points. For example, in a reflective vest area with a size of 256×128 pixels, 48 candidate points are detected, and 32 key points with uniform spatial distribution are retained after non-maximum suppression, including the intersection point of the 3rd horizontal and 2nd vertical reflective strips (coordinates (80,60)), the centroid of the reflective particle area in the lower right corner (coordinates (220,100)), etc. It should be noted that the key point detection needs to adapt to different lighting conditions. For example, in a low-illumination environment, an adaptive threshold adjustment strategy is adopted to dynamically reduce the Harris threshold to 0.03 to maintain the detection sensitivity, and at the same time, the broken parts of the reflective strips are filled through morphological closing operations to ensure the integrity of geometric features.

[0120] Step S612: Perform optical flow estimation on the set of key points and the corresponding reflective vest area in the previous frame of historical monitoring image to obtain the displacement vector of each key point between adjacent frames.

[0121] Optical flow estimation refers to calculating the pixel-level displacement of key points between adjacent frames based on the assumptions of brightness constancy and small motion. The displacement vector contains components in the horizontal and vertical directions. Specifically, the pyramid Lucas-Kanade algorithm is used to construct a three-layer image pyramid (scaling factor 0.5), and iterative optical flow calculations are performed on each layer to improve the accuracy of large-displacement tracking. For example, the key point K09 had coordinates (150,90) in the previous frame, and it was found to move to (155,88) in the current frame after pyramid calculation, and the displacement vector is (+5,-2). For the center point of the high-reflectivity area, due to the possible brightness mutation caused by the reflective characteristics, reverse optical flow verification needs to be introduced: that is, trace back from the current frame to the previous frame, and if the two-way tracking error exceeds 1 pixel, it is determined as a failure point and excluded. For example, a certain key point has a forward optical flow displacement of (+6,+3), and the reverse tracking error is (+1,-1), which meets the error tolerance, and the displacement vector is valid; if the error reaches (+4,-2), it is regarded as a mismatching and excluded.

[0122] Step S613: Screen out the matching key point pairs according to the displacement vector, and construct an initial motion trajectory segment based on the coordinate offsets of the matching key point pairs.

[0123] A matching key-point pair refers to a pair of points that are successfully tracked between two frames and whose displacements conform to the constraints of rigid body motion. An initial motion trajectory segment is composed of the displacement sequences of the same key point in consecutive frames. Specifically, the results of optical flow estimation are screened by RANSAC (Random Sample Consensus), a global motion model (such as an affine transformation) is fitted, and the outlier points that deviate from the model by more than 2 pixels are removed. For example, among 50 candidate displacement vectors, 42 conform to the affine transformation model (translation + rotation), and the remaining 8 are removed due to abnormal displacements caused by local occlusion. Subsequently, the remaining key points are linked in chronological order to form an initial trajectory segment: the displacement sequence of key point K22 in frames 1 - 5 is [(0,0), (+2,+1), (+3,+2), (+5,+3), (+7,+4)], showing a uniform motion pattern in the lower right direction. This process can effectively suppress trajectory breaks caused by specular reflection, flicker, or temporary occlusion. For example, if a key point is lost in frame 3 due to a short-term overexposure, its trajectory segment can be completed through an interpolation algorithm.

[0124] Step S614: Perform a continuity check on the initial motion trajectory segment, remove the abnormal trajectory points whose displacement directions do not conform to the human body motion posture constraints, and generate a corrected candidate motion trajectory.

[0125] Continuity check means verifying whether the trajectory displacement direction conforms to the motion laws of parts such as the torso and arms according to the human joint kinematic model. Specifically, a human body posture constraint rule library is constructed: the horizontal movement speed of the shoulder ≤ 5 pixels / frame, the vertical movement speed ≤ 3 pixels / frame; the change rate of the waist rotation angle ≤ 2° / frame. For example, if a displacement of (+15,+0) occurs between frames 4 - 5 in a certain trajectory segment, exceeding the shoulder horizontal speed threshold, it is determined as abnormal (possibly caused by specular artifacts) and needs to be removed from the candidate trajectory. The corrected candidate motion trajectory needs to meet the condition of motion smoothness: the change in the angle between adjacent frame displacement vectors ≤ 30°, and the acceleration ≤ 5 pixels / frame². For example, the displacement of K22 in frame 3 of the corrected trajectory is adjusted from (+3,+2) to (+4,+2), reducing the angle between frames 2 - 3 from 18° to 12°, meeting the smoothness requirement.

[0126] Step S615: According to the spatial distribution density of the candidate motion trajectory in consecutive multiple frames, extract the motion direction consistency parameter and the area change rate of the trajectory coverage area in the occlusion area.

[0127] The motion direction consistency parameter refers to the average angular variance of the displacement vectors of all candidate trajectories, reflecting the degree of concentration of the overall movement direction in the occluded area; the area change rate of the trajectory coverage area is obtained by calculating the time derivative of the convex hull area of the trajectory points. Specifically, in 10 frames of images, the average displacement angle of 20 candidate trajectories is 85°, the variance is 5°, and the direction consistency parameter is 0.92 (1 - variance / 90°); the convex hull area of the trajectory points increases from 200 pixels² in frame 1 to 800 pixels² in frame 10, and the area change rate is (600 / 500) / 2 = 0.6 / s. This parameter can distinguish different types of occlusions: the sliding of the tool kit results in high direction consistency (>0.9) and continuous area growth, while the fluttering of the clothes by the wind may result in a direction variance >20° and area fluctuations.

[0128] Step S616: Based on the motion direction consistency parameter and the area change rate of the trajectory coverage area, fit the motion trend curve of the occluded area, and determine the starting position, ending position, and path shape of the motion trajectory.

[0129] The motion trend curve refers to the spatio-temporal distribution function of the trajectory points fitted by polynomial regression or spline interpolation, and the path shape is characterized by the curvature and extension direction of the curve. Specifically, the displacement data of 20 candidate trajectories are fitted using a cubic polynomial to obtain the parametric equations x(t) = 0.5t² + 2t, y(t) = 0.2t³ - 0.1t² + 3t, R² = 0.98, indicating that the occluded area expands along an approximate parabola. The starting position is the average value of the trajectory points in frame 1 (120 ± 5, 80 ± 3), the ending position is (250 ± 8, 110 ± 5) in frame 10, and the path shape is detected as linear (curvature < 0.01) by the Hough transform. This fitting result is used to predict future motion: if the trend curve shows continuous extension towards the core area of the reflective vest (such as the center of the chest), a warning is triggered; if the path turns towards the edge and the curvature increases, it may be a temporary occlusion.

[0130] Step S620: Calculate the area change rate and shape similarity between adjacent frames of the occluded area.

[0131] The area change rate between adjacent frames refers to the relative increase or decrease ratio of the pixel area of the occluded region per unit time. The shape similarity measures the morphological consistency of the occluded region contours between consecutive frames through the Hausdorff distance. Specifically, in the t-th frame and the (t + 1)-th frame, the binary masks of the occluded regions determined by key-point tracking are extracted respectively. The difference in their pixel areas is calculated as ΔA = |A_{t + 1}-A_t|, and the change rate R = ΔA / ((A_t+A_{t + 1}) / 2) / Δt, where Δt is the frame interval time of 0.2 seconds. For example, if the area of a certain occluded region increases from 1500 pixels in the 3rd frame to 1800 pixels in the 4th frame, the change rate R=(300 / 1650) / 0.2≈0.909 / s, indicating a rapid expansion of the occlusion. When calculating the shape similarity, the contours of the two-frame masks are discretized into sequences of polygon vertices, and the Hausdorff distance of the bidirectional maximum and minimum vertex distances is calculated. For example, the contour distance between the 5th frame and the 6th frame is 8 pixels, and after normalization, the similarity S = 1-(8 / 256)=0.969 (image size 256×256). This metric can effectively distinguish the stability of the occlusion morphology. For example, the bar-shaped occlusion caused by the sliding of a safety rope and the triangular occlusion formed by the temporary blowing of a corner of a clothes by the wind. The former has a shape similarity higher than 0.9, while the latter may be lower than 0.7.

[0132] Step S630: Construct a time-series feature vector based on the motion trajectory, area change rate, and shape similarity.

[0133] A time-series feature vector refers to encoding multi-dimensional time-series data into a numerical vector of a fixed dimension, which is used to characterize the pattern characteristics of occlusion evolution. Specifically, within a 2-second time window, the motion trajectory direction angle, area change rate, and shape similarity are collected every 0.2 seconds, generating a total of 10 sets of original data. First, a sliding window average filter (window size 3 frames) is applied to the motion trajectory direction angle to eliminate instantaneous jitter noise. For example, the direction angle sequence [85°, 88°, 82°] in the 2nd - 4th frames is filtered to output [85°, 85°, 83°]. Subsequently, the area change rates are arranged in chronological order as [0.2, 0.5, 0.9, 1.1, 0.8,...], and the variance value in the first-order difference sequence is calculated to reflect the fluctuation intensity of the change rate. The shape similarity sequence extracts the main frequency component through the fast Fourier transform. For example, a periodic fluctuation of 0.5 Hz is detected, indicating that the occlusion morphology is affected by regular external forces (such as equipment vibration). Finally, 12 statistics such as the direction angle mean, area change rate variance, and main frequency amplitude are aligned according to the time stamps to generate a 10×12-dimensional feature matrix, which is then mapped to the [0,1] interval through maximum-minimum normalization to form a 120-dimensional time-series feature vector. This vector can comprehensively describe the trend, volatility, and morphological rules of the occlusion motion. For example, a continuously expanding occlusion is characterized by a stable direction angle, a positive cumulative area change rate, and a high shape similarity.

[0134] As an implementation manner, in step S630, according to the motion trajectory, the area change rate, and the shape similarity, a time series feature vector is constructed, including:

[0135] Step S631: Perform segmented sampling on the direction change pattern of the motion trajectory, extract the sequence of direction angle offsets within each time window, and generate direction change pattern parameters.

[0136] The direction change pattern parameters refer to the mean, variance, and autocorrelation characteristics of the trajectory direction angle statistically calculated within a sliding time window, quantifying the regularity of the motion. Specifically, set the window size to 3 frames (0.6 seconds) and the step size to 1 frame, and perform segmented sampling on the 10-frame direction angle sequence [82°, 85°, 83°, 87°, 84°, 88°, 85°, 86°, 84°, 87°] to obtain 8 window data. For example, the mean of the direction angles in windows 1 - 3 is 83.3°, the variance is 1.25, and the autocorrelation coefficient is 0.78; the mean in windows 2 - 4 is 85°, the variance is 2.0, and the autocorrelation coefficient is 0.65. Concatenate the 3 statistics of each window to generate a 24-dimensional direction feature sub-vector, reflecting the stability of the direction change. For example, the autocorrelation coefficient of a continuously linear motion window > 0.7, while that of a random fluctuation < 0.3.

[0137] Step S632: Based on the time series data of the area change rate, calculate the relative change gradient between adjacent time windows, and accumulate the area increase and decrease trends at different time scales to generate the area cumulative change amount.

[0138] The area cumulative change amount refers to integrating the area change rate at multiple time scales (such as short-term 1 second, medium-term 2 seconds, long-term 3 seconds), characterizing the persistence of occlusion expansion. Specifically, for the area change rate sequence [0.2, 0.5, 0.9, 1.1, 0.8, 1.2, 1.3, 1.4, 1.5, 1.6] / s, calculate its cumulative amount at the 1-second scale: the cumulative amount of the first 5 frames (1 second) is 0.2 + 0.5 + 0.9 + 1.1 + 0.8 = 3.5, and the cumulative amount of the last 5 frames is 1.2 + 1.3 + 1.4 + 1.5 + 1.6 = 7.0, with an increase rate of 100%. At the same time, calculate the gradient between adjacent windows: the gradient of windows 1 - 2 is (0.5 - 0.2) / 0.2 = 1.5, and the gradient of windows 2 - 3 is (0.9 - 0.5) / 0.2 = 2.0, reflecting the increase in the growth rate. This parameter can distinguish between short-term fluctuations (alternating positive and negative gradients) and continuous growth (monotonically increasing gradients).

[0139] Step S633: Perform frequency domain conversion on the continuous fluctuations of the shape similarity, extract the energy distribution ratio of the low-frequency component and the high-frequency component, and generate a shape stability index.

[0140] Exemplarily, the shape stability index decomposes the shape similarity time series signal into a spectrum through fast Fourier transform, and calculates the ratio of low-frequency energy from 0 to 0.5 Hz to high-frequency energy from 0.5 to 2 Hz. For example, after the FFT of a 10-frame shape similarity sequence [0.95, 0.94, 0.93, 0.92, 0.91, 0.90, 0.89, 0.88, 0.87, 0.86], the low-frequency energy accounts for 85% and the high-frequency energy accounts for 15%. The index value is 85 / 15 ≈ 5.67, indicating that the morphological change is slow and stable. If the sequence contains [0.7, 0.9, 0.6, 0.8, 0.5, …], then the high-frequency energy accounts for > 40% and the index value < 1.5, indicating violent morphological fluctuations.

[0141] Step S634: Align the direction change pattern parameters, the area cumulative change amount, and the shape stability index according to the time stamp, and perform normalization processing to generate a standardized time series data block.

[0142] The standardized time series data block refers to aligning multi-dimensional heterogeneous data according to a unified time reference, and eliminating the dimensional difference through Z-score normalization. Specifically, at the time stamp t = 2.0 s, the direction feature sub-vector is [83.3°, 1.25, 0.78], the area cumulative amount is 3.5, and the shape stability index is 5.67, which are normalized to [0.75, 0.12, 0.81], 0.62, and 0.92 (based on the global maximum value) respectively. The normalized data is arranged in chronological order as a 10×5 matrix, and missing values (such as when the optical flow tracking fails for a certain frame) are filled using linear interpolation or the mean value of adjacent frames.

[0143] Step S635: Concatenate the direction change pattern parameters, the area cumulative change amount, and the shape stability index corresponding to the same time stamp in the standardized time series data block to form a multi-dimensional feature segment.

[0144] The multi-dimensional feature segment is a vertical concatenation vector of all features at each time stamp. For example, at t = 2.0 s, the direction parameter [0.75, 0.12, 0.81], the area cumulative amount 0.62, and the shape index 0.92 are concatenated into a 5-dimensional vector [0.75, 0.12, 0.81, 0.62, 0.92]. The vectors at 10 consecutive time stamps are vertically stacked into a 10×5 matrix to form a basic feature segment.

[0145] Step S636: Perform sliding window aggregation on the multi-dimensional feature segments of consecutive time stamps, and statistically calculate the mean value, variance, and maximum value of each dimension within the window to generate the time series feature vector.

[0146] Sliding window aggregation means traversing the basic feature segments with a window size of 3 and a step size of 1, and calculating the statistics of each feature within each window. For example, the mean parameter of the direction in windows 1 - 3 is (0.75 + 0.68 + 0.72) / 3 ≈ 0.72, the variance is 0.03; the maximum value of the area cumulative amount is 0.65. The statistics of the 3 windows are concatenated as (0.72, 0.03, 0.65, …) × 3 windows = 45 dimensions, and finally a 45 × 3 = 135 - dimensional time - series feature vector is generated. This vector enhances the model's ability to recognize trend patterns by capturing the time - domain statistical characteristics of the features. For example, the variance of the continuously increasing area cumulative amount approaches 0, while the variance of the fluctuating state increases significantly.

[0147] Step S640: Input the time - series feature vector into a long short - term memory network to predict the evolution path of the occluded area in the next several frames, and generate the dynamic change trend according to the evolution path.

[0148] The long short - term memory network refers to a recurrent neural network structure with a gating mechanism, which is used to model the long - term dependencies in time - series data. The evolution path prediction iteratively outputs the occlusion parameters at future time steps in an autoregressive manner. Specifically, a network model containing 3 layers of LSTM units is constructed. The input layer receives a 120 - dimensional feature vector, the hidden layer dimension is 256, and the output layer generates the predicted values of the direction angle, area change rate, and shape similarity for the next 5 frames (1 second). During training, 10,000 groups of time - series samples in the historical data are used to optimize the network parameters through the mean square error loss function. For example, when inputting the feature vector of a continuously sliding safety rope, the network outputs the predicted area change rate for the next 5 frames as [1.2, 1.3, 1.4, 1.5, 1.6] / s, the direction angle remains at 82° ± 2°, and the shape similarity remains above 0.91. Based on this, a dynamic trend curve of the linearly expanding occluded area along a fixed direction is generated. This trend will be used to judge whether the recognition result needs to be corrected: if the predicted area exceeds the 60% threshold within 3 seconds and the direction points to the core reflective area, the un - worn state correction is triggered; if the predicted area drops below 10% and the shape returns to stability, it is corrected to the normal wearing state.

[0149] Step S700: If the dynamic change trend indicates that the occluded area continues to expand, correct the recognition result to the un - worn state.

[0150] Exemplarily, continuous expansion means that the area of the occluded region shows a monotonically increasing trend within a time window and the average growth rate exceeds a preset threshold (e.g., ≥3% / s), while the centroid moves towards the core region of the reflective vest (e.g., moving more than 20 pixels from the edge to the center). For example, due to improper wearing, a worker's reflective vest gradually slips off within 5 seconds, the occluded area linearly increases from 10% to 60%, the centroid moves from (100, 80) to (160, 90), the growth rate reaches 10% / s and the moving direction is opposite to the normal coverage area of the vest. At this time, the system determines that the occlusion is irreversible and the overall reflection function fails, corrects the partial occlusion state to the non-wearing state, and triggers a secondary alarm signal. It should be noted that the correction logic needs to exclude temporary occlusion interference (such as a waving action briefly covering the reflective strip), and improve the judgment robustness by setting the minimum number of consecutive frames for the growth rate persistence (such as the growth rate ≥2% / s for 4 consecutive frames) to avoid incorrect correction.

[0151] Step S800: If the dynamic change trend indicates that the occluded area returns to the normal range within a preset time, correct the recognition result to the normal wearing state.

[0152] The preset time is set to 8 seconds according to the reasonable duration of the human body adjustment action, and the normal range means that the occluded area ≤10% and the centroid is located in the non-core area at the edge of the reflective vest (e.g., the distance from the boundary ≤15 pixels). Specifically, after a worker carries a tool bag for 3 seconds and actively adjusts to the side waist position, the occluded area drops from the peak value of 40% to 8% within the subsequent 5 frames, and the centroid withdraws from (140, 85) to (180, 120) (the lower right corner boundary coordinate of the reflective vest is (200, 130)). The system detects that the area regression rate reaches -6.4% / s and the centroid displacement direction deviates from the core area, and the dynamic trend classifier outputs a "return to normal" signal, triggering the correction of the recognition result. It should be noted that the recovery determination needs to meet two conditions: one is that the difference between the occluded area of the final frame and the initial normal state does not exceed 5 percentage points; the other is that there is no fluctuation in the secondary occlusion expansion during the recovery process (such as the area rising by more than 3 percentage points) to ensure the stability of the correction result. For example, due to the area fluctuation (35%→28%→40%→15%) caused by the safety rope blown by the wind for a temporary occlusion, the observation window needs to be extended to 10 seconds due to the rising phenomenon until the fluctuation disappears and stabilizes within the threshold before the correction is executed.

[0153] As an optional implementation manner, after step S400, according to the matching degree between the target wearing state feature and the preset wearing state threshold, to determine the recognition result of the reflective vest wearing state of the to-be-detected human object, the method may further include:

[0154] Step S400A: When the recognition result is in a partially occluded state, extract the local occlusion features related to the contour of the occluded area from the target wearing state features, and obtain the auxiliary verification images synchronously captured by the multi-angle acquisition devices of the real-time monitoring image.

[0155] The local occlusion features are sub-feature vectors segmented from the target wearing state features and related to the geometric shape and reflection characteristics of the occluded area. The auxiliary verification images refer to the multi-view image set captured at the same time by the optical acquisition devices deployed at different spatial positions. Specifically, in the construction site safety monitoring system, when the main camera detects an occluded area with an area ratio of 35% on the right shoulder of a worker's reflective vest, the system immediately triggers three auxiliary cameras distributed on the left, top, and rear of the working area to synchronously collect images. The local occlusion features of the main camera include the boundary coordinates of the occluded area (x1 = 180, y1 = 90, x2 = 220, y2 = 120), the average reflection intensity value of 85 (in the range of 0 - 255), and the contour polygon vertex sequence. The resolution and frame rate of the auxiliary verification images are the same as those of the main camera to ensure the spatio-temporal alignment accuracy. For example, the left camera captures the side image of the worker at a 45-degree angle, the top-down camera records the flat-expanded state of the reflective vest, and the rear camera supplements the view of the back area, forming a four-way synchronous image stream with the timestamp error controlled within ±10 ms, providing a multi-dimensional data source for cross-view verification.

[0156] Step S400B: Perform perspective alignment processing on the auxiliary verification images to generate multi-angle verification images consistent with the human posture in the real-time monitoring image, and extract the corresponding reflective vest area verification features from the multi-angle verification images respectively.

[0157] The perspective alignment processing maps the multi-angle images to a unified human coordinate system through 3D pose estimation and affine transformation. The reflective vest area verification features refer to the standardized feature vectors extracted from each perspective image. Specifically, the OpenPose algorithm is used to extract the human bone key points (such as shoulders, hips, and knees) from the main camera image, construct a 3D pose model, and calculate the projection matrix in the view of the auxiliary camera. For example, when the worker is in a standing posture in the main view and the left arm is raised 30 degrees, causing wrinkles in the reflective vest, after the projection transformation of the left camera image, the reflective vest area is corrected to the flat-expanded state consistent with the main view. Subsequently, the same reflective vest area detection process is applied to the corrected auxiliary images: the boundary coordinates of the reflective area are detected in the top-down image (x1 = 175, y1 = 85, x2 = 215, y2 = 115), and the average reflection intensity is 82; in the rear image, due to the backpack occlusion, the area ratio of the reflective area drops to 25%. The verification features of each perspective are normalized to generate 512-dimensional feature vectors with the same dimension for cross-view consistency verification.

[0158] Step S400C: Perform cross - perspective feature matching between the local occlusion feature and the reflective vest area verification feature, and calculate the visibility confidence of the occluded area at different perspectives.

[0159] Cross - perspective feature matching refers to measuring the semantic consistency between the local occlusion feature and the verification features of each perspective through cosine similarity. The visibility confidence reflects the reproducibility probability of the occluded area at multiple perspectives. Specifically, the similarity between the local occlusion feature of the main perspective and the verification feature of the left perspective is 0.78, 0.85 with the top perspective, and 0.35 with the rear perspective. The visibility confidence CV is calculated as the weighted harmonic mean of the similarities of each perspective, and the weights are determined by the angle between the camera spatial position and the main perspective: the weight of the left perspective is 0.4 (angle 45 degrees), the top perspective is 0.3 (90 degrees), and the rear perspective is 0.3 (180 degrees). Then CV=(0.4×0.78 + 0.3×0.85 + 0.3×0.35) / (0.4 + 0.3 + 0.3)=0.66. It should be noted that when the similarity of a certain perspective is lower than 0.2, the data of that perspective is determined to be invalid and the weights are re - allocated. For example, when the similarity of the rear perspective is 0.1 due to complete occlusion, its weight is transferred to the valid perspectives to ensure the robustness of the confidence calculation.

[0160] Step S400D: If the visibility confidence is lower than the preset verification threshold, it is determined that the occluded area is covered by a temporary interference object, and the recognition result is corrected to the normal wearing state.

[0161] The preset verification threshold is set to 0.7 according to the analysis of historical false alarm rates. Temporary interference object coverage refers to non - fixed occluders that exist briefly (such as splashing water stains, flying insects passing by briefly). For example, when a worker is welding, the main camera detects 15% occlusion in the chest reflective area. The similarities between the verification features of the left and top perspectives are 0.85 and 0.88 respectively, and the similarity of the rear perspective is 0.25 due to the interference of arc strong light. The visibility confidence CV after weight adjustment is 0.72, still lower than the threshold 0.7. The system determines it as an instantaneous light spot interference rather than an entity occluder, corrects the recognition result to the normal wearing state, and suppresses the triggering of false alarms. The correction logic needs to combine time - persistence verification: if CV of three consecutive frames is lower than the threshold and the standard deviation of the fluctuation of the occluded area area > 10%, then correction is performed; if only a single frame is lower than the threshold but the area is stable, the partial occlusion determination is maintained.

[0162] Step S400E: If the visibility confidence is higher than the preset verification threshold, based on the shape consistency parameter of the occluded area in the multi - perspective verification images, generate a speculation result of the occluder type, and perform a similarity search on the speculation result with a preset occluder database to determine whether the occluder is an acceptable safety accessory device.

[0163] The shape consistency parameter calculates the morphological matching degree of the multi-view occlusion area contour through the Hausdorff distance. The occlusion object database contains typical contour templates of preset safety devices such as safety rope buckles, tool kits, walkie-talkies, etc. For example, the Hausdorff distance between the occlusion area of the main view and the left view is 5 pixels (similarity 0.95), and that with the top view is 8 pixels (0.92). The shape consistency parameter SC = 1 - (5 + 8) / (256 + 256) = 0.97. Matching the sequence of contour vertices with the database, the similarity of the tool kit template is 0.89, that of the safety rope buckle is 0.65, and that of the walkie-talkie is 0.42. It is determined that the occlusion object type is a tool kit. The acceptable device list in the database includes essential protective equipment such as safety rope buckles and breathing masks. If the matching result belongs to this list and the similarity > 0.8, it is regarded as a compliant occlusion; otherwise, further inspection is triggered. For example, the tool kit is not included in the acceptable list, and even though the similarity is 0.89, it is still determined as an unacceptable device.

[0164] Step S400F: When the occlusion object is determined to be an unacceptable safety accessory device, generate alarm trigger data including the position and type information of the occlusion object, and bind the alarm trigger data with the timestamp and spatial coordinates of the real-time monitoring image to generate a structured alarm record.

[0165] The alarm trigger data includes the occlusion object bounding box coordinates (x1 = 180, y1 = 90, x2 = 220, y2 = 120), type label "tool kit", confidence score 0.89, and multi-view verification image index. The timestamp is accurate to the millisecond level (e.g., 2023-09-15 14:23:45.789), and the spatial coordinates are converted into the construction site three-dimensional coordinate system (X = 35.2m, Y = 12.7m, Z = 1.5m) through the laser rangefinder and the internal and external parameter matrices of the camera. The structured alarm record is encapsulated in JSON format and contains the above fields and the original image storage path. For example, the following encapsulation example can be referred to:

[0166] json

[0167] {

[0168] "alert_id": "20230915142345789_001",

[0169] "timestamp": "2023-09-15 14:23:45.789",

[0170] "coordinates": {"X":35.2, "Y":12.7, "Z":1.5},

[0171] "obstruction_type": "toolkit",

[0172] "confidence": 0.89,

[0173] "image_paths": [" / cam1 / frame_789.jpg", " / cam2 / frame_789.jpg"],

[0174] "status": "pending"

[0175] }

[0176] This record is written into the distributed time-series database in real time, supporting millisecond-level concurrent writing and multi-condition retrieval, and providing a standardized data interface for subsequent auditing and behavior analysis.

[0177] Step S400G: Sort the structured alarm records by priority, send warning signals of different levels to the target terminal according to the preset alarm response rules, and store the alarm records in the historical database for subsequent calls by the behavior analysis model.

[0178] The priority sorting is dynamically calculated based on the risk level of the obstruction type, the sensitivity of the spatial location, and the duration. The preset rule library defines that the risk level of toolkit obstruction is level 2 (medium), which is upgraded to level 1 (high risk) if it is located in the high-altitude operation area (Z > 5m), triggering an audible and visual alarm and pushing it to the handheld terminal of the on-site safety officer; if it is located in the ordinary ground area and the duration < 30 seconds, it is marked as level 3 (low risk), only logging without triggering a real-time alarm. For example, in an alarm record, if the toolkit obstruction is located in the scaffolding area 8 meters above the ground, the system immediately activates the warning light in this area to flash red and sends a warning message including the location map to the safety officer's PDA within a radius of 50 meters. The historical database adopts a columnar storage structure and establishes a composite index according to time partitions (day / month / year) and spatial grids (10m × 10m grids). The behavior analysis model can efficiently retrieve alarm records within a specific time period and area, mine the spatio-temporal distribution law of illegal wearing patterns. For example, the incidence rate of toolkit obstruction events from 9 to 10 am on Mondays is 300% higher than that in other time periods, and based on this, the pre-job equipment inspection process can be optimized.

[0179] Based on the same principle as the method shown in Figure 1 In the embodiment of the present invention, a reflective vest wearing state recognition device 10 based on deep learning is also provided. As shown in Figure 2 shown, the device 10 includes:

[0180] An image acquisition module 11, configured to acquire a real-time monitoring image of the target scene, where the real-time monitoring image includes at least one human object to be detected;

[0181] A feature extraction module 12 is configured to extract an image of a reflective vest area corresponding to the human object to be detected from the real-time monitoring image, and perform multi-scale feature extraction on the image of the reflective vest area to generate an initial wearing state feature.

[0182] A feature analysis module 13 is configured to perform regional distribution analysis on the initial wearing state feature based on a pre-trained wearing state recognition model to generate a target wearing state feature. The pre-trained wearing state recognition model is obtained by fusing multi-modal training data, and the multi-modal training data includes sample images of reflective vests under different lighting conditions.

[0183] A state recognition module 14 is configured to determine a recognition result of the wearing state of the reflective vest of the human object to be detected according to the matching degree between the target wearing state feature and a preset wearing state threshold. The recognition result includes normal wearing, not wearing, or partially blocked states.

[0184] The above embodiments introduce the reflective vest wearing state recognition device 10 based on deep learning from the perspective of virtual modules. The following introduces a computer system from the perspective of physical modules, as follows:

[0185] An embodiment of the present invention provides a computer system, as Figure 3 shown, the computer system 100 includes a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the computer system 100 may further include a transceiver 104. It should be noted that in practical applications, the transceiver 104 is not limited to one, and the structure of the computer system 100 does not constitute a limitation to the embodiments of the present invention.

[0186] The processor 101 may be a CPU, a general-purpose processor, a GPU, a DSP, an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present invention. The processor 101 may also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0187] The bus 102 may include a path for transmitting information between the above components. The bus 102 may be a PCI bus or an EISA bus, etc. The bus 102 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0188] The memory 103 can be a ROM or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or an EEPROM, a CD-ROM, or other optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0189] The memory 103 is used to store the application program code for implementing the solution of the present invention and is controlled by the processor 101 for execution. The processor 101 is used to execute the application program code stored in the memory 103 to implement the content shown in any of the foregoing method embodiments.

[0190] An embodiment of the present invention provides a computer system. The computer system in the embodiment of the present invention includes: one or more processors; a memory; one or more computer programs, where one or more computer programs are stored in the memory and are configured to be executed by one or more processors. When the one or more programs are executed by the processor, the above method is implemented.

[0191] An embodiment of the present invention provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program runs on the processor, the processor can execute the corresponding content in the foregoing method embodiments.

[0192] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps is not strictly limited in order and can be executed in other orders. Moreover, at least some of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0193] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for recognizing the wearing state of a reflective vest based on deep learning, characterized in that, The method includes: Obtaining a real-time monitoring image of a target scene, where the real-time monitoring image includes at least one human object to be detected; Extracting a reflective vest region image corresponding to the human object to be detected from the real-time monitoring image, and performing multi-scale feature extraction on the reflective vest region image to generate an initial wearing state feature; Dividing the initial wearing state feature into multiple spatial grid units, and calculating the similarity between the feature vector of each spatial grid unit and a preset reflective feature template; according to the distribution of the similarities, generating a spatial weight matrix, where the spatial weight matrix is used to identify the distribution probability of effective reflective regions in the reflective vest region image; multiplying the spatial weight matrix and the initial wearing state feature element by element to obtain a weighted intermediate feature; performing non-maximum suppression processing on the intermediate feature to remove redundant feature regions with an overlap rate higher than a preset threshold, and generating a target wearing state feature; where the pre-trained wearing state recognition model is obtained by fusing multi-modal training data, and the multi-modal training data includes reflective vest sample images under different lighting conditions; Determining the recognition result of the wearing state of the reflective vest of the human object to be detected according to the matching degree between the target wearing state feature and a preset wearing state threshold; where the recognition result includes normal wearing, not wearing, or partially blocked state.

2. The method according to claim 1, wherein The extracting the reflective vest region image corresponding to the human object to be detected from the real-time monitoring image includes: Performing human contour segmentation on the real-time monitoring image to obtain the human contour bounding box of the human object to be detected; Based on the human contour bounding box, locating the torso region of the human object to be detected, and performing color space conversion on the torso region to generate a first candidate region; Performing high-reflection region detection on the first candidate region to screen out a pixel set with a reflection intensity higher than a preset threshold; According to the connectivity distribution of the pixel set, determining the boundary coordinates of the reflective vest region image, and cropping the image region corresponding to the boundary coordinates from the real-time monitoring image.

3. The method according to claim 2, characterized in that, The performing multi-scale feature extraction on the reflective vest region image to generate an initial wearing state feature includes: Inputting the reflective vest region image into a multi-branch feature extraction network, where the multi-branch feature extraction network includes a first branch network, a second branch network, and a third branch network; where the first branch network is used to extract low-resolution texture features of the reflective vest region image, the second branch network is used to extract medium-resolution edge features, and the third branch network is used to extract high-resolution detail features; Performing cross-level feature fusion on the low-resolution texture features, medium-resolution edge features, and high-resolution detail features to generate a fused feature map; Performing channel attention weighting on the fused feature map, enhancing the feature channel weights related to the reflective material, and compressing the spatial dimension to generate the initial wearing state feature.

4. The method according to claim 3, wherein The training process of the pre-trained wearing state recognition model includes: Obtain multiple groups of training samples, where each group of training samples includes a reflective vest sample image, a corresponding wearing state label, and a lighting condition label; Perform adaptive lighting enhancement on the reflective vest sample image to generate an enhanced sample image; wherein, the adaptive lighting enhancement includes dynamically adjusting the brightness compensation parameter according to the lighting condition label; Input the enhanced sample image into the initial recognition model and output the predicted wearing state feature; Calculate the classification loss between the predicted wearing state feature and the wearing state label, and update the parameters of the initial recognition model based on gradient backpropagation until the classification loss converges; Perform end-to-end joint training on the converged initial recognition model and the multi-branch feature extraction network, fix the parameters of the multi-branch feature extraction network, and optimize the parameters of the region distribution analysis layer of the initial recognition model.

5. The method according to claim 4, wherein The adaptive lighting enhancement includes the following steps: Determine the lighting intensity level of the current sample according to the lighting condition label; If the lighting intensity level is low lighting condition, perform local contrast enhancement on the reflective vest sample image and superimpose random noise to simulate low lighting interference; If the lighting intensity level is high lighting condition, perform overexposed area suppression on the reflective vest sample image and reduce the global brightness to a preset range; If the lighting intensity level is dynamic change condition, perform multi-frame temporal fusion on the reflective vest sample image to generate an enhanced sample image with balanced lighting.

6. The method according to claim 1, wherein Determining the recognition result of the reflective vest wearing state of the human object to be detected according to the matching degree between the target wearing state feature and the preset wearing state threshold includes: Extract a first statistic related to the covered area of the reflective material, a second statistic related to the continuity of the reflective stripes, and a third statistic related to the matching degree of the human body posture in the target wearing state feature; Input the first statistic, the second statistic, and the third statistic into a preset multi-condition decision tree model to generate a comprehensive matching score; If the comprehensive matching score is higher than the first threshold, it is determined as the normal wearing state; If the comprehensive matching score is lower than the second threshold, it is determined as the not-worn state; If the comprehensive matching score is between the first threshold and the second threshold, it is determined as the partially occluded state; Among them, the construction process of the multi-condition decision tree model includes: Collect historical recognition data, where the historical recognition data includes multiple groups of the first statistic, the second statistic, the third statistic, and the corresponding true wearing state labels; Perform clustering analysis on the historical recognition data to determine the distribution boundaries of different wearing state categories in the three-dimensional statistic space; Generate multiple decision rules according to the distribution boundaries, and each decision rule corresponds to a hyperplane segmentation condition; Optimize the order arrangement of the decision rules based on the Gini coefficient, and construct hierarchical decision nodes until each leaf node only contains samples of a single wearing state category.

7. The method according to claim 1, wherein The method further includes a dynamic correction step for the recognition result: When the recognition result is the partially occluded state, obtain consecutive multi-frame historical monitoring images of the human object to be detected; Extract temporal features from the historical monitoring images to generate the dynamic change trend of the wearing state of the reflective vest; If the dynamic change trend indicates that the occluded area continues to expand, correct the recognition result to the non-wearing state; If the dynamic change trend indicates that the occluded area returns to the normal range within a preset time, correct the recognition result to the normal wearing state.

8. The method according to claim 7, characterized in that The extracting temporal features from the historical monitoring images to generate the dynamic change trend of the wearing state of the reflective vest includes: Perform key-point tracking on the reflective vest area in multiple consecutive frames of historical monitoring images to determine the motion trajectory of the occluded area; Calculate the area change rate and shape similarity of the occluded area between adjacent frames; Construct a time-series feature vector based on the motion trajectory, area change rate, and shape similarity; Input the time-series feature vector into a long short-term memory network to predict the evolution path of the occluded area in several future frames, and generate the dynamic change trend according to the evolution path.

9. A computer system, characterized in that, Including: One or more processors; A memory; One or more computer programs; Wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • High-altitude operation lifeline early warning method

    CN119068412A