System for consistent detection of human body and hat

By integrating a positioning module and a camera module into the smart safety helmet, and combining them with the AI ​​algorithm of the cloud management platform, a deep fusion of positioning and identity recognition is achieved. This solves the problem of the inability to detect the consistency between the person and the helmet in real time in existing technologies, and provides an efficient violation alarm and credential retention mechanism, which is suitable for safety management in multi-person operation scenarios.

CN121963258AActive Publication Date: 2026-05-01云筑信息科技(成都)有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
云筑信息科技(成都)有限公司
Filing Date
2026-04-02
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

The existing smart safety helmets' positioning function cannot verify the wearer's identity, leading to violations such as "the helmet is there but the person is not" or "someone else is wearing it on their behalf." The positioning data and identity recognition data are processed independently and cannot be linked in real time. The lack of a systematic design makes it impossible to achieve consistency judgment between the person and the helmet and to provide timely alarms.

Method used

The system uses a first safety helmet to broadcast positioning signals, a second safety helmet to capture facial images and generate a video stream, receives signals from a base station and uploads them to a cloud management platform for face recognition and consistency determination, and uses AI algorithms for face detection, quality assessment, feature extraction and comparison to achieve dual positioning verification and alarm linkage.

Benefits of technology

It achieves deep integration of positioning and identity recognition, determines the consistency between person and hat in real time, is suitable for multi-person operation scenarios, improves detection efficiency and accuracy, and provides third-party supervision and tamper-proof evidence retention of violations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963258A_ABST
    Figure CN121963258A_ABST
Patent Text Reader

Abstract

The invention discloses a system for human-helmet consistency detection. The system comprises a first helmet, a second helmet, a receiving base station and a cloud management platform. The first safety helmet is worn by a to-be-detected person and broadcasts a first positioning signal; and the second safety helmet is worn by a detector, broadcasts a second positioning signal, collects a face image of the to-be-detected person, generates a video stream and uploads the video stream. And the receiving base station receives the positioning signal and generates positioning data containing the unique identifier of the safety helmet, the signal strength value, the receiving time and the base station number. The cloud management platform receives the video stream and the positioning data, carries out face recognition and helmet consistency judgment through an AI algorithm, and triggers an alarm when it is judged that the helmet and the helmet are inconsistent or a stranger is recognized. According to the invention, deep linkage of positioning and identity recognition is realized, the consistency of people and caps can be accurately detected, and the method is suitable for high-risk operation scenes such as buildings, mines and chemical engineering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart wearables and engineering safety management technology, specifically to a system for detecting the consistency between a person and their hat. Background Technology

[0002] In high-risk work scenarios such as construction, mining, and chemical production, safety helmets are the core protective equipment for ensuring the safety of workers' heads, and wearing safety helmets correctly is a basic requirement for safe production. With the increasing demand for intelligent management, smart safety helmets with positioning functions have emerged, using technologies such as GPS, Bluetooth, or UWB to track the helmet's location, making it easier for managers to monitor the distribution of workers. Simultaneously, some work sites have deployed facial recognition systems to verify personnel identity and attendance status.

[0003] However, existing technologies have significant shortcomings: Traditional smart safety helmets can only obtain the helmet's location information, but cannot verify the wearer's identity. This leads to violations such as "helmet present but person not present" or "someone else wearing it on their behalf," and the location data is severely out of sync with the actual status of the person. Independently deployed facial recognition systems can only identify whether a person is present, but cannot link whether the person is wearing a safety helmet or whether the helmet is their own, making it difficult to achieve effective "person-helmet matching" verification.

[0004] Furthermore, the existing positioning and identity recognition systems operate independently, lacking a unified system architecture and data linkage mechanism. Positioning data and identity recognition data are stored and processed separately, making it impossible to perform correlation analysis at the same time dimension. This results in the inability to complete the real-time determination of the consistency between the person and their hat. When violations such as the separation of the person from their hat or wearing the hat on behalf of another person occur, the system fails to trigger alarms in a timely manner, and there is no on-site evidence retained, creating a security management loophole and hindering subsequent violation tracing and responsibility determination.

[0005] While some technical solutions attempt to combine positioning and identification functions, they suffer from problems such as high communication latency, rudimentary judgment logic, and lack of systematic design, making it difficult to meet the stringent requirements for real-time performance and accuracy in high-risk scenarios.

[0006] In summary, existing technologies urgently need a systematic solution that can achieve deep integration of positioning and identity recognition, real-time determination of the consistency between the person and the hat, and automatic alarm and credential retention functions. Summary of the Invention

[0007] This invention aims to solve the technical problems in the prior art, such as the disconnect between safety helmet positioning and identity recognition, the inability to detect the consistency between the person and the helmet in real time, the untimely warning of violations, and the lack of evidence retention, and provides a system for detecting the consistency between the person and the helmet.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A system for detecting human-hat consistency includes: One or more first safety helmets, each first safety helmet being worn by one person to be tested, including a first positioning module, the first positioning module being used to broadcast a first positioning signal containing the unique identifier of the safety helmet; One or more second safety helmets, each for one inspector, include a second positioning module and a camera module. The second positioning module is used to broadcast a second positioning signal containing the unique identifier of the safety helmet. The camera module is used to capture facial images of multiple inspectors wearing the first safety helmet and generate video stream data, which are then encoded, compressed, and uploaded as a compressed video stream. Multiple receiving base stations are deployed in a preset detection area. Each receiving base station has fixed base station location information and is used to receive the first positioning signal and the second positioning signal. For each received positioning signal, positioning data containing the unique identifier of the corresponding safety helmet, signal strength value, reception time and base station number are generated. The cloud management platform includes a video access module, a location data access module, an inference and analysis module, a judgment module, and an alarm linkage module; The video access module is used to receive compressed video streams via the WebRTC protocol, perform real-time decoding to obtain video frames, and then perform format conversion on the video frames to output standard image tensors. The location data access module is used to receive location data; The inference and analysis module is used to process standard image tensors using AI algorithms, including face detection, quality assessment, feature extraction and comparison, to obtain face comparison results; The determination module is used to determine the consistency between the face and the hat based on the face comparison results and the location data. The alarm linkage module is used to trigger an alarm when the hat and person are not consistent or when a stranger is identified.

[0009] Furthermore, the receiving base station supports the same communication protocol as the first and second safety helmets, and the communication protocol operates in the 2.4GHz frequency band. After receiving the first or second positioning signal, the receiving base station parses and obtains the unique identifier of the corresponding safety helmet, and records the signal reception time and the base station number of the receiving base station. The unique identifier of the safety helmet, the reception time, and the base station number are combined to form positioning data and uploaded to the cloud management platform.

[0010] Furthermore, in the alarm linkage module, when the determination result of the determination module is inconsistent between the person and the hat, an alarm command is generated and sent to the corresponding first and second safety helmets, and an alarm message is generated and pushed to the third-party business system; when the face comparison result of the reasoning and analysis module identifies a stranger, an alarm command is generated and sent to the corresponding second safety helmet, and the video frame at the corresponding time is obtained from the video access module according to the determination time, watermarked and stored as a scene photo, and an alarm message is generated and pushed to the third-party business system.

[0011] Furthermore, the video access module deploys a WebRTC signaling server and a media server, establishes a connection with the second safety helmet via the WebRTC protocol, completes SDP signaling interaction, ICE candidate address negotiation, and DTLS encryption verification, and then receives the compressed video stream uploaded by the second safety helmet through the RTP parser; the video access module calls the hardware accelerated decoding library to decode the H.264 / VP8 format compressed video stream in real time and outputs YUV420P format video frames; The video access module converts YUV420P format to RGB format through pixel format conversion, and then converts it into image data presented in the form of NumPy multidimensional array or OpenCV Mat matrix. Then, it scales the image data to match the input dimension of the inference analysis module through size normalization, adjusts the image data to the floating-point 0-1 range through data type conversion and numerical normalization, converts it to channel-first format through channel rearrangement, and finally adds batch dimension to generate standard image tensors.

[0012] Furthermore, the logic for determining the consistency between the person and the hat in the determination module includes: Obtain the face comparison result of the person to be detected from the inference analysis module; if the face comparison result is that a stranger is identified, then the person is determined to be a stranger; if the face comparison result is that the person is a registered person in the project, then obtain the person's identity information and the unique identifier of the corresponding first safety helmet, as well as the time of person identification when the person is identified. The location data of the second safety helmet at the moment of personnel identification is obtained from the location data access module, and the location of the detection personnel is determined based on the signal strength value of the second location signal received by multiple receiving base stations. Based on the person's identity information and the unique identifier of the corresponding first safety helmet, multiple location data of the first safety helmet within the time window at the moment of person identification are obtained from the location data access module; Based on the acquired multiple location data, it is determined whether there is at least one receiving base station near the moment of personnel identification that received the first location signal broadcast by the first safety helmet; If it does not exist, the result is determined to be inconsistent with the hat; If it exists, then based on the multiple positioning data of the first safety helmet and the base station location information of the receiving base station, the location of the first safety helmet at the time of personnel identification is calculated based on the signal strength value. The distance between the location of the personnel identification and the location of the detection personnel is calculated. If the distance is less than a preset threshold, the result is determined to be a match between the helmet and the person; otherwise, the result is determined to be a mismatch between the helmet and the person.

[0013] Furthermore, the reasoning and analysis module includes: A face detection model is used to detect face regions from standard image tensors and output face detection results, which include face bounding boxes, key points, and confidence scores. A face alignment model is used to perform affine transformations based on key points to normalize face images to a preset size. A face quality assessment model is used to evaluate the sharpness and lighting quality of normalized face images, filtering out frames with unacceptable sharpness and lighting. The feature extraction model is used to extract features from face images that have passed quality assessment and generate feature vectors. A face comparison model is used to calculate the similarity between feature vectors and a pre-stored face feature database.

[0014] Furthermore, the methods for evaluating the quality of normalized face images using the face quality assessment model include: The normalized face image is converted into a grayscale image; the Laplacian operator is applied to the grayscale image for convolution; the variance is calculated after taking the absolute value of the convolution result; if the variance is lower than the preset first threshold, the image is judged to be unqualified in terms of sharpness and is filtered out. Five indicators are calculated based on all pixels of the grayscale image: average brightness, proportion of dark pixels, proportion of bright pixels, standard deviation of brightness, and horizontal brightness balance of the left and right faces. The weighted scores of each indicator are combined to obtain a comprehensive lighting score. If the score is lower than the preset second threshold, it is judged as unqualified for lighting quality assessment and filtered out. Face images that pass both the sharpness assessment and the lighting quality assessment are defined as face images that pass the quality assessment.

[0015] Furthermore, the feature extraction model loads the ArcFace ONNX model, scales the qualified face images to 112×112 pixels and inputs them into the ArcFace model, performs forward inference through a deep convolutional neural network, and outputs a 512-dimensional feature vector from the fully connected layer.

[0016] Furthermore, the broadcast signal frequency of the first positioning module and the second positioning module is 5 times / second.

[0017] Furthermore, the first safety helmet also includes: The first microprocessor, electrically connected to the first positioning module, is used to control the broadcast frequency of the first positioning signal; The first alarm notification module is used to receive alarm commands from the cloud management platform and issue audible and visual alerts. The first power supply module provides power to the first positioning module and the first microprocessor; The second safety helmet also includes: The second microprocessor, electrically connected to the second positioning module and the camera module, is used to control the broadcast frequency of the second positioning signal and the image acquisition of the camera module, and to encode and compress the video stream data acquired by the camera module to generate a compressed video stream; The second alarm notification module is used to receive alarm commands from the cloud management platform and issue audible and visual alerts; The second power supply module provides power to the second positioning module, the camera module, and the first microprocessor.

[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention deeply integrates positioning and identity recognition functions through a system architecture consisting of a first safety helmet (worn by the person being tested), a second safety helmet (worn by the testing personnel), a receiving base station, and a cloud management platform. The first safety helmet broadcasts a positioning signal at a high frequency of 5 times per second, while the second safety helmet captures facial images of the person being tested in real time and generates a video stream for uploading. The receiving base station is deployed in the testing area, receiving signals from each safety helmet and generating positioning data containing a unique helmet identifier, signal strength value, reception time, and base station number. The cloud management platform uniformly receives the video stream and positioning data, and uses AI algorithms to perform facial recognition and helmet-person consistency determination. This architecture fundamentally solves the problem in existing technologies where positioning data and identity recognition data are independent and cannot be correlated for verification, realizing a complete data link from terminal acquisition and local transmission to cloud processing.

[0019] 2. The video access module establishes a connection with the second safety helmet via the WebRTC protocol. After completing signaling interaction and encryption verification, it receives the compressed video stream, calls the hardware-accelerated decoding library to decode and output YUV420P format video frames in real time, and generates standard image tensors through preprocessing such as pixel format conversion, size normalization, data type conversion, numerical normalization, and channel rearrangement, to achieve millisecond-level low-latency video transmission and adapt to network fluctuations in mobile scenarios.

[0020] 3. The inference and analysis module adopts a five-level cascaded architecture of face detection, alignment, quality assessment, feature extraction, and comparison. The face detection model improves detection speed through anchorless design and lightweight modifications; the alignment model achieves face normalization based on affine transformation of key points; the quality assessment model filters low-quality frames through multi-index scoring of sharpness and illumination; the feature extraction model loads ArcFace to generate 512-dimensional feature vectors; and the face comparison model performs identity matching based on similarity, maintaining high recognition accuracy in complex environments.

[0021] 4. The face quality assessment model employs a comprehensive scoring mechanism combining sharpness and illumination. Sharpness assessment calculates variance through convolution, filtering out blurry frames below a first threshold. Illumination quality assessment calculates five indicators based on grayscale images: average brightness, proportion of excessively dark pixels, proportion of excessively bright pixels, brightness standard deviation, and brightness balance between the left and right sides of the face. These indicators are weighted and fused to obtain a comprehensive illumination score, filtering out frames below a second threshold that fail to meet illumination standards. This ensures that only images of acceptable quality are included in feature extraction, significantly reducing the false alarm rate.

[0022] 5. Existing technology uses a self-portrait detection method, which places the camera on the safety helmet of the person being inspected and allows the person to identify themselves. This method is only suitable for single-person, independent work scenarios and cannot enable supervision by others. Furthermore, the location verification only verifies the position of the safety helmet and cannot rule out violations such as someone else wearing it on their behalf.

[0023] This invention employs a detection method using personnel photography. A camera is mounted on a second safety helmet worn by the inspector (manager), allowing the manager to actively capture facial images of the personnel being inspected, thus implementing a third-party proactive supervision mechanism. Each second safety helmet can inspect multiple personnel corresponding to the first safety helmets. In other words, a manager, wearing the second safety helmet, can sequentially or simultaneously perform helmet-person consistency checks on multiple personnel (workers), suitable for multi-person work scenarios and meeting the safety management needs of high-risk work environments. This design allows one manager to cover multiple workers, effectively improving detection efficiency, reducing labor costs, and enabling proactive supervision by the manager, overcoming the limitation of self-portrait-based detection methods that cannot achieve external supervision. Regarding location verification, this invention simultaneously verifies the distance relationship between the inspector's position and the inspected personnel's position, forming a dual location guarantee.

[0024] Based on the above innovations, the determination process of this invention is as follows: First, the manager's second safety helmet captures facial images of the personnel being inspected, and their identity is confirmed through facial recognition. If a stranger is identified, an alarm is triggered directly; if a registered employee of the project is identified, their identity information and the unique identifier of their corresponding first safety helmet are obtained.

[0025] Subsequently, dual positioning verification is performed: the location of the inspector is determined based on the positioning data of the second safety helmet; the positioning data of the first safety helmet near the time of personnel identification is obtained based on its unique identifier acquired after identification. If no base station signal is received, it is determined that the person and helmet are inconsistent; if a signal is present, the distance between the actual location of the first safety helmet and the location of the inspector is calculated. If the distance is less than a threshold, the person and helmet are determined to be consistent; otherwise, they are determined to be inconsistent.

[0026] This collaborative mechanism provides dual protection through a combination of self-photographed identity recognition and dual-location verification. It not only solves the problem that self-photographed detection cannot achieve supervision by others, but also effectively prevents violations such as the hat being on the person but not being there, or someone else wearing it on their behalf, by verifying the location distance between the inspector and the person being inspected. At the same time, it avoids misjudgments within the coverage area of ​​the base station, significantly improving the accuracy and reliability of the consistency determination of the person and hat.

[0027] 6. When the alarm linkage module determines that the helmet and the person are not the same, it sends alarm commands to the first and second helmets. When a stranger is identified, it sends an alarm command to the second helmet, which then issues an audible and visual alert. At the same time, it obtains the video frame at the corresponding moment based on the determination time, adds a watermark and saves it as a scene photo, and pushes alarm information to third-party business systems through the API interface, forming a closed-loop management system of "detection-alarm-save" to provide an tamper-proof and effective basis for tracing violations. Attached Figure Description

[0028] Figure 1 This is a block diagram of the system composition of the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0030] In the description of this invention, it should be noted that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0031] like Figure 1 As shown, the present invention provides a system for detecting the consistency of a person's helmet and hat, which consists of three parts: multiple safety helmets, a receiving base station deployed in the work area, and a cloud management platform.

[0032] Multiple safety helmets are included, comprising a first safety helmet and a second safety helmet. The first safety helmet, worn by the person being tested, includes a first positioning module for broadcasting a first positioning signal containing the helmet's unique identifier. The second safety helmet, worn by the testing personnel, includes a second positioning module and a camera module. The second positioning module broadcasts a second positioning signal containing the helmet's unique identifier, and the camera module captures facial images of multiple individuals wearing the first safety helmet and generates a video stream, which is then encoded, compressed, and uploaded as a compressed video stream. The positioning module of each smart safety helmet broadcasts the positioning signal at a preset frequency (e.g., 5 times / second).

[0033] Receiving base stations are deployed at a preset density within a predetermined detection area. Each base station has fixed location information, supports the 2.4GHz communication protocol, and receives positioning signals broadcast by all first and second safety helmets within its coverage area in real time. Upon receiving a positioning signal, the receiving base station parses and obtains the unique identifier of the safety helmet, records the reception time of the signal, and its own base station number. It combines the unique helmet identifier, signal strength value, reception time, and base station number to form positioning data, which is then uploaded to the cloud management platform via wired network or 4G / 5G communication. The receiving base station uses the same 2.4GHz communication protocol with both the first and second safety helmets. This protocol has optimized anti-interference mechanisms for industrial environments and, compared to standard protocols such as Bluetooth, WiFi, and Zigbee, exhibits stronger anti-interference capabilities and higher positioning accuracy in complex environments such as construction sites, mines, and chemical plants.

[0034] The cloud management platform communicates with all receiving base stations to receive the first positioning signal, the compressed video stream uploaded by the second safety helmet, and the positioning data uploaded by each base station. The cloud management platform uses AI algorithms for face recognition and helmet-person consistency determination, specifically including: decoding and preprocessing the compressed video stream through the video access module to output a standard image tensor; performing face detection, quality assessment, feature extraction, and comparison on the standard image tensor through the inference analysis module to obtain face comparison results; determining helmet-person consistency based on the face comparison results and positioning data through the determination module; and triggering an alarm through the alarm linkage module when a helmet-person inconsistency is determined or a stranger is identified, sending an alarm to the corresponding first or second safety helmet, while simultaneously saving on-site photos and pushing alarm information to third-party business systems.

[0035] The cloud management platform includes: The video access module is used to receive compressed video streams via the WebRTC protocol, perform real-time decoding to obtain video frames, and then perform format conversion on the video frames to output standard image tensors. The location data access module is used to receive location data; The inference and analysis module is used to process standard image tensors using AI algorithms, including face detection, quality assessment, feature extraction and comparison, to obtain face comparison results; The determination module is used to determine the consistency between the face comparison results and the location data, and output the determination result and determination time. The alarm linkage module is used to trigger an alarm when the hat and person are not consistent or when a stranger is identified.

[0036] In one specific implementation, within the alarm linkage module, when the determination result shows an inconsistency between the person and the helmet, an alarm command is generated and sent to the corresponding first and second safety helmets to warn the violating personnel and inspection personnel. Simultaneously, an alarm message is generated and pushed to a third-party business system, and an alarm notification is sent to a pre-set management personnel's mobile phone via SMS. When the face comparison result from the inference analysis module identifies a stranger, an alarm command is generated and sent to the corresponding second safety helmet. Based on the determination time, the corresponding video frame is retrieved from the video access module, watermarked, and saved as a scene photo. An alarm message is then pushed to the third-party business system via an API interface. The third-party business system includes a management system associated with a computer or mobile phone to achieve remote monitoring and recording.

[0037] In one specific implementation, the video access module deploys a WebRTC signaling server and a media server. It establishes a connection with the smart helmet via the WebRTC protocol, completes SDP (Session Description Protocol) signaling interaction, ICE (Interactive Connection Establishment) candidate address negotiation, and DTLS (Datagram Transport Layer Security) encryption verification, and then receives the compressed video stream uploaded by the smart helmet through an RTP (Real-Time Transport Protocol) parser. The video access module calls hardware-accelerated decoding libraries (such as OpenH264 and libvpx) to decode the H.264 / VP8 format compressed video stream in real time and outputs YUV420P format video frames.

[0038] YUV420P format video frames are continuous one-dimensional byte streams in the data layer, stored consecutively in the order of Y, U, and V channels. The pixel resolution of the Y channel is "H×W" (H is the image height and W is the image width), and the U and V channels are "H / 2×W / 2". The overall byte stream length is H×W×1.5 bytes, without dimension markings or pixel coordinate mapping.

[0039] In a more specific implementation, the video access module further leverages the memory-sharing mechanism between OpenCV and NumPy to map the structured data of the OpenCV Mat matrix into a NumPy multidimensional array. Specifically, the two-dimensional matrix structure of OpenCV cv::Mat ("H rows × W columns × 3 channels") is reconstructed into a three-dimensional array of NumPy [H, W, 3], with the dimension order strictly matching the input requirements of the inference and analysis module. The data buffer of the NumPy array directly points to the RGB pixel data block pointed to by the data pointer of cv::Mat, and the two share the same memory space, achieving zero-copy cross-frame compatibility. The NumPy array inherits the data type uint8 from cv::Mat, and the channel order remains R→G→B, with the row step and total number of bytes completely consistent with cv::Mat. At the same time, a multidimensional array access interface is encapsulated, supporting direct access to the single channel value of a single pixel through an index (such as arr[y, x, c], where y is the row, x is the column, and c is the channel).

[0040] After completing the above data format conversion, the video access module performs a size normalization operation on the image data. This operation is based on the input size set by the face detection model in the inference analysis module, such as 640×640 pixels required by the face detection model. Specifically, the scaling factors scale_w = input_w / img_w and scale_h = input_h / img_h are calculated (input_w and input_h are the input width and input height set by the face detection model, / img_w and / img_h are the actual width and actual height of the original input image, and scale_w and scale_h are the width and height scaling factors, respectively). The minimum scaling factor is taken to ensure that the image is not stretched. The original image is scaled proportionally, and after scaling, grayscale values ​​are filled in the blank areas of the image (pad_value=114, pad_value is the filling pixel value), generating a square image with a size of input_h×input_w, with the data format kept as uint8 and the shape as [input_h, input_w, 3].

[0041] Next, the video access module performs data type conversion and numerical normalization operations. The image data type is converted from uint8 (the original image data type) to float32 (the converted data type) to reduce the loss of network computational precision. Simultaneously, pixel value normalization is performed on the float32 image data, mapping pixel values ​​from 0-255 to the 0-1 range. The calculation formula is pixel_norm = pixel_raw / 255.0 (where pixel_norm and pixel_raw are the normalized pixel value and the original pixel value, respectively).

[0042] Subsequently, the video access module performs a channel rearrangement operation, converting the image data from the [H,W, 3] dimension order of [height, width, channel] to the [3, H, W] channel-first format of [channel, height, width] in the AI ​​framework standard, adapting to the channel-first storage rules of deep learning frameworks such as PyTorch / TensorFlow.

[0043] Finally, the video access module performs a tensor encapsulation operation, adding a batch dimension to the image data and encapsulating it into a model inference tensor. The batch size is set to 1 (batch_size is the batch size, i.e., the number of images processed in one inference operation; it is set to 1 for inference on a single image). The final output standard inference tensor format is [1, 3, input_h, input_w]. Taking a 640×640 input as an example, the output tensor shape is [1, 3, 640, 640], dtype is float32, and the pixel value range is [0,1]. This tensor can be directly input into the face detection model of the inference analysis module for forward inference.

[0044] To address the dynamic differences in frame rate and resolution between video streams from different sources, the video access module integrates the aforementioned adaptive preprocessing mechanism. Through intelligent frame sampling and dynamic size normalization algorithms, it ensures that the data format input to the core algorithm model is highly consistent, and the tensor size strictly matches the input dimension of the face detection model, eliminating the need for secondary format conversion and improving inference efficiency.

[0045] In one specific implementation, the reasoning analysis module includes: A face detection model is used to detect face regions from standard image tensors and output face detection results, which include face bounding boxes, key points, and confidence scores. A face alignment model is used to perform affine transformations based on key points to normalize face images to a preset size. A face quality assessment model is used to evaluate the sharpness and lighting quality of normalized face images, filtering out frames with unacceptable sharpness and lighting. The feature extraction model is used to extract features from face images that have passed quality assessment and generate feature vectors. A face comparison model is used to calculate the similarity between feature vectors and a pre-stored face feature database.

[0046] In a more specific implementation, the face detection model adopts the YOLOv8-Face architecture. The input is a standard image tensor output by the video access module, such as a floating-point tensor with shape=[1,3,640,640] and pixel values ​​ranging from [0,1]. As a face-specific detection model optimized based on the YOLOv8 general detection architecture, YOLOv8-Face inherits the anchor-free core design of YOLOv8, abandoning the traditional anchor-box predefinition and matching process. It directly regresses the offset and size of the face bounding box through the detection head. At the same time, it optimizes the backbone network downsampling, neck feature fusion, and head prediction layer output for face detection scenarios. The network is divided into three parts: the Backbone feature extraction layer, the Neck multi-scale feature fusion layer, and the Head detection prediction layer.

[0047] The backbone layer uses the C2f-YOLOv8-Face module, which is a lightweight modification of the C2f module of YOLOv8, reducing the number of branches to adapt to low-texture facial features. The input tensor [1,3,640,640] is passed through a convolutional layer with a 6×6 kernel, stride of 2, and padding of 2 to complete channel upscaling and initial downsampling, and output feature map [1,64,320,320]. Subsequently, three sets of customized C2f modules are sequentially passed through the face detection layers. Each set includes 1×1 convolutional channel compression, 3×3 convolutional feature extraction, and residual connections. Downsampling is then performed using a 3×3 convolution with a stride of 2, ultimately outputting three core scale feature maps: F1 [1,128,160,160] corresponds to a 4x downsampling for large-scale face detection; F2 [1,256,80,80] corresponds to an 8x downsampling for medium-scale face detection; and F3 [1,512,40,40] corresponds to a 16x downsampling for small-scale face detection. The final C2f-Face module with a stride of 2 is then applied to F3, outputting F4 [1,1024,20,20] for ultra-small-scale face detection. To address the simple features and low texture of faces, the backbone layer reduces the number of convolutional branches to decrease computation. Simultaneously, the SiLU activation function is used instead of ReLU to solve the gradient vanishing problem under low-texture features.

[0048] The Neck layer employs an improved PAN-FPN structure to address the scale variability in face detection, enabling top-down high-level semantic feature transfer and bottom-up low-level detail feature aggregation. First, F4 is compressed to 512 channels via a 1×1 convolution, then upsampled to [1,512,40,40] using nearest neighbor upsampling. This is concatenated with F3 and fused using the C2f-Face module to obtain F3'. Similarly, F3' is upsampled and fused with F2 to obtain F2'; F2' is upsampled and fused with F1 to obtain F1'. Then, bottom-up fusion is performed: F1' is downsampled via a 2-stride convolution and concatenated with F2' to obtain F2''; F2'' is downsampled and concatenated with F3' to obtain F3''; F3'' is downsampled to obtain F4''. The final output consists of fused feature maps at three scales: P1[1,128,160,160], P2[1,256,80,80], and P3[1,512,40,40]. The neck layer changes the upsampling method from transposed convolution to nearest-neighbor upsampling to improve inference speed, and adds a channel attention module after concatenation and fusion to adaptively strengthen the feature weights of the face region.

[0049] The Head layer is a customized anchorless detection layer. P1 / P2 / P3 are respectively convolved with 1×1 to map the channel dimension. Each pixel of each feature map predicts 4 bounding box regression parameters and 1 confidence parameter, and outputs prediction tensors T1[1,5,160,160], T2[1,5,80,80], and T3[1,5,40,40]. The dimensions of the three prediction tensors are rearranged from "batch, channel, height, width" [1, 5, H, W] to "batch, number of feature points, prediction parameters", where the number of feature points = H × W. This results in T1': [1, 160 × 160, 5] = [1, 25600, 5][1, 25600, 5], T2': [1, 80 × 80, 5] = [1, 6400, 5], and T3': [1, 40 × 40, 5] = [1, 1600, 5]. After concatenating the feature point count dimension, the global face prediction tensor T is obtained: shape = [1, 25600 + 6400 + 1600, 5] = [1, 33600, 5]. This achieves the transformation from feature representation to face detection result prediction. The confidence parameter is mapped to the [0, 1] interval using the Sigmoid activation function, and the bounding box regression parameters are directly output. The detection head uses single-class prediction to reduce prediction parameters by more than 40%, and the anchor-free design adapts to arbitrary changes in face size. Depth-separable convolution further reduces the amount of computation.

[0050] The global face prediction tensor undergoes post-processing to obtain the final face detection result. Post-processing includes: prediction tensor parsing and coordinate restoration. The five prediction parameters [dx, dy, dw, dh, conf] (representing the x-offset, y-offset, width, height, and face confidence of the face bounding box, respectively) for each feature point in the global face prediction tensor are combined with the feature point grid center and downsampling step size to calculate the absolute coordinates of the face bounding box in the 640×640 image. The coordinates are restored to the original image using scaling factors and padding values ​​applied during preprocessing. For the 33,600 face bounding boxes in the candidate set array, low-confidence boxes are filtered based on a confidence threshold (default 0.5, adjustable to 0.3 for smart helmet scenarios) to reduce subsequent computation. A customized NMS-Face algorithm (IoU threshold 0.6) is used to remove duplicate boxes, and the face bounding boxes are cropped and their sizes fine-tuned. The final output is the accurate face detection result of the original image, in NumPy array format, which includes the coordinates of the face bounding box (x1, y1, x2, y2) (x1 and y1 represent the x-axis coordinates and y-axis coordinates of the upper left corner of the face bounding box in the original image, respectively, and x2 and y2 represent the x-axis coordinates and y-axis coordinates of the lower right corner of the face bounding box in the original image, respectively), the coordinates of 5 key points (such as left eye, right eye, nose tip, left corner of mouth, and right corner of mouth) and confidence score.

[0051] In a more specific implementation, the face alignment model is implemented through the "face_align" module. It receives the face bounding box and key points output by the face detection model, performs affine / similarity transformations on the detected face region, and normalizes the face image to the standard size expected by the ArcFace feature extraction model. Specifically, an affine transformation matrix is ​​calculated based on the correspondence between the coordinates of five key points and the positions of the target standard face key points. This matrix is ​​applied to transform and scale the face region image to 112×112 pixels. Bilinear interpolation is used during the transformation process to ensure image quality, and the normalized standard face image is output.

[0052] In more specific implementations, the face quality assessment model uses methods to assess the quality of normalized face images, including sharpness assessment and illumination quality assessment.

[0053] The specific steps for implementing sharpness assessment are as follows: First, the input normalized face image is converted into a grayscale image. The conversion formula is based on the ITU-R BT.601 standard and fully considers the differences in human eye sensitivity to different color channels. Specifically: f_gray(m,n) = 0.299×R(m,n) + 0.587 ×G(m,n) + 0.114×B(m,n) Where R(m,n), G(m,n), and B(m,n) represent the red, green, and blue channel pixel values ​​of the face image at coordinates (m,n), respectively, and f_gray(m,n) is the converted grayscale value.

[0054] Then, the Laplacian operator is applied to the grayscale image for convolution to extract the high-frequency components. The Laplacian operator uses a 3×3 convolution kernel, which can highlight the edges and details in the image. The absolute value of the convolution result is taken to eliminate the directional influence, resulting in the absolute value of the Laplacian response. The mean of the absolute value of the Laplacian response is calculated, and then the variance of the absolute value of the Laplacian response is calculated from the mean. This variance reflects the richness of the high-frequency components of the image; the larger the variance, the sharper the image edges and the richer the details, i.e., the clearer the image.

[0055] Finally, a sharpness determination is performed: if the variance is greater than or equal to a preset first threshold T, the frame is determined to be sharp and retained; if the variance is less than the first threshold T, the frame is determined to be blurry and discarded. The first threshold T is an empirical value that can be adaptively adjusted according to the specific lighting conditions and motion characteristics of the smart helmet shooting scene.

[0056] Lighting quality assessment is based on all pixels of the grayscale image. Let the pixel value of the face grayscale image be I(x,y), the image height be H1, the width be W1, the total number of pixels be N1 = H1 × W1, and the grayscale value range be [0, 255]. The assessment includes the following six quantitative indicators: mean brightness, dark ratio, bright ratio, brightness standard deviation, and lateral balance of brightness between the left and right sides of the face.

[0057] The average brightness metric is used to determine whether the lighting is too dim. Average brightness reflects the overall brightness level of the face area; the smaller the value, the dimmer the lighting. Its physical meaning lies in quantifying the overall illumination intensity, making it a fundamental indicator of whether the lighting is sufficient. Specifically, average brightness is calculated using the `np.mean(gray_face)` function from the NumPy library, which directly calculates the arithmetic mean of all pixels in the grayscale image.

[0058] The percentage of excessively dark pixels is used to help determine whether the lighting is too dark. The percentage of excessively dark pixels refers to the percentage of pixels with a brightness value less than 50 out of the total pixels. A higher percentage indicates darker lighting and more dark areas. The calculation first uses OpenCV's cv2.calcHist function to calculate a grayscale histogram hist[0...255], where hist[k] represents the number of pixels with a grayscale value of k. Then, the histogram is normalized so that the sum of the percentages of all intervals is 1. Specifically, the percentage of excessively dark pixels is calculated using np.sum(hist[:50]). 100 is the sum of the normalized proportions of grayscale values ​​0-49, converted to a percentage.

[0059] The overexposed pixel percentage metric is used to help determine if an image is overexposed. The overexposed pixel percentage refers to the percentage of pixels with a brightness value greater than 200 out of the total pixels. A higher percentage indicates more overexposed areas, indirectly reflecting an imbalance in lighting. The overexposed pixel percentage is specifically calculated using `np.sum(hist[200:])`. 100 is the percentage obtained by summing the normalized proportions of grayscale values ​​between 200 and 255.

[0060] The luminance standard deviation is a core metric used to determine whether the lighting is balanced. It reflects the dispersion of pixel brightness in the face region; a smaller value indicates a more concentrated brightness distribution and more uniform lighting, while a larger value indicates greater differences in brightness and more unbalanced lighting (such as localized shadows or strong light). Specifically, the luminance standard deviation is calculated using the `np.std(gray_face)` function from the NumPy library, which directly calculates the overall standard deviation of all pixels in the grayscale image without any degree of freedom correction.

[0061] The horizontal brightness balance index for the left and right faces is used to determine whether the lighting is balanced. The horizontal brightness balance index refers to the percentage difference in brightness between the left and right halves of the face. A smaller value indicates that the brightness of the left and right faces is closer and the lighting is more balanced; a larger value indicates a greater difference in brightness between the left and right faces (such as unilateral lighting or unilateral shadow), and a more unbalanced lighting. In specific calculations, the grayscale image of the face is divided into two halves along the vertical midline. When the image width W1 is even, W / 2 = W1 / 2 = W1 / 2, and the number of pixels in the left and right halves is equal; when the image width W1 is odd, W / 2 = W1 / 2 + 1, and the right half has one more column of pixels than the left half. Next, the average brightness of the left half (meanleft) and the average brightness of the right half (meanright) are calculated. The horizontal brightness of the left and right faces is specifically calculated using abs(meanleft - meanright) / max(meanleft, meanright). 100, calculate the percentage difference in brightness between the left and right sides of the face.

[0062] Finally, the weighted scores of the above five indicators (approximately 100 points in total) are combined to obtain a comprehensive lighting score. A higher score indicates more uniform lighting, more moderate brightness, and no excessive darkness or overexposure. Weights are assigned to different indicators, with higher weights indicating a greater impact on lighting quality. Negative corrections are applied to the percentage of excessive darkness, excessive brightness, and brightness standard deviation; that is, the larger the percentage or standard deviation, the more points are deducted. The weighting of each indicator is as follows: average brightness 0.4, excessive darkness percentage 0.15, excessive brightness percentage 0.15, brightness standard deviation 0.15, and horizontal brightness balance of the left and right sides of the face 0.15.

[0063] The weighted score is calculated as follows: score = 0.4 × (mean_brightness / 255 × 100) + 0.15 × (100 - dark_ratio) + 0.15 × (100 - bright_ratio) + 0.15 × (100 - min(brightness_std, 100)) + 0.15 × (100 - min(lateral_balance, 100)).

[0064] When calculating the weighted score, the average brightness item normalizes the brightness values ​​in the range of 0-255 to 0-100 points to reflect the basic brightness level; the proportion of overly dark pixels and the proportion of overly bright pixels are calculated by subtracting the proportion from 100. The higher the proportion, the lower the score for this item, thus penalizing overly dark and overexposed pixels; the brightness standard deviation and the horizontal brightness balance of the left and right faces are limited to the maximum deduction value, taking min(brightness_std,100) and min(lateral_balance, 100) to avoid the standard deviation and brightness balance being too large, which would cause the score for this item to be negative.

[0065] In some embodiments of this invention, the inference analysis module further includes a face tracking model. This model is used to perform cross-frame association and identification management of faces detected in a continuous video frame sequence composed of standard image tensors output by the video access module. Specifically, the face tracking model is implemented through the "face_track" module. Combined with the deep-sort-realtime algorithm model, it receives face detection results from different video frames output by the face detection model and performs ID preservation and trajectory updates for the detected faces. The face tracking model compresses the frequency of "detection-feature extraction" to only when necessary, triggering feature extraction only for consistently occurring faces that meet quality standards. This significantly reduces overall computational power consumption and improves temporal stability and the reliability of the face comparison algorithm. The face tracking model ultimately outputs face detection results with tracking labels, including face bounding boxes, key points, confidence scores, and tracking labels.

[0066] In some embodiments of the present invention, the feature extraction model loads the ArcFace ONNX model, scales the qualified face image to 112×112 pixels and inputs it into the ArcFace model, performs forward inference through a deep convolutional neural network, and outputs a 512-dimensional feature vector from the fully connected layer.

[0067] In one specific implementation, the face comparison model is implemented through the "faceapi" module. It calculates similarity based on feature vectors and matches them against a pre-stored face feature database. The model retrieves the most similar feature vector from the MongoDB face vector feature database. This database pre-stores records of registered face feature vectors (feature database vectors) and corresponding identity information for all personnel registered in the project. Each record includes fields such as personnel ID, name, position, and a unique identifier for the associated first safety helmet. Cosine similarity is used to calculate the similarity between the current feature vector and the feature database vector, with the formula: similarity = cos(θ) = (A·B) / (||A||×||B||). Since the feature vectors are normalized, the similarity is simplified to a vector dot product. A dynamic threshold strategy is set, which can be adaptively adjusted based on the quality score output by the face quality assessment model. If the highest similarity exceeds the dynamic threshold, a successful match is determined, and the personnel's identity information is output; if all similarities are below the threshold, the person is determined to be a stranger.

[0068] In one specific implementation, the logic for determining the consistency between the person and the hat in the determination module includes: First, the face comparison results of the person to be detected are obtained from the inference analysis module. Specifically, the inference analysis module performs cascaded processing on the standard image tensor output by the video access module: the face detection model detects face regions in the image and outputs face bounding boxes and key points; the face alignment model performs affine transformation based on the key points to normalize the face image to a preset size; the face quality assessment model evaluates the sharpness and lighting quality of the normalized face image and filters low-quality frames; the feature extraction model extracts features from the face images that pass the quality assessment and generates a 512-dimensional feature vector; the face comparison model calculates the similarity between the feature vector and the pre-stored face feature library. If the highest similarity exceeds the dynamic threshold, the match is successful, and the identity information of the corresponding person and the unique identifier of the first safety helmet associated with them are output. At the same time, the time when the recognition result is generated is recorded as the time of person recognition; if the highest similarity is lower than the dynamic threshold, it is determined to be a stranger.

[0069] Secondly, if the face comparison result is a registered person in the project (or a currently registered person in the project), the location data of the second safety helmet at the time of personnel identification is obtained from the location data access module. Based on the signal strength values ​​of the second location signal received by multiple receiving base stations, the location of the detected personnel is calculated and determined by the location algorithm.

[0070] Next, based on the unique identifier of the first safety helmet corresponding to the person, multiple location data of the first safety helmet within the time window of the person identification moment are obtained from the location data access module. The time window can be set according to the actual scenario, such as 0.5 seconds before and after the person identification moment.

[0071] Then, based on the acquired location data, it is determined whether at least one receiving base station received the first location signal broadcast by the first safety helmet near the time of personnel identification. If no base station received the signal, the result is directly determined to be inconsistent between the person and the helmet.

[0072] If at least one base station receives the signal, then based on the multiple positioning data of the first safety helmet and the base station location information of each receiving base station, the location of the first safety helmet at the moment of personnel identification is calculated using a positioning algorithm based on the signal strength value, and then the distance between the location and the location of the detection personnel is calculated.

[0073] Finally, the calculated distance is compared with a preset threshold: if the distance is less than the preset threshold, the result is determined to be a match between the person and the hat; if the distance is greater than or equal to the preset threshold, the result is determined to be a mismatch between the person and the hat. This preset threshold can be set according to the safety requirements of the actual working scenario, for example, set to 1 meter or 2 meters.

[0074] In summary, the logic for determining the consistency of the person and hat in this embodiment can be summarized as follows: First, the identity of the person to be tested is confirmed through facial recognition. If it is a stranger, an alarm is triggered directly. If it is a registered person, dual positioning verification is performed—first, the location of the person to be tested is determined, then the positioning signal of the safety helmet corresponding to the person is obtained and its actual location is calculated. Finally, the distance between the two is calculated and compared with a preset threshold. If the distance is less than the threshold, it is determined that the person and hat are consistent; otherwise, it is determined that the person and hat are inconsistent.

[0075] This logic provides dual protection through both identity verification and location authentication: First, it is crucial to prevent situations where the helmet is present but the person is not. If the person to be tested is not present, facial recognition will fail; if the helmet is placed far away, the distance between the helmet and the person being tested will be too great to pass distance verification. Only when the person is present and the helmet is nearby can it be determined that the person and helmet match.

[0076] Secondly, it prevents others from wearing the helmet on behalf of others. When person B wears person A's helmet, the system captures person B's facial image. If person B is a stranger, an alarm is triggered directly; if person B is a registered person, the system queries person B's own helmet location signal, but person B is actually wearing person A's helmet. The two do not match, resulting in the helmet position not matching the detected person's position, and thus failing the location distance verification.

[0077] Therefore, only when the person being tested is present at the scene and their helmet is nearby can it be determined that the person and helmet match, thus effectively preventing violations such as the helmet being present but the person not being there, or someone else wearing it on their behalf.

[0078] The judgment module outputs the final judgment result and judgment time to the alarm linkage module for subsequent alarm triggering and credential retention.

[0079] In some embodiments of the present invention, the first safety helmet and the second safety helmet are specifically implemented as follows: The first and second positioning modules employ low-power radio frequency chips, supporting the 2.4G communication protocol, to broadcast a signal containing the unique identifier of the safety helmet at a preset frequency. This identifier is an unchangeable unique ID written to each safety helmet at the factory, used to uniquely identify the safety helmet in the system. In a preferred embodiment, the first and second positioning modules broadcast the signal at a frequency of 5 times per second. This high-frequency broadcast design ensures the real-time nature of the positioning data, enabling the receiving base station to receive the signal 5 times per second, providing sufficient data sampling points for subsequent real-time positioning and helmet-person consistency determination.

[0080] The second safety helmet's camera module uses a wide-angle, high-definition camera with a resolution of at least 1080P. It is positioned at the front brim of the helmet to capture real-time facial images of the person being inspected and generate a video stream. The camera module's field of view is at least 120° to ensure clear capture of the person's facial features. The first safety helmet may not be equipped with a camera module, or may be equipped with one but used only for other purposes.

[0081] The first safety helmet also includes a first microprocessor electrically connected to the first positioning module, used to control the broadcast frequency of the first positioning signal. The second safety helmet also includes a second microprocessor electrically connected to the second positioning module and the camera module, used to control the broadcast frequency of the second positioning signal and the image acquisition timing of the camera module, and simultaneously encode and compress the raw video stream data acquired by the camera module to generate a compressed video stream. In a preferred embodiment, the first and second microprocessors use embedded ARM architecture chips, connected to the corresponding positioning module and camera module through interfaces such as SPI and MIPI, to perform H.264 hardware encoding and compression processing on the video stream data, reducing data transmission bandwidth requirements.

[0082] The first safety helmet also includes a first alarm module, used to receive alarm commands from the cloud management platform and issue audible and visual alerts. The alarm command is generated and sent when the determination result shows that the helmet and person do not match. The second safety helmet also includes a second alarm module, used to receive alarm commands from the cloud management platform and issue audible and visual alerts. The alarm command is generated and sent when a stranger is detected. In a preferred embodiment, the alarm module includes an LED indicator and a buzzer, electrically connected to a corresponding microprocessor. After receiving the alarm command, the microprocessor controls the LED indicator to flash and drives the buzzer to sound, continuing until the wearer manually resets it, ensuring that the wearer and management personnel can promptly detect abnormalities.

[0083] The first safety helmet also includes a first power module to power the first positioning module and the first microprocessor. The second safety helmet also includes a second power module to power the second positioning module, the camera module, and the second microprocessor. In a preferred embodiment, the power module uses a rechargeable lithium battery, and a power management chip provides stable power to each module to ensure normal operation. The power module capacity is designed to meet the requirement of a minimum 10-hour battery life for the entire device, adapting to the long-term use needs of high-risk work scenarios.

[0084] In a preferred embodiment, the modules of the first and second safety helmets are integrated using in-mold injection molding to ensure that the waterproof and dustproof rating reaches the IP67 standard, meeting the usage requirements of high-risk work scenarios.

[0085] Finally, it should be noted that the above embodiments are merely preferred embodiments of the present invention used to illustrate the technical solutions of the present invention, and are not intended to limit the invention, nor are they intended to limit the patent scope of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. That is to say, any changes or refinements made to the main design concept and spirit of the present invention that are not of substantial significance, but whose technical problems are still consistent with the present invention, should be included within the protection scope of the present invention. In addition, the direct or indirect application of the technical solutions of the present invention to other related technical fields are similarly included within the patent protection scope of the present invention.

Claims

1. A system for detecting the consistency of a person's hat, characterized in that, include: One or more first safety helmets, each first safety helmet being worn by one person to be tested, including a first positioning module, the first positioning module being used to broadcast a first positioning signal containing the unique identifier of the safety helmet; One or more second safety helmets, each for one inspector, include a second positioning module and a camera module. The second positioning module is used to broadcast a second positioning signal containing the unique identifier of the safety helmet. The camera module is used to capture facial images of multiple inspectors wearing the first safety helmet and generate video stream data, which are then encoded, compressed, and uploaded as a compressed video stream. Multiple receiving base stations are deployed in a preset detection area. Each receiving base station has fixed base station location information and is used to receive the first positioning signal and the second positioning signal. For each received positioning signal, positioning data containing the unique identifier of the corresponding safety helmet, signal strength value, reception time and base station number are generated. The cloud management platform includes a video access module, a location data access module, an inference and analysis module, a judgment module, and an alarm linkage module; The video access module is used to receive compressed video streams via the WebRTC protocol, perform real-time decoding to obtain video frames, and then perform format conversion on the video frames to output standard image tensors. The location data access module is used to receive location data; The inference and analysis module is used to process standard image tensors using AI algorithms, including face detection, quality assessment, feature extraction and comparison, to obtain face comparison results; The determination module is used to determine the consistency between the face and the hat based on the face comparison results and the location data. The alarm linkage module is used to trigger an alarm when the hat and person are not consistent or when a stranger is identified.

2. The system according to claim 1, characterized in that, The receiving base station supports the same communication protocol as the first and second safety helmets, and the communication protocol operates in the 2.4GHz frequency band. After receiving the first or second positioning signal, the receiving base station parses and obtains the unique identifier of the corresponding safety helmet, and records the signal reception time and the base station number of the receiving base station. The unique identifier of the safety helmet, the reception time, and the base station number are combined to form positioning data and uploaded to the cloud management platform.

3. The system according to claim 1, characterized in that, In the alarm linkage module, when the determination result of the determination module is inconsistent with the helmet, an alarm command is generated and sent to the corresponding first and second safety helmets, and an alarm message is generated and pushed to the third-party business system. When the face comparison result of the reasoning analysis module identifies a stranger, an alarm command is generated and sent to the corresponding second safety helmet. Based on the judgment time, the video frame at the corresponding moment is obtained from the video access module, watermarked and saved as a scene photo, and an alarm message is generated and pushed to the third-party business system.

4. The system according to claim 1, characterized in that, The video access module deploys a WebRTC signaling server and a media server. It establishes a connection with the second safety helmet through the WebRTC protocol, completes SDP signaling interaction, ICE candidate address negotiation, and DTLS encryption verification, and then receives the compressed video stream uploaded by the second safety helmet through the RTP parser. The video access module calls the hardware acceleration decoding library to decode the H.264 / VP8 format compressed video stream in real time and output YUV420P format video frames; The video access module converts YUV420P format to RGB format through pixel format conversion, and then converts it into image data presented in the form of NumPy multidimensional array or OpenCV Mat matrix; Then, the image data is scaled to match the input dimension of the inference analysis module through size normalization. The image data is adjusted to the floating-point 0-1 range through data type conversion and numerical normalization. Then, it is converted to the channel-first format through channel rearrangement. Finally, batch dimensions are added to generate a standard image tensor.

5. The system according to claim 1, characterized in that, The logic for determining the consistency of the person's hat in the determination module includes: Obtain the face comparison result of the person to be detected from the inference analysis module; if the face comparison result is that a stranger is identified, then the person is determined to be a stranger; if the face comparison result is that the person is a registered person in the project, then obtain the person's identity information and the unique identifier of the corresponding first safety helmet, as well as the time of person identification when the person is identified. The location data of the second safety helmet at the moment of personnel identification is obtained from the location data access module, and the location of the detection personnel is determined based on the signal strength value of the second location signal received by multiple receiving base stations. Based on the person's identity information and the unique identifier of the corresponding first safety helmet, multiple location data of the first safety helmet within the time window at the moment of person identification are obtained from the location data access module; Based on the acquired multiple location data, it is determined whether there is at least one receiving base station near the moment of personnel identification that received the first location signal broadcast by the first safety helmet; If it does not exist, the result is determined to be inconsistent with the hat; If it exists, then based on the multiple positioning data of the first safety helmet and the base station location information of the receiving base station, the location of the first safety helmet at the time of personnel identification is calculated based on the signal strength value. The distance between the location of the personnel identification and the location of the detection personnel is calculated. If the distance is less than a preset threshold, the result is determined to be a match between the helmet and the person; otherwise, the result is determined to be a mismatch between the helmet and the person.

6. The system according to claim 5, characterized in that, The reasoning and analysis module includes: A face detection model is used to detect face regions from standard image tensors and output face detection results, which include face bounding boxes, key points, and confidence scores. A face alignment model is used to perform affine transformations based on key points to normalize face images to a preset size. A face quality assessment model is used to evaluate the sharpness and lighting quality of normalized face images, filtering out frames with unacceptable sharpness and lighting. The feature extraction model is used to extract features from face images that have passed quality assessment and generate feature vectors. A face comparison model is used to calculate the similarity between feature vectors and a pre-stored face feature database.

7. The system according to claim 6, characterized in that, Methods for face quality assessment models to evaluate the quality of normalized face images include: The normalized face image is converted into a grayscale image; the Laplacian operator is applied to the grayscale image for convolution; the variance is calculated after taking the absolute value of the convolution result; if the variance is lower than the preset first threshold, the image is judged to be unqualified in terms of sharpness and is filtered out. Five indicators are calculated based on all pixels of the grayscale image: average brightness, proportion of dark pixels, proportion of bright pixels, standard deviation of brightness, and horizontal brightness balance of the left and right faces. The weighted scores of each indicator are combined to obtain a comprehensive lighting score. If the score is lower than the preset second threshold, it is judged as unqualified for lighting quality assessment and filtered out. Face images that pass both the sharpness assessment and the lighting quality assessment are defined as face images that pass the quality assessment.

8. The system according to claim 6, characterized in that, The feature extraction model loads the ArcFace ONNX model. After scaling the qualified face images to 112×112 pixels, they are input into the ArcFace model. After forward inference through a deep convolutional neural network, a 512-dimensional feature vector is output from the fully connected layer.

9. The system according to claim 1, characterized in that, The broadcast signal frequency of the first positioning module and the second positioning module is 5 times / second.

10. The system according to claim 1, characterized in that, The first safety helmet also includes: The first microprocessor, electrically connected to the first positioning module, is used to control the broadcast frequency of the first positioning signal; The first alarm notification module is used to receive alarm commands from the cloud management platform and issue audible and visual alerts. The first power supply module provides power to the first positioning module and the first microprocessor; The second safety helmet also includes: The second microprocessor, electrically connected to the second positioning module and the camera module, is used to control the broadcast frequency of the second positioning signal and the image acquisition of the camera module, and to encode and compress the video stream data acquired by the camera module to generate a compressed video stream; The second alarm notification module is used to receive alarm commands from the cloud management platform and issue audible and visual alerts; The second power supply module provides power to the second positioning module, the camera module, and the first microprocessor.

Citation Information

Patent Citations

  • Safety helmet, safety system and personnel management method

    CN109938439A

  • Human-helmet matching method, system and equipment based on intelligent safety helmet

    CN121188493A