Construction site safety monitoring method, system and device based on intelligent safety helmet and medium
Patent Information
- Application Number
- CN202610301777.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-03-12
AI Technical Summary
然而,现有技术在实际应用中基于单一感知维度的方法,如果外来人员或无资质人员未佩戴定位卡直接潜入危险区域,系统将无法通过定位信号察觉;同时,由于动态巡检视角下的视频画面与三维空间坐标缺乏有效映射,系统难以将画面中的实体目标与实际的身份资质进行精准匹配,对于复杂工地环境下的无卡人员非法入侵与低资质人员违规跨区域作业,存在漏报与盲区管控风险
[0058]通过构建集成人员身份识别、作业区域空间定位、视频智能分析以及安全预警处理于一体的工地安全监测技术体系,将智能安全帽作为前端感知终端,对施工现场人员信息、空间位置及现场作业情况进行协同采集与综合分析,实现人员身份信息、作业资质信息与现场作业环境数据之间的联动管理,从而形成覆盖人员准入核验、现场行为监测以及异常情况预警的全过程安全监管机制;同时利用电子地图与空间定位信息对施工区域进行数字化建模,使监测结果能够与具体作业区域进行精准关联,提高施工现场管理的可视化程度与信息化水平,并通过自动化预警与数据记录机制减少人工巡检工作量,提升工地安全管理的智能化程度与管理效率。
Smart Images

Figure CN122200766B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart construction site safety management technology, and in particular to a construction site safety monitoring method, system, equipment and medium based on a smart safety helmet. Background Technology
[0002] With the rapid development of digital transformation and smart construction sites in the construction industry, the scale and complexity of construction sites continue to rise, and the mobility of workers and cross-regional collaborative operations are constantly increasing. How to efficiently monitor, intelligently identify, and manage the real-time location, identity, qualifications, and on-site safety status of construction personnel has become a core requirement for ensuring construction safety and improving construction site management efficiency.
[0003] In existing technologies, safety monitoring and intrusion warnings for construction site personnel are typically achieved through monitoring systems based on a single sensing method. For example, this might rely on on-site video surveillance cameras for personnel identification and counting, or simply on location signal cards worn by workers for area headcount and unauthorized access alarms. However, in practical applications, existing technologies based on a single sensing dimension are ineffective if unauthorized personnel or those without proper identification cards enter dangerous areas without being detected by the system's location signals. Furthermore, because the video footage from a dynamic inspection perspective lacks effective mapping with three-dimensional spatial coordinates, the system struggles to accurately match physical targets in the footage with their actual identities and qualifications. This leads to risks of missed detections and blind spots in complex construction site environments, particularly regarding unauthorized intrusions by unauthorized personnel and unauthorized cross-regional operations by unqualified individuals. Summary of the Invention
[0004] In view of this, this application provides a construction site safety monitoring method, system, equipment and medium based on a smart safety helmet to solve the above problems.
[0005] Firstly, a construction site safety monitoring method based on a smart safety helmet is provided, the method comprising:
[0006] The system acquires facial feature data of the wearer collected by the smart safety helmet, compares the facial feature data with a preset qualification database, and extracts the corresponding qualification level of the wearer from the qualification database.
[0007] The system can acquire video stream data and spatial coordinate data of the target work area collected by the smart safety helmet in real time, and also acquire signal data of personnel positioning cards located within the target work area.
[0008] The target video frame is obtained by extracting frames from the video stream data using a preset sampling frequency.
[0009] The spatial coordinate data is matched with the preset electronic map to determine the virtual work area where the smart safety helmet is currently located, and the pre-configured access level requirements of the virtual work area are obtained from the electronic map.
[0010] Obtain the camera imaging parameters of the smart safety helmet, and calculate the field of view coverage corresponding to the target video frame based on the camera imaging parameters and spatial coordinate data;
[0011] The number of personnel positioning card signal data within the field of view coverage is counted to obtain the first number of on-duty personnel, and the number of personnel entity targets in the target video frame is counted using a preset deep neural network model to obtain the second number of on-duty personnel.
[0012] When the number of employees on duty differs from the number of employees on duty, or when the qualification level is lower than the access level requirement, a safety hazard warning instruction is generated and output.
[0013] The above technical solution automatically identifies the wearer and obtains the corresponding qualification level by comparing the facial features collected by the smart safety helmet with the qualification database. At the same time, it combines video stream data, spatial coordinate data, and personnel positioning card signal data to perform dual statistics on the number of personnel in the target work area and compares the statistical results with the access level requirements of the work area. This enables real-time monitoring of the identity, number, and work qualifications of personnel in the construction area. When the number of personnel is abnormal or the personnel qualifications do not meet the requirements, a safety hazard warning is generated in a timely manner, improving the timeliness and accuracy of safety supervision at the construction site.
[0014] Optionally, the target video frame is obtained by performing frame extraction on the video stream data using a preset sampling frequency, specifically including:
[0015] Establish a buffer queue with a preset capacity, and configure a first independent thread and a second independent thread for data interaction based on the buffer queue;
[0016] The first independent thread acquires video frames from the video stream data according to the preset connection strategy and determines whether the buffer queue has reached the preset full load state.
[0017] If the buffer queue is full, remove the historical video frame at the head of the buffer queue and store the currently acquired video frame at the tail of the buffer queue; otherwise, store the currently acquired video frame directly into the buffer queue.
[0018] The target video frame is obtained by retrieving video frames from the buffer queue using a second independent thread based on the sampling frequency.
[0019] The above technical solution achieves asynchronous reading and processing of video stream data by establishing a buffer queue with a fixed capacity and setting up two independent threads to perform video frame acquisition and frame extraction processing respectively. When the buffer queue is full, the oldest video frame is automatically removed to store the latest video frame, which can ensure continuous updating of video data. At the same time, video frames are extracted from the queue according to a preset sampling frequency, thereby reducing the amount of video processing computation while ensuring real-time performance and improving system processing efficiency and stability.
[0020] Optionally, video frames from the video stream data are acquired via a first independent thread according to a preset connection strategy, specifically including:
[0021] Initiate a connection request for the video stream data and obtain the connection status feedback.
[0022] When the connection status is successful, extract video frames from the video stream data according to the normal acquisition branch in the connection strategy.
[0023] When the feedback connection status is a connection failure status, the video stream data is reconnected according to the abnormal reconnection branch in the connection strategy, using the preset exponential backoff strategy.
[0024] When the number of reconnection operations exceeds a preset threshold, the video capture resources currently occupied by the first independent thread are released and a new connection request is initiated so that video frames can be extracted after the connection is restored to a successful state.
[0025] The above technical solution sets a connection strategy in the first independent thread to detect the video stream connection status. When the connection is successful, video frames are extracted normally. When the connection fails, an exponential backoff strategy is used to reconnect. When the number of reconnections exceeds a threshold, video capture resources are released and a new connection request is initiated. This allows the video stream acquisition to be automatically restored when there is a network anomaly or the device connection is interrupted, ensuring the continuity of video data acquisition and the stability of system operation.
[0026] Optionally, the camera imaging parameters of the smart helmet are obtained, and the field of view coverage corresponding to the target video frame is calculated based on the camera imaging parameters and spatial coordinate data, specifically including:
[0027] Acquire real-time attitude data of the smart safety helmet and extract the horizontal and vertical field of view parameters from the camera imaging parameters;
[0028] A three-dimensional projection cone is generated by using spatial coordinate data as vertices and the direction vector determined by real-time attitude data as the central axis, combined with horizontal and vertical field of view parameters.
[0029] Calculate the intersecting polygonal region between the 3D projection cone and the plane of the preset electronic map, and determine the intersecting polygonal region as the field of view coverage area corresponding to the target video frame.
[0030] The above technical solution, by acquiring the camera imaging parameters and real-time attitude data of the smart safety helmet, and combining them with spatial coordinate data to construct a three-dimensional projection cone, and then calculating the intersection area of the cone with the electronic map plane, can accurately determine the actual field of view coverage corresponding to the current video frame, thereby clarifying the spatial corresponding area of the video image in the construction site and improving the spatial positioning accuracy of on-site personnel statistics and safety monitoring.
[0031] Optionally, a pre-defined deep neural network model is used to detect human entities in the target video frame, and the number of human entities is counted to obtain a second number of on-duty personnel. Specifically, this includes:
[0032] The target video frame is input into a deep neural network model for feature extraction to obtain the spatial feature vector of the person.
[0033] Perform target classification and bounding box regression on the spatial feature vectors of personnel to generate candidate personnel prediction boxes and corresponding classification confidence scores;
[0034] Filter candidate prediction boxes with a classification confidence level greater than a preset confidence level threshold, count the number of candidate prediction boxes after filtering, and obtain the number of entities in the initial screening.
[0035] Based on the detection results of historical video frames, the number of multiple initial screening entities generated consecutively within a preset time window is obtained. Consistency verification is performed on the number of initial screening entities to generate a stable number of entities, and the stable number of entities is determined as the second number of on-duty personnel.
[0036] The above technical solution utilizes a deep neural network model to detect people in the target video frame and filters the predicted bounding boxes of the detected candidates according to their confidence levels. At the same time, it combines the detection results of historical video frames within a preset time window for consistency verification. This reduces misjudgments caused by instantaneous detection errors or occlusions and improves the stability and accuracy of the statistical results of the number of people.
[0037] Optionally, the target video frames are input into a deep neural network model for feature extraction to obtain spatial feature vectors of people, specifically including:
[0038] The target video frame input to the deep neural network model is divided into grid regions of a preset size, and the initial spatial feature matrix of each grid region is extracted by performing a preset convolution operation.
[0039] Each initial spatial feature matrix is input into the spatial attention layer of the deep neural network model to perform feature weight redistribution operations, generating an enhanced spatial feature matrix corresponding to each grid region.
[0040] The spatial arrangement sequence of each grid region in the target video frame is obtained, and the enhanced spatial feature matrices are flattened and stitched according to the spatial arrangement sequence to obtain the spatial feature vector of the personnel.
[0041] The above technical solution divides the target video frame into multiple grid regions and extracts convolutional features. Then, it uses a spatial attention layer to redistribute the feature weights of each grid region, enabling the model to enhance spatial features related to the person target. The enhanced features are flattened and stitched together according to the spatial arrangement sequence to generate a unified spatial feature vector of the person, thereby improving the model's ability to express the features of the person target and enhancing the accuracy of person detection.
[0042] Optionally, when the number of employees on duty at the first post differs from the number of employees on duty at the second post, or when the qualification level is lower than the entry level requirement, a safety hazard warning instruction will be generated and output, specifically including:
[0043] Calculate the difference between the first number of people on duty and the second number of people on duty. When the difference is not zero, extract the target video frame and spatial coordinate data as on-site evidence data and generate the first sub-early warning instruction containing the illegal intrusion identifier and the on-site evidence data.
[0044] Obtain the first quantitative value corresponding to the qualification level and the second quantitative value corresponding to the access level requirement. When the first quantitative value is less than the second quantitative value, obtain the device identifier of the smart safety helmet, and associate the facial feature data, qualification level and device identifier to obtain personnel identity verification data, and generate a second sub-warning instruction containing the unauthorized operation identifier and personnel identity verification data.
[0045] The generated first and / or second sub-early warning instructions are encapsulated to generate a safety hazard warning instruction, which is then output to a preset management terminal.
[0046] The above technical solution determines the differences in the number of personnel and their qualification levels. When the number of personnel is inconsistent, it generates an illegal intrusion warning message containing on-site video frames and spatial coordinate data. When the personnel's qualifications are below the access level requirements, it generates an unauthorized operation warning message containing facial feature data, qualification level, and equipment identification. The relevant data is then packaged and sent to the management terminal, thereby providing managers with data-driven warning information containing on-site evidence, which facilitates the rapid verification and handling of safety hazards at the construction site.
[0047] Secondly, a construction site safety monitoring system based on a smart safety helmet is provided, the system comprising:
[0048] The qualification verification module is configured to acquire the facial feature data of the wearer collected by the smart safety helmet, compare the facial feature data with the preset qualification database, and extract the corresponding qualification level of the wearer from the qualification database.
[0049] The data acquisition module is configured to acquire video stream data and spatial coordinate data of the target work area collected by the smart safety helmet in real time, and to acquire signal data of personnel positioning cards located in the target work area.
[0050] The video frame extraction module is configured to extract target video frames from video stream data using a preset sampling frequency.
[0051] The area matching module is configured to match spatial coordinate data with a preset electronic map to determine the virtual work area where the smart safety helmet is currently located, and to obtain the pre-configured access level requirements of the virtual work area from the electronic map.
[0052] The field of view calculation module is configured to acquire camera imaging parameters for the smart safety helmet and calculates the field of view coverage corresponding to the target video frame based on the camera imaging parameters and spatial coordinate data.
[0053] The dual-mode counting module is configured to count the number of personnel positioning card signal data within the field of view coverage to obtain the first number of on-duty personnel, and to use a preset deep neural network model to detect personnel entity targets in the target video frame and count the number of personnel entity targets to obtain the second number of on-duty personnel.
[0054] The safety early warning module is configured to generate and output a safety hazard warning command when the number of employees on duty differs from the number of employees on duty, or when the qualification level is lower than the access level requirement.
[0055] Thirdly, an electronic device is provided, including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the above.
[0056] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing instructions that, when executed, perform the method as described in any of the preceding claims.
[0057] In summary, implementing one or more technical solutions provided in this application has at least the following technical effects or advantages:
[0058] By constructing a construction site safety monitoring technology system that integrates personnel identification, work area spatial positioning, intelligent video analysis, and safety early warning processing, and using smart safety helmets as front-end sensing terminals, the system collaboratively collects and comprehensively analyzes personnel information, spatial location, and on-site work conditions. This enables the coordinated management of personnel identity information, work qualification information, and on-site work environment data, thereby forming a comprehensive safety supervision mechanism covering personnel access verification, on-site behavior monitoring, and abnormal situation early warning. Simultaneously, by utilizing electronic maps and spatial positioning information to digitally model the construction area, the monitoring results can be accurately correlated with specific work areas, improving the visualization and informatization level of construction site management. Furthermore, automated early warning and data recording mechanisms reduce the workload of manual inspections, enhancing the intelligence and efficiency of construction site safety management. Attached Figure Description
[0059] Figure 1 This is an exemplary system architecture diagram of a construction site safety monitoring method or a construction site safety monitoring system based on a smart safety helmet, which applies the present application.
[0060] Figure 2 This is a flowchart illustrating a construction site safety monitoring method based on a smart safety helmet disclosed in this application;
[0061] Figure 3 This is a schematic diagram of a construction site safety monitoring system based on a smart safety helmet disclosed in this application;
[0062] Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in this application.
[0063] Explanation of reference numerals in the attached diagram: 100, System architecture; 101, First terminal device; 102, Second terminal device; 103, Third terminal device; 104, Network; 105, Server; 301, Qualification verification module; 302, Data acquisition module; 303, Video frame extraction module; 304, Region matching module; 305, Field of view calculation module; 306, Dual-mode counting module; 307, Security early warning module; 401, Processor; 402, Communication bus; 403, User interface; 404, Network interface; 405, Memory. Detailed Implementation
[0064] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0065] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0066] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0067] Figure 1 This paper illustrates an exemplary system architecture diagram of an embodiment of a construction site safety monitoring method or a construction site safety monitoring system based on a smart safety helmet, to which this application can be applied.
[0068] like Figure 1 As shown, the system architecture 100 may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium to provide communication links between the terminal devices 101, 102, 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0069] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as model training applications, video recognition applications, web browser applications, social platform software, etc.
[0070] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptops, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.
[0071] When terminals 101, 102, and 103 are hardware devices, video capture devices can also be installed on them. These video capture devices can be various devices capable of capturing video, such as cameras, sensors, etc. Users can use the video capture devices on terminals 101, 102, and 103 to capture video.
[0072] Server 105 can be a server that provides various services, such as a backend server for processing data displayed on terminal devices 101, 102, and 103. The backend server can analyze and process the received data and can feed back the processing results (such as recognition results) to the terminal devices.
[0073] It should be noted that a server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services), or as a single software program or software module. No specific limitations are made here.
[0074] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. In particular, if the target data does not need to be obtained remotely, the above system architecture may exclude the network and include only terminal devices or servers.
[0075] Figure 2This is a flowchart illustrating a construction site safety monitoring method based on a smart safety helmet, as described in this application. This method can be implemented using a computer program, a microcontroller, or run on a construction site safety monitoring system based on a smart safety helmet. The computer program can be integrated into the application or run as a standalone utility application. The specific steps of the construction site safety monitoring method based on a smart safety helmet are described in detail below.
[0076] S201: Acquire the facial feature data of the wearer collected by the smart safety helmet, compare the facial feature data with the preset qualification database, and extract the qualification level corresponding to the wearer from the qualification database.
[0077] In this embodiment of the application, the preset qualification database refers to a pre-built and maintained structured data set used to record the identity of all legal on-site personnel and their corresponding work safety permissions. For example, the database may store the facial feature vector of a worker and qualification level labels such as "Special Electrician Level 1 Qualification" or "Ordinary Worker Without Special Qualification" that are bound to his / her identity.
[0078] Specifically, a visual sensor integrated into the front of the smart safety helmet collects facial images upon initial wear or periodic wake-up. Addressing the complex lighting and occlusion environments of construction sites, the collected facial images undergo image quality enhancement and region masking analysis. The system assesses the light intensity and blurriness of the facial image. In backlit or low-light environments, adaptive histogram equalization or the Retinex algorithm is used for illumination compensation and dehazing to restore facial texture details. Simultaneously, region masking detection is added for other accessories such as masks, dust masks, and polarized goggles. When large areas of facial occlusion are detected, the feature extraction weights are adaptively adjusted, focusing the comparison on high-frequency features in exposed areas such as the eye area and brow bone. Based on the enhanced image, liveness detection and keypoint localization are performed to extract unique facial features of the wearer. The system collects facial feature data from a structured database; it converts the acquired facial feature data into a multi-dimensional feature vector, and compares it with multiple sets of legal personnel facial feature templates pre-stored in a pre-set qualification database to calculate the vector similarity between the two; when the calculated similarity score is greater than or equal to the set security pass threshold, it determines that the current wearer's identity is legal and the comparison is passed, and then, based on the matched identity index information, it directly reads and extracts the wearer's corresponding qualification level from the pre-set qualification database, thereby completing the accurate conversion and extraction from personnel physical and biological characteristics to security business permission characteristics.
[0079] S202: Real-time acquisition of video stream data and spatial coordinate data of the target work area collected by the smart safety helmet, and acquisition of personnel positioning card signal data located in the target work area.
[0080] In this embodiment of the application, the personnel positioning card signal data refers to digital carrier information continuously broadcast by the wireless radio frequency beacon worn by the construction site workers, which contains a unique identification and signal strength characteristics. For example, the signal data can be a data packet with a worker ID number and a Received Signal Strength Indication (RSSI) parameter sent based on UWB (Ultra Wide Band) technology or Bluetooth Low Energy technology.
[0081] Specifically, by activating the image acquisition component configured at the front of the smart safety helmet and the built-in satellite positioning module (such as the RTK high-precision positioning module), the system continuously captures and records the dynamic physical scene in front of the wearer's field of vision at a preset frame rate during construction, forming a continuous video stream of the target work area. Simultaneously, it records spatial coordinate data representing the current absolute geographical location or relative three-dimensional coordinate system position of the smart safety helmet. At the same time, through the wireless communication gateway configured at the work site or the radio frequency receiving antenna in the smart safety helmet, it scans and listens to the wireless radio frequency broadcasts in the target work area. After signal demodulation at the underlying physical layer and protocol parsing at the link layer, it obtains the personnel positioning card signal data carried in the radiated electromagnetic waves, thereby completing the synchronous aggregation of multi-source perception data such as visual images of the work site environment, the spatial orientation of the smart safety helmet itself, and the distribution characteristics of surrounding personnel.
[0082] S203: Obtain the target video frame by performing frame extraction processing on the video stream data using a preset sampling frequency.
[0083] For example, this step aims to address the resource matching challenge between high-frequency visual acquisition streams and limited AI analysis computing power in edge computing scenarios at construction sites. Considering that the front-end camera of a smart safety helmet typically outputs continuous images at a high frame rate (e.g., 25 or 30 frames per second), and that there is significant visual environmental redundancy between adjacent frames within a short period, requiring the system to perform frame-by-frame deep learning detection on the entire video stream would not only result in ineffective use of underlying computing resources but also inevitably lead to severe memory backlog and delayed warnings. Therefore, the system introduces a dynamic temporal downsampling mechanism, using a reasonable rhythm (e.g., extracting 3 to 5 frames per second) adapted to the actual inference speed of the downstream neural network to "freeze" dynamic slices of the work site. This strategy is equivalent to extracting the most representative video frames from the continuous video data stream as needed, significantly reducing the concurrent computing pressure on the system hardware and ensuring from the source that every image subsequently fed into the AI model for security assessment possesses high real-time performance and timely prevention capabilities.
[0084] In one possible implementation, the target video frame is obtained by extracting frames from the video stream data using a preset sampling frequency. Specifically, this includes: establishing a buffer queue with a preset capacity, and configuring a first independent thread and a second independent thread for data interaction based on the buffer queue; the first independent thread acquiring video frames from the video stream data according to a preset connection strategy, and determining whether the buffer queue has reached a preset full-load state; if the buffer queue is full, removing the historical video frame at the head of the buffer queue and storing the currently acquired video frame at the tail of the buffer queue; otherwise, directly storing the currently acquired video frame in the buffer queue; and the second independent thread extracting video frames from the buffer queue based on the sampling frequency to obtain the target video frame.
[0085] In this embodiment of the application, a buffer queue refers to a first-in-first-out data storage structure with a fixed physical length allocated in memory space. It is used to represent an asynchronous decoupling relay station built between the high-frequency acquisition end (producer) and the low-frequency algorithm processing end (consumer) of the video stream. For example, the buffer queue can be a circular buffer with a capacity limit of 30 frames, which is specifically used to temporarily store the live scene data that has just been captured by the camera but has not yet been extracted and analyzed by the subsequent model.
[0086] Specifically, memory blocks are allocated in the runtime environment to establish a buffer queue with a preset capacity. Under the task scheduling mechanism of the underlying operating system, a first independent thread and a second independent thread are configured for data interaction based on the buffer queue, thereby achieving concurrent execution of data retrieval and data analysis without blocking each other. The first independent thread continuously decodes and acquires real-time video frames from the video stream data at a preset decoding rate. Before each write operation, the capacity monitoring interface is called to determine whether the current occupancy rate of the buffer queue has reached the preset full load state (i.e., the number of stored frames equals the preset quantity). If the buffer queue is detected to be full, in order to ensure the absolute real-time performance of the security monitoring screen at the business level and avoid computing power bottlenecks... The backlog of video frames caused by bottlenecks will trigger an active preservation mechanism. This mechanism removes the oldest historical video frame from the head of the buffer queue using pointer operations to free up a storage slot, and stores the latest video frame at the tail of the buffer queue. Otherwise, if the full load threshold is not reached, the currently acquired video frame is directly stored in the buffer queue. At the same time, a second independent thread responsible for downstream AI inference calculations periodically extracts currently resident video frames from the buffer queue on demand based on its own set sampling frequency, which is usually lower than the acquisition frequency. It removes redundant data that has not been sampled, so that through this asynchronous frame extraction process, the target video frame that meets the algorithm's computing power requirements and maintains extremely high real-time performance is finally obtained.
[0087] In one possible implementation, a first independent thread acquires video frames from the video stream data according to a preset connection strategy. Specifically, this includes: initiating a connection request for the video stream data and obtaining the feedback connection status; when the feedback connection status is successful, extracting video frames from the video stream data according to the normal acquisition branch of the connection strategy; when the feedback connection status is failed, performing a reconnection operation for the video stream data using a preset exponential backoff strategy according to the abnormal reconnection branch of the connection strategy; and when the number of reconnection operations exceeds a preset threshold, releasing the video capture resources currently occupied by the first independent thread and re-initiating the connection request to extract video frames after the connection status is restored to successful.
[0088] In this embodiment, the preset exponential backoff strategy refers to a retry algorithm model in which the waiting interval for the control system to re-attempt to establish a connection increases exponentially when a network communication failure or link disconnection occurs. It is used to represent a dynamic backoff protection mechanism set to avoid network bandwidth congestion or exhaustion of streaming media server resources due to high-frequency blind reconnection. For example, the strategy can be set to wait 1 second before the first retry, 2 seconds after the second failure, 4 seconds after the third failure, and so on, until the waiting time reaches the set maximum limit.
[0089] Specifically, the first independent thread, responsible for fetching underlying streaming media data, initiates a connection request for video stream data (such as RTSP / RTMP streams) to the remote smart helmet or edge gateway. After the underlying network protocol handshake interaction is completed, it listens for and obtains the feedback connection status of the connection request returned by the socket. When the feedback connection status at the underlying layer indicates a successful connection after the handshake is established, the control program execution flow enters and follows the normal acquisition branch in the preset connection strategy, continuously parsing the payload in the network data packets and stably extracting video frames from the video stream data through the decoding process. Conversely, if the feedback connection status is detected as a network timeout, connection refusal, or heartbeat packet loss, the exception handling logic is immediately triggered, and the exception reconnection branch in the preset connection strategy is followed to suspend unintentional connections. For high-frequency requests, the aforementioned preset exponential backoff strategy is adopted. As the number of consecutive failures increases, the sleep interval is gradually extended, thereby rhythmically and non-blockingly executing reconnection operations for video stream data. Furthermore, when the number of reconnection operations recorded by the counter exceeds a preset threshold (e.g., 10 consecutive attempts still fail), it is determined that the current underlying communication handle may be "fake dead" or resource deadlocked. At this time, in order to prevent system memory leaks and port exhaustion, the video capture resources currently occupied by the first independent thread are forcibly cleared and completely released (including destroying the underlying decoder instance and disconnecting the network connection handle). After completing the resource cleanup, the connection request is re-initiated from scratch. Through this self-healing method of completely resetting the link, video frames can continue to be reliably extracted after the network environment recovers and the feedback is that the connection is successful.
[0090] S204: Match the spatial coordinate data with the preset electronic map to determine the virtual work area where the smart safety helmet is currently located, and obtain the pre-configured access level requirements of the virtual work area from the electronic map.
[0091] In this embodiment of the application, the virtual work area refers to a logical block with specific management attributes that is divided into closed geometric boundaries (such as a set of polygon vertex coordinates) in a digital two-dimensional or three-dimensional space. It is used to represent the electronic mapping of different physical construction zones in the construction site in the digital twin system, which are divided according to the degree of danger of the actual environment, the construction stage, or the confidentiality level of personnel. For example, the virtual work area can be a "high-risk excavation area for deep foundation pits", "tower crane lifting fall radius area" or "no-go zone for live work" delineated in the Geographic Information System (GIS) or BIM (Building Information Modeling).
[0092] Specifically, by analyzing the spatial coordinate data containing three-dimensional orientation information such as longitude, latitude, and elevation received in real time by the underlying positioning module, and considering the multipath effect and signal drift problems that exist at construction sites, especially near deep foundation pits or tall steel structures, a time-series coordinate sliding window is first established before position matching. This is combined with the nine-axis inertial measurement unit (IMU) built into the smart safety helmet. The acceleration and angular velocity data output by the IMU (Integrated Measurement Unit) are used to smooth the motion trajectory and remove outliers from the raw spatial coordinate data output by the underlying positioning module using an extended Kalman filter algorithm. During this filtering process, the raw spatial coordinate data (longitude, latitude, and elevation) are used as observation variables, and the position and velocity of the smart helmet are used as state variables. State updates and covariance predictions are performed by combining the acceleration and angular velocity output by the IMU. After obtaining the filtered and stable spatial coordinate data, this data is used as a dynamic anchor point and projected onto a unified coordinate system plane containing a pre-built and rendered electronic map based on the construction site aerial survey model for geometric topology and position matching. During the spatial position mapping process, the underlying spatial computing engine calls algorithms such as ray casting or vector cross product to determine the inclusion relationship between points and polygons, calculating in real time whether the spatial coordinate point completely falls within one or more defined electronic fences. Within the boundary, spatial locking is triggered only when the filtered stable coordinates fall within the virtual work area for multiple consecutive sampling periods. This allows for the precise and dynamic determination of the physical location of the worker wearing the smart safety helmet, mapped to a specific virtual work area in the digital space, without human intervention. After spatial locking, the unique block identifier corresponding to the area is extracted, and a relational query command is used to retrieve the security configuration field mounted on the preset electronic map attribute tree. This accurately retrieves the pre-configured access level requirements of the triggered virtual work area from the electronic map (e.g., the access control policy corresponding to the area is "only personnel holding advanced special operation qualifications are allowed access"). This completely transforms the geographical location of the physical space into the permission threshold of the logical space, providing a rigid judgment benchmark for subsequent personnel identity verification and hazard interception.
[0093] S205: Obtain the camera imaging parameters of the smart safety helmet, and calculate the field of view coverage corresponding to the target video frame based on the camera imaging parameters and spatial coordinate data.
[0094] For example, this step aims to establish a precise mapping bridge between the mobile device's first-person dynamic visual image and the objective physical geographic space. Considering that the wearer's head gaze direction is constantly changing during construction and operation, the system abandons the traditional static blind spot estimation. Instead, it utilizes the camera's inherent optical angular characteristics and the current absolute positioning anchor point to dynamically reverse-engineer and delineate the actual physical surface outline covered by the current video frame on a two-dimensional digital base map. This is equivalent to projecting a virtual light spot boundary representing the "currently visible area" onto the smart helmet in real time within a global electronic fence, thereby establishing a mapping relationship between the relative visual coordinate system and the absolute geographic coordinate system. This ensures that the system can accurately pinpoint which specific construction site area is currently under the monitoring line of sight, providing a rigid geographic constraint for subsequent precise delineation and inventory of positioning card beacons within the field of view.
[0095] In one possible implementation, the camera imaging parameters of the smart safety helmet are obtained, and the field of view coverage corresponding to the target video frame is calculated based on the camera imaging parameters and spatial coordinate data. Specifically, this includes: obtaining the real-time attitude data of the smart safety helmet and extracting the horizontal and vertical field of view parameters from the camera imaging parameters; generating a three-dimensional projection cone with the spatial coordinate data as the vertex and the direction vector determined by the real-time attitude data as the central axis, combined with the horizontal and vertical field of view parameters; calculating the intersecting polygonal region between the three-dimensional projection cone and the plane of the preset electronic map, and determining the intersecting polygonal region as the field of view coverage corresponding to the target video frame.
[0096] In this embodiment of the application, the three-dimensional projection cone refers to a virtual three-dimensional geometric space model that extends outward along the main optical axis of the lens in a three-dimensional physical coordinate system, with the physical spatial location of the image acquisition device as the vertex, and is constrained by the inherent optical angular properties of the lens. It is used to represent the absolute boundary of the real-world three-dimensional visible area that the camera can capture in the current specific physical pose state. For example, the three-dimensional projection cone can be a three-dimensional perception area in the shape of a square pyramid with the center of the camera lens on the smart safety helmet worn by the worker as the fixed point, projecting forward and downward in a wide angle like a flashlight beam.
[0097] Specifically, the system reads data in real time from attitude sensing components such as the nine-axis IMU or high-precision gyroscope built into the smart helmet via the underlying communication bus. This data is used to calculate the helmet's yaw, pitch, and roll angles in the spatial coordinate system, obtaining the helmet's real-time attitude data. The system then calls the underlying hardware device description file to extract the horizontal and vertical field-of-view parameters (e.g., the lens's horizontal FOV) from the pre-calibrated camera imaging parameters. (View, field of view) is 120 degrees, vertical FOV is 90 degrees). In the constructed three-dimensional digital twin coordinate system, the spatial coordinate data representing the absolute position of the smart safety helmet calculated by the high-precision positioning module is used as the vertex, and the direction vector representing the front of the lens, calculated in space based on the real-time attitude data obtained above, is used as the central axis. Combined with the horizontal and vertical field of view parameters as geometric divergence boundary conditions, the above-mentioned three-dimensional projection cone is generated through a spatial analytical geometry algorithm. The spatial section solution matrix equation is established, and the two-dimensional geographic layer containing the actual geographic information of the construction site surface is used as the spatial section. The intersection polygon area between the three-dimensional projection cone and the plane of the preset electronic map is calculated. The process of establishing the spatial section solution matrix equation includes: constructing a camera intrinsic parameter matrix based on camera imaging parameters (such as horizontal and vertical field of view). A rotation matrix is constructed based on real-time attitude data (yaw, pitch, roll). ; and use spatial coordinate data as the camera's translation vector. Assuming the pixel coordinates at the edge of the field of view in the target video frame are (u, v), and their corresponding real-world 3D spatial coordinates are (X, Y, Z), then the two satisfy the following projection matrix equation:
[0098]
[0099] Where s is the depth scaling factor. Further, the plane of the preset electronic map is set as the horizontal plane of the Earth's surface in the world coordinate system. That is, Z=0 is used as the spatial tangent condition and substituted into the above equation. Since the pixel coordinates (u, v) of the four corner points of the target video frame are known, the two-dimensional mapping coordinates (X, Y) of these four corner points on the real physical surface can be solved through matrix inverse operations. The area formed by connecting these four two-dimensional coordinate points sequentially is determined as the intersecting polygonal region (field of view coverage) corresponding to the target video frame. This calculation process reduces the three-dimensional stereoscopic view to two-dimensional surface polygonal "spots." After plane coordinate system transformation, the intersecting polygonal region containing multiple two-dimensional coordinate vertices is determined as the precise field of view coverage of the currently extracted target video frame in the real physical world, thus accurately mapping the camera's visual image to the physical territory on the two-dimensional map.
[0100] S206: Count the number of personnel positioning card signal data within the field of view coverage to obtain the first number of on-duty personnel, and use a preset deep neural network model to detect personnel entity targets in the target video frame, and count the number of personnel entity targets to obtain the second number of on-duty personnel.
[0101] For example, this step aims to construct a "dual-mode counting" mechanism based on cross-validation of IoT spatial perception and computer vision. On the one hand, the system counts the total number of location card beacons currently within the physical polygon range covered by the smart safety helmet's lens, based on geospatial mapping relationships (i.e., the first number of on-duty personnel, representing the "theoretically expected number of personnel" recorded by the system); on the other hand, it directly calls AI vision algorithms as a visual verification mechanism to perform target instance-level feature stripping and statistics on the real two-dimensional images captured by the lens (i.e., the second number of on-duty personnel, representing the "actual number of people seen" captured by the camera). Through this dual-track counting method of "radio frequency digital signals" and "physical visual images," it can effectively overcome the business pain points of traditional single-monitoring methods, such as "the card is there but the person is not" (e.g., workers hang location cards or safety helmets on construction scaffolds to clock in) or the pure vision solution being easily affected by complex lighting and blind spots. This provides a dual comparison benchmark with high fault tolerance and high reliability for subsequent hazard assessment.
[0102] In one possible implementation, the number of personnel positioning card signal data within the field of view coverage area is counted to obtain the first number of on-duty personnel. Specifically, this includes: parsing the received personnel positioning card signal data using positioning base stations (such as UWB base station arrays) deployed in the target work area; calculating the absolute global two-dimensional coordinates of each personnel positioning card at the current moment using time difference of arrival or time-of-flight positioning algorithms; extracting the set of boundary vertex coordinates of the intersecting polygonal regions corresponding to the field of view coverage area; uniformly converting the absolute global two-dimensional coordinates of each personnel positioning card to the plane coordinate system of the electronic map; and using a preset spatial inclusion determination algorithm (such as ray intersection method or corner method) to calculate whether the absolute global two-dimensional coordinates of each personnel positioning card fall within the boundary of the intersecting polygonal region; filtering and counting the effective number of personnel positioning cards whose coordinate points fall within the intersecting polygonal region, and determining the effective number as the first number of on-duty personnel.
[0103] In one possible implementation, a preset deep neural network model is used to detect personnel entities in the target video frame, and the number of personnel entities is counted to obtain the second number of on-duty personnel. Specifically, this includes: inputting the target video frame into the deep neural network model for feature extraction to obtain a personnel spatial feature vector; performing target classification and bounding box regression processing on the personnel spatial feature vector to generate candidate personnel prediction boxes and corresponding classification confidence scores; filtering candidate personnel prediction boxes with classification confidence scores greater than a preset confidence score threshold, counting the number of candidate personnel prediction boxes after filtering, and obtaining the initial number of entities; based on the detection results of historical video frames, obtaining the number of multiple initial entities generated consecutively within a preset time window, performing consistency verification processing on each initial entity number to generate a stable number of entities, and determining the stable number of entities as the second number of on-duty personnel.
[0104] In this embodiment, consistency verification processing refers to a computer processing mechanism that smooths and denoises multiple discrete algorithm recognition results in the time domain. It is used to represent an anti-interference process to eliminate sudden jumps in the number of data caused by video image jitter, brief occlusion by people, or false detection of a single frame by the model. For example, the consistency verification processing may extract the recognition count of 15 consecutive single frames generated in the past 3 seconds and calculate the statistical mode. When there are multiple modes with the same frequency or the data is extremely scattered, a weighted time average algorithm within a sliding window (giving higher weight to frames closer to the current time) is further used to smooth the values, thereby transforming the unstable values into a continuous and stable value that best matches the actual situation on site.
[0105] Specifically, after preprocessing operations such as image scaling and normalization, the extracted target video frames are input into a deep neural network model based on a convolutional architecture (such as YOLO or Faster R-CNN series) for multi-scale feature extraction. The deep convolutional kernels within the network capture the edge, texture, and high-level semantic distribution patterns between pixels, reducing the dimensionality to obtain a spatial feature vector of people containing high-dimensional spatial features. Through the prediction head structure at the end of the network, target classification (determining the probability that a feature region belongs to the "person" class) and bounding box regression processing (calculating the center point coordinates and length and width offsets of the feature region in the original image) are performed on this spatial feature vector. This generates a large number of candidate person prediction boxes surrounding suspected targets in the image, and simultaneously calculates the classification confidence of each prediction box for whether it corresponds to a real worker. To eliminate false positives caused by complex backgrounds at the construction site (such as scaffolding and reflective objects), non-maximum suppression (NMS) or hard threshold comparison logic is used to rigorously screen the classifications. Candidate prediction boxes with a confidence level greater than a preset confidence threshold (e.g., setting a lower limit of 0.8) are generated, and an accumulation counting function is called to count the number of candidate prediction boxes after screening, thus obtaining the preliminary result of personnel identification in the current single frame, i.e., the number of initially screened entities. Furthermore, considering that the numerical fluctuations in a single frame are easily caused by local occlusion, and the single result cannot be directly accepted, based on the detection results of historical video frames temporarily stored in the memory queue, the number of multiple initially screened entities generated continuously within a preset duration window (e.g., 2 consecutive seconds or 60 frames) is retrieved. Using the temporal filtering or mode smoothing algorithm explained above, the consistency verification processing of each initially screened entity number collected due to the time sequence is performed, and abnormally high or low instantaneous values caused by sudden environmental interference are removed, generating a stable number of entities after smoothing and denoising. This stable number of entities is determined as the second number of on-duty personnel at the current moment, thereby completing the accurate leap from the bottom visual pixels to the high-level business supervision indicators.
[0106] In one possible implementation, the target video frame is input into a deep neural network model for feature extraction to obtain a spatial feature vector of people. Specifically, this includes: dividing the target video frame input into the deep neural network model into grid regions of a preset size, and extracting the initial spatial feature matrix of each grid region by performing a preset convolution operation; inputting each initial spatial feature matrix into the spatial attention layer in the deep neural network model to perform a feature weight redistribution operation to generate an enhanced spatial feature matrix corresponding to each grid region; obtaining the spatial arrangement sequence of each grid region in the target video frame, and flattening and stitching each enhanced spatial feature matrix according to the spatial arrangement sequence to obtain the spatial feature vector of people.
[0107] In the embodiments of this application, the spatial attention layer refers to a network structure embedded inside a deep neural network model, specifically used to dynamically adjust pixel-level response values based on the local spatial importance of input features. It represents a weight learning mechanism that enables the model to adaptively "focus" on people and entities in the image while suppressing background clutter (such as scaffolding, tower cranes, and other irrelevant environments at a construction site). For example, the spatial attention layer can be a spatial attention submodule based on the Convolutional Block Attention Module (CBAM) architecture. It generates a spatial attention weight map by performing max pooling and average pooling operations along the channel dimension, and then performs a dot product operation with the original features to highlight the core features.
[0108] Specifically, in the forward propagation phase of the object detection task, the target video frame, after pre-scale scaling and tensor quantization, is input into the deep neural network model and divided into grid regions of a preset size on the physical or logical receptive field according to a set aspect ratio (e.g., divided into a 13×13 or 26×26 two-dimensional matrix topology). Multiple deep convolutional kernels cascaded in the model's backbone network slide across the image matrix, performing preset convolution operations (including linear inner product calculation, batch normalization, and non-linear activation function mapping) to extract and dimensionality-reducedly fuse the initial semantic information containing edges, textures, and shallow local semantics within each grid region layer by layer. Initial spatial feature matrices are used as input tensors, considering that the complex background of the construction site can easily interfere with the accurate extraction of personnel features. These initial spatial feature matrices are then fed into the spatial attention layer of the deep neural network model. Inside this layer, the contribution of each spatial pixel to the specific detection target of "personnel entity" is evaluated, and feature weight redistribution is performed on the input tensors accordingly. This means that grid regions containing personnel feature contours are given higher response weights, while grid regions containing only irrelevant background features are penalized and suppressed. This filters out redundant environmental noise and generates enhanced spatial feature matrices with high signal-to-noise ratios for each grid region.
[0109] Furthermore, to ensure that the relative positional relationships of personnel targets in the real physical two-dimensional world are not lost in the subsequent feature fusion stage, the absolute topological adjacency relationships and spatial arrangement sequences of each grid region in the original target video frame plane are obtained through a coordinate mapping table (e.g., using a row-by-row scanning index sorting rule from left to right and from top to bottom). Under the premise of strictly locking the spatial topological structure, each enhanced spatial feature matrix after attention enhancement processing is one-dimensionalized according to the spatial arrangement sequence, and flattening and splicing processes are performed sequentially along a specific data channel dimension. In this way, a personnel spatial feature vector that integrates high-order semantic information and completely preserves the original spatial position mapping structure is constructed in the underlying memory, providing solid data support for subsequent accurate bounding box regression and classification actions.
[0110] S207: When the number of employees on duty is inconsistent with the number of employees on duty in the second place, or when the qualification level is lower than the access level requirement, generate and output a safety hazard warning instruction.
[0111] For example, this step constitutes the final decision-making and response hub of the entire intelligent construction site monitoring method. By cross-checking the results of IoT radio frequency sensing (locating the number of people on the card) and AI visual sensing (identifying the number of people in the body), the system can overcome the blind spots of a single sensor and effectively identify complex on-site violation scenarios such as "person and card separation," "lost location card," or "unauthorized personnel following and intruding without a card." At the same time, with the help of dynamic permission comparison based on geofencing, the system can accurately predict and intercept "overstepping construction" or "unauthorized entry" behaviors in special high-risk operation areas. Once any of the above-mentioned hazard judgment rules are met, the system will immediately aggregate and tag the current abnormal state, the triggering reason (such as abnormal number of people or insufficient permissions), and related on-site environmental information, generate standardized safety hazard warning instructions, and push them to the back-end monitoring screen or the communication device of the safety specialist to drive subsequent on-site intervention, voice warnings to drive away people, and violation management, thereby realizing an automated closed loop from perception and judgment to alarm.
[0112] In one possible implementation, when the number of people on duty at the first position is inconsistent with the number of people on duty at the second position, or when the qualification level is lower than the access level requirement, a safety hazard warning instruction is generated and output. Specifically, this includes: calculating the difference between the number of people on duty at the first position and the number of people on duty at the second position; when the difference is not zero, extracting the target video frame and spatial coordinate data as on-site evidence data, and generating a first sub-warning instruction containing an illegal intrusion identifier and on-site evidence data; obtaining a first quantitative value corresponding to the qualification level and a second quantitative value corresponding to the access level requirement; when the first quantitative value is less than the second quantitative value, obtaining the device identifier of the smart safety helmet, and associating the facial feature data, qualification level, and device identifier to obtain personnel identity evidence data, and generating a second sub-warning instruction containing an unauthorized operation identifier and personnel identity evidence data; encapsulating the generated first sub-warning instruction and / or second sub-warning instruction to generate a safety hazard warning instruction, and outputting the safety hazard warning instruction to a preset management terminal.
[0113] In this embodiment of the application, the on-site evidence data refers to a collection of electronic evidence that is automatically captured and solidified by the underlying program at the moment the security violation rule is triggered. This collection contains multimodal information such as objective visual images and absolute geographical location. It is used to represent the most direct and tamper-proof evidence for tracing violations when anomalies such as illegal intrusion occur at the construction site. For example, the on-site evidence data may be a high-definition photo captured by surveillance with precise latitude and longitude watermarks and timestamps, as well as a GPS coordinate parameter file packaged together.
[0114] Specifically, after completing the cross-comparison and verification of multimodal data, the final early warning judgment and triggering logic is entered. When it is continuously monitored that the first number of on-duty personnel representing physical radio frequency signals is inconsistent with the second number of on-duty personnel representing visual computing features, or when the qualification level of the personnel's identity and permissions is lower than the access level requirement set by the current electronic fence, the early warning generation logic is immediately triggered to generate and output a safety hazard early warning command. Specifically, this includes: firstly executing the personnel matching and verification branch, and accurately calculating the difference between the first and second number of on-duty personnel through arithmetic operations to check whether there are any abnormal personnel who are not carrying positioning devices or whose vision has been missed. When the arithmetic difference is not zero, it is determined that there is an intrusion or equipment malfunction in the current area. At this time, the cache interface is immediately invoked to extract the real-time target video frame at the moment of the discrepancy in the number of people and the current absolute spatial coordinate data as on-site evidence data with legal validity and traceability value. When generating the evidence data, the current absolute spatial coordinate data, the precise timestamp of the time of occurrence, and the unique device identifier of the smart safety helmet are directly embedded into the pixel matrix of the real-time target video frame as an invisible digital watermark using a steganography algorithm (such as the LSB algorithm) to prevent visual tampering. Then, combined with the underlying alarm communication protocol, a packet is generated. The first sub-warning instruction contains a specific alarm error code, i.e., an illegal intrusion identifier, and the aforementioned on-site evidence data; simultaneously or in parallel, the personnel permission verification branch executes a first quantitative value (e.g., mapping "basic operation qualification" to the value 1) that can be used for low-level mathematical comparison of the qualification level in character form through a preset level mapping dictionary, and obtains a second quantitative value (e.g., mapping "limited to intermediate and above qualification access" to the value 2) that corresponds to the access level requirement. When the numerical comparison determines that the first quantitative value is less than the second quantitative value, a violation of the level requirement is confirmed, and at this time, the underlying device interface is immediately used to obtain... The smart safety helmet hardware has a unique device identifier (such as a MAC address or device serial number). It deeply associates the previously extracted facial feature data, specific qualification level, and the device identifier with the device identifier through a hash algorithm or structured message. It then uses a preset asymmetric encryption algorithm private key to calculate and generate a digital signature for the associated data, or synchronously reports the hash digest of the data to a preset blockchain evidence storage node. This results in personnel identity evidence data that is used to accurately identify the responsible person and has absolute unforgeability and judicial evidentiary effect. In turn, it generates an unauthorized operation identifier containing the corresponding error code and a second sub-warning instruction for the personnel identity evidence data.
[0115] Furthermore, to ensure the efficiency and integrity of the early warning data transmission, the first and / or second sub-early warning instructions (which may be a single triggered instruction or a set of two triggered instructions, depending on the actual violation) are formatted, encoded, and encapsulated according to a preset network transmission protocol (such as MQTT or TCP / IP). This generates the final structured safety hazard early warning instruction, which aggregates a multi-dimensional, tamper-proof evidence chain. The instruction is then sent in real time to a preset management terminal (such as the central control monitoring screen at the construction site or the mobile device of the safety specialist) via a wireless communication link. This enables the accurate detection of safety hazards at the construction site, automatic evidence consolidation, and a closed-loop management system.
[0116] Figure 3 This is a schematic diagram of a construction site safety monitoring system based on a smart safety helmet, as described in an embodiment of this application. This system can be implemented through software, hardware, or a combination of both, forming all or part of a larger system. Figure 3 As shown, the system includes:
[0117] The qualification verification module 301 is configured to acquire the facial feature data of the wearer collected by the smart safety helmet, compare the facial feature data with the preset qualification database, and extract the qualification level corresponding to the wearer from the qualification database.
[0118] The data acquisition module 302 is configured to acquire video stream data and spatial coordinate data of the target work area collected by the smart safety helmet in real time, and to acquire personnel positioning card signal data located in the target work area;
[0119] The video frame extraction module 303 is configured to extract target video frames from video stream data using a preset sampling frequency.
[0120] The area matching module 304 is configured to match spatial coordinate data with a preset electronic map to determine the virtual work area where the smart safety helmet is currently located, and to obtain the pre-configured access level requirements of the virtual work area from the electronic map.
[0121] The field of view calculation module 305 is configured to acquire the camera imaging parameters of the smart safety helmet and calculate the field of view coverage corresponding to the target video frame based on the camera imaging parameters and spatial coordinate data.
[0122] The dual-mode counting module 306 is configured to count the number of personnel positioning card signal data within the field of view coverage to obtain the first number of on-duty personnel, and to use a preset deep neural network model to detect personnel entity targets in the target video frame and count the number of personnel entity targets to obtain the second number of on-duty personnel.
[0123] The safety early warning module 307 is configured to generate and output a safety hazard early warning command when the number of employees on duty is inconsistent with the number of employees on duty, or when the qualification level is lower than the access level requirement.
[0124] Based on the above embodiments, as an optional embodiment, the video frame extraction module 303 is specifically used to: establish a buffer queue with a preset capacity, and configure a first independent thread and a second independent thread for data interaction based on the buffer queue; obtain video frames from the video stream data through the first independent thread according to a preset connection strategy, and determine whether the buffer queue has reached a preset full load state; if the buffer queue is full, remove the historical video frame at the head of the buffer queue and store the currently obtained video frame at the tail of the buffer queue; otherwise, directly store the currently obtained video frame into the buffer queue; and extract video frames from the buffer queue based on the sampling frequency through the second independent thread to obtain the target video frame.
[0125] Based on the above embodiments, as an optional embodiment, the video frame extraction module 303 is specifically configured to: initiate a connection request for video stream data and obtain the feedback connection status of the connection request; when the feedback connection status is a successful connection status, extract video frames from the video stream data according to the normal acquisition branch in the connection strategy; when the feedback connection status is a failed connection status, perform a reconnection operation for the video stream data according to the abnormal reconnection branch in the connection strategy using a preset exponential backoff strategy; when the number of reconnection operations exceeds a preset threshold, release the video capture resources currently occupied by the first independent thread and re-initiate the connection request so as to extract video frames after restoring to a successful connection status.
[0126] Based on the above embodiments, as an optional embodiment, the field of view calculation module 305 is specifically used to: acquire the real-time attitude data of the smart safety helmet, and extract the horizontal field of view parameters and vertical field of view parameters from the camera imaging parameters; generate a three-dimensional projection cone with the spatial coordinate data as the vertex and the direction vector determined by the real-time attitude data as the central axis, combined with the horizontal field of view parameters and the vertical field of view parameters; calculate the intersecting polygonal region between the three-dimensional projection cone and the plane where the preset electronic map is located, and determine the intersecting polygonal region as the field of view coverage range corresponding to the target video frame.
[0127] Based on the above embodiments, as an optional embodiment, the dual-mode counting module 306 is specifically used for: inputting the target video frame into a deep neural network model for feature extraction to obtain a personnel spatial feature vector; performing target classification and bounding box regression processing on the personnel spatial feature vector to generate candidate personnel prediction boxes and corresponding classification confidence scores; filtering candidate personnel prediction boxes with classification confidence scores greater than a preset confidence score threshold, counting the number of candidate personnel prediction boxes after filtering, and obtaining the number of initially screened entities; based on the detection results of historical video frames, obtaining the number of multiple initially screened entities continuously generated within a preset duration window, performing consistency verification processing on each initially screened entity number, generating a stable entity number, and determining the stable entity number as the second number of on-duty personnel.
[0128] Based on the above embodiments, as an optional embodiment, the dual-mode counting module 306 is specifically used for: dividing the target video frame input to the deep neural network model into grid regions of a preset size, and extracting the initial spatial feature matrix of each grid region by performing a preset convolution operation; inputting each initial spatial feature matrix into the spatial attention layer in the deep neural network model to perform a feature weight redistribution operation to generate an enhanced spatial feature matrix corresponding to each grid region; obtaining the spatial arrangement sequence of each grid region in the target video frame, and flattening and splicing each enhanced spatial feature matrix according to the spatial arrangement sequence to obtain the personnel spatial feature vector.
[0129] Based on the above embodiments, as an optional embodiment, the safety warning module 307 is specifically used for: calculating the difference between the first number of on-duty personnel and the second number of on-duty personnel; when the difference is not zero, extracting the target video frame and spatial coordinate data as on-site evidence data, and generating a first sub-warning instruction containing an illegal intrusion identifier and on-site evidence data; obtaining a first quantitative value corresponding to the qualification level and a second quantitative value corresponding to the access level requirement; when the first quantitative value is less than the second quantitative value, obtaining the device identifier of the smart safety helmet, and associating the facial feature data, qualification level, and device identifier to obtain personnel identity evidence data, and generating a second sub-warning instruction containing an unauthorized operation identifier and personnel identity evidence data; encapsulating the generated first sub-warning instruction and / or second sub-warning instruction to generate a safety hazard warning instruction, and outputting the safety hazard warning instruction to a preset management terminal.
[0130] It should be noted that the system provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0131] This embodiment also discloses an electronic device, as shown in the reference. Figure 4 The electronic device may include: at least one processor 401, at least one communication bus 402, user interface 403, network interface 404, and at least one memory 405.
[0132] The communication bus 402 is used to enable communication between these components.
[0133] The user interface 403 may include a display screen and a camera. Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.
[0134] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0135] The processor 401 may include one or more processing cores. The processor 401 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 405, and by calling data stored in memory 405. Optionally, the processor 401 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 401 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 401.
[0136] The memory 405 may include random access memory (RAM) or read-only memory. Optionally, the memory 405 may include a non-transitory computer-readable storage medium. The memory 405 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 405 may also be at least one storage device located remotely from the aforementioned processor 401. Figure 4 As shown, the memory 405, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a construction site safety monitoring method based on a smart safety helmet.
[0137] exist Figure 4 In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 401 can be used to call an application program stored in the memory 405 that is a construction site safety monitoring method based on a smart safety helmet. When executed by one or more processors 401, the electronic device executes one or more methods as described in the above embodiments.
[0138] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0139] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0140] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.
[0141] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0142] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0143] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 405 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory 405 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.
[0144] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are considered exemplary only, and the scope of this application is defined by the claims.
Claims
1. A construction site safety monitoring method based on a smart safety helmet, characterized in that, The method includes: The facial feature data of the wearer collected by the smart safety helmet is acquired, and the facial feature data is compared with a preset qualification database to extract the qualification level corresponding to the wearer from the qualification database. The system acquires video stream data and spatial coordinate data of the target work area collected by the smart safety helmet in real time, and acquires the signal data of the personnel positioning card located in the target work area. The target video frame is obtained by extracting frames from the video stream data using a preset sampling frequency. The spatial coordinate data is matched with a preset electronic map to determine the virtual work area where the smart safety helmet is currently located, and the pre-configured access level requirements of the virtual work area are obtained from the electronic map. Obtain the camera imaging parameters of the smart safety helmet, and calculate the field of view coverage of the target video frame based on the camera imaging parameters and the spatial coordinate data; The number of personnel positioning card signal data within the field of view coverage is counted to obtain the first number of on-duty personnel, and the number of personnel entity targets in the target video frame is detected using a preset deep neural network model to obtain the second number of on-duty personnel. When the number of employees on duty is different from the number of employees on duty, or when the qualification level is lower than the access level requirement, a safety hazard warning instruction is generated and output.
2. The method according to claim 1, characterized in that, The step of obtaining the target video frame by performing frame extraction processing on the video stream data at a preset sampling frequency specifically includes: Establish a buffer queue with a preset capacity, and configure a first independent thread and a second independent thread for data interaction based on the buffer queue; The first independent thread acquires video frames from the video stream data according to a preset connection strategy, and determines whether the buffer queue has reached a preset full load state. If the buffer queue is in the full state, remove the historical video frame at the head of the buffer queue and store the currently acquired video frame at the tail of the buffer queue; otherwise, store the currently acquired video frame directly into the buffer queue. The second independent thread extracts the video frame from the buffer queue based on the sampling frequency to obtain the target video frame.
3. The method according to claim 2, characterized in that, The step of obtaining video frames from the video stream data through the first independent thread according to a preset connection strategy specifically includes: Initiate a connection request for the video stream data and obtain the feedback connection status of the connection request; When the feedback connection status is a successful connection status, the video frame is extracted from the video stream data according to the normal acquisition branch in the connection strategy. When the feedback connection status is a connection failure status, the video stream data is reconnected according to the abnormal reconnection branch in the connection strategy, using a preset exponential backoff strategy. When the number of reconnection operations exceeds a preset threshold, the video capture resources currently occupied by the first independent thread are released and the connection request is re-initiated so as to extract the video frames after the connection is restored to a successful state.
4. The method according to claim 1, characterized in that, The step of acquiring the camera imaging parameters of the smart safety helmet and calculating the field of view coverage corresponding to the target video frame based on the camera imaging parameters and the spatial coordinate data specifically includes: The real-time posture data of the smart safety helmet is obtained, and the horizontal field of view and vertical field of view parameters are extracted from the camera imaging parameters. Using the spatial coordinate data as the vertex and the direction vector determined by the real-time attitude data as the central axis, a three-dimensional projection cone is generated by combining the horizontal field of view parameter and the vertical field of view parameter. Calculate the intersecting polygonal region between the three-dimensional projection cone and the plane where the preset electronic map is located, and determine the intersecting polygonal region as the field of view coverage area corresponding to the target video frame.
5. The method according to claim 1, characterized in that, The step of using a preset deep neural network model to detect human entities in the target video frame and counting the number of human entities to obtain a second number of on-duty personnel specifically includes: The target video frame is input into the deep neural network model for feature extraction to obtain a spatial feature vector of people. Target classification and bounding box regression are performed on the spatial feature vector of the personnel to generate candidate personnel prediction boxes and corresponding classification confidence scores. Filter the candidate prediction boxes whose classification confidence is greater than a preset confidence threshold, count the number of candidate prediction boxes after filtering, and obtain the number of entities initially screened. Based on the detection results of historical video frames, the number of multiple initial screening entities generated consecutively within a preset time window is obtained. Consistency verification processing is performed on each of the initial screening entity numbers to generate a stable entity number, and the stable entity number is determined as the second number of on-duty personnel.
6. The method according to claim 5, characterized in that, The step of inputting the target video frame into the deep neural network model for feature extraction to obtain a spatial feature vector of people specifically includes: The target video frame input to the deep neural network model is divided into grid regions of a preset size, and the initial spatial feature matrix of each grid region is extracted by performing a preset convolution operation. Each of the initial spatial feature matrices is input into the spatial attention layer of the deep neural network model to perform a feature weight redistribution operation, thereby generating an enhanced spatial feature matrix corresponding to each of the grid regions. Obtain the spatial arrangement sequence of each grid region in the target video frame, and flatten and stitch each enhanced spatial feature matrix according to the spatial arrangement sequence to obtain the personnel spatial feature vector.
7. The method according to claim 1, characterized in that, When the number of employees on duty differs from the number of employees on duty, or when the qualification level is lower than the access level requirement, a safety hazard warning instruction is generated and output, specifically including: Calculate the difference between the first number of people on duty and the second number of people on duty. When the difference is not zero, extract the target video frame and the spatial coordinate data as on-site evidence data, and generate a first sub-early warning instruction containing an illegal intrusion identifier and the on-site evidence data. Obtain the first quantitative value corresponding to the qualification level and the second quantitative value corresponding to the access level requirement. When the first quantitative value is less than the second quantitative value, obtain the device identifier of the smart safety helmet, and associate the facial feature data, the qualification level and the device identifier to obtain personnel identity verification data, and generate a second sub-warning instruction containing an unauthorized operation identifier and the personnel identity verification data. The generated first sub-early warning instruction and / or second sub-early warning instruction are encapsulated to generate the safety hazard warning instruction, and the safety hazard warning instruction is output to a preset management terminal.
8. A construction site safety monitoring system based on a smart safety helmet, characterized in that, The system includes: The qualification verification module is configured to acquire facial feature data of the wearer collected by the smart safety helmet, compare the facial feature data with a preset qualification database, and extract the qualification level corresponding to the wearer from the qualification database. The data acquisition module is configured to acquire in real time the video stream data and spatial coordinate data of the target work area collected by the smart safety helmet, and to acquire the personnel positioning card signal data located in the target work area; A video frame extraction module is configured to extract target video frames from the video stream data using a preset sampling frequency. The area matching module is configured to match the spatial coordinate data with a preset electronic map to determine the virtual work area where the smart safety helmet is currently located, and to obtain the pre-configured access level requirements of the virtual work area from the electronic map. The field of view calculation module is configured to acquire the camera imaging parameters of the smart safety helmet and calculate the field of view coverage corresponding to the target video frame based on the camera imaging parameters and the spatial coordinate data. The dual-mode counting module is configured to count the number of personnel positioning card signal data within the field of view coverage to obtain a first number of on-duty personnel, and to use a preset deep neural network model to detect personnel entity targets in the target video frame and count the number of personnel entity targets to obtain a second number of on-duty personnel. The safety early warning module is configured to generate and output a safety hazard early warning command when the first number of on-duty personnel is inconsistent with the second number of on-duty personnel, or when the qualification level is lower than the access level requirement.
9. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Safety management system for tunnel constructors
CN117829533A
Building construction potential safety hazard real-time monitoring system based on AI visual identification
CN121415323A