A construction site safety helmet state monitoring method and system for multi-modal image recognition
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]现有的,传统的安全帽佩戴监测方法一般通过人工定时巡查的方式对工人的安全帽状态进行监测,或者通过感应器对佩戴状态进行监测,而上述现有技术对安全帽的状态监测较为单一,并且依赖施工人员的自觉性,存在监测不到位的情况,从而降低电力施工时的整体安全性,因此需要改进
[0014] By employing the aforementioned technical solution and integrating multimodal data such as visible light images, thermal infrared images, and RFID tag information for construction site safety helmet status monitoring, accurate and real-time safety helmet status identification and early warning are achieved. Utilizing the 3D model of the construction site and point features extracted from the construction plan, precise spatial positioning and construction background information are provided for safety helmet status monitoring. Simultaneously, the multimodal fusion model based on a self-learning algorithm continuously optimizes its safety helmet status identification capability, adapting to different construction scenarios and personnel changes. The introduction of RFID tags enhances the identification and management of construction personnel, improving the accuracy and relevance of early warnings. Overall, this method effectively improves the safety management level of power construction sites, reduces the risks faced by construction workers due to improper safety helmet wearing, thereby enhancing the safety of power construction and providing strong protection for power construction safety.
Smart Images

Figure CN121278309B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of intelligent monitoring of engineering safety, and in particular to a method and system for monitoring the status of construction safety helmets using multimodal image recognition. Background Technology
[0002] In the field of power construction, ensuring the safety and health of construction workers is of paramount importance. As a basic piece of personal protective equipment, the correct wearing of a construction helmet can effectively reduce the risk of head injuries during construction.
[0003] Existing traditional methods for monitoring the wearing of safety helmets generally involve manual, periodic inspections to monitor the status of workers' safety helmets, or the use of sensors to monitor the wearing status. However, these existing technologies are relatively simple in their monitoring of the safety helmet status and rely on the self-discipline of construction workers, which can lead to inadequate monitoring and reduce the overall safety of power construction. Therefore, improvements are needed. Summary of the Invention
[0004] To improve the safety of power construction, this application provides a method and system for monitoring the status of safety helmets at construction sites using multimodal image recognition.
[0005] Firstly, the above-mentioned inventive objective of this application is achieved through the following technical solution:
[0006] A method for monitoring the condition of construction site safety helmets using multimodal image recognition, the method comprising the following steps:
[0007] Acquire image data of the construction site and construct an overall three-dimensional model of the construction site based on the image data;
[0008] The construction plan is obtained to extract the features of the construction points, and the coordinates of the points are marked on the overall three-dimensional model.
[0009] Multimodal data of construction workers' safety helmets are collected, and the multimodal data is spatiotemporally aligned and preprocessed to obtain a multimodal frame sequence in a unified coordinate system. The multimodal data of the safety helmets includes visible light images, thermal infrared images and RFID tag information of the construction workers.
[0010] The acquired visible light and thermal infrared images are preprocessed to generate a multimodal image dataset;
[0011] The RFID tag information is used to perform spatiotemporal calibration on the multimodal image dataset, and the calibrated multimodal image data is associated and mapped with personnel identities to establish an image-identity dataset;
[0012] The pre-set multimodal fusion model uses a self-learning algorithm to extract and fuse features from calibrated image data in order to identify the status information of the safety helmet;
[0013] The pre-set early warning response model uses a self-learning algorithm to comprehensively analyze RFID tag information and status classification results, and determines whether to trigger the corresponding early warning mechanism.
[0014] By employing the aforementioned technical solution and integrating multimodal data such as visible light images, thermal infrared images, and RFID tag information for construction site safety helmet status monitoring, accurate and real-time safety helmet status identification and early warning are achieved. Utilizing the 3D model of the construction site and point features extracted from the construction plan, precise spatial positioning and construction background information are provided for safety helmet status monitoring. Simultaneously, the multimodal fusion model based on a self-learning algorithm continuously optimizes its safety helmet status identification capability, adapting to different construction scenarios and personnel changes. The introduction of RFID tags enhances the identification and management of construction personnel, improving the accuracy and relevance of early warnings. Overall, this method effectively improves the safety management level of power construction sites, reduces the risks faced by construction workers due to improper safety helmet wearing, thereby enhancing the safety of power construction and providing strong protection for power construction safety.
[0015] In a preferred example, this application can be further configured such that the step of extracting and fusing features from calibrated image data based on a self-learning algorithm using a pre-set multimodal fusion model to identify the status information of the safety helmet includes the following steps:
[0016] The visible light feature extraction subnetwork extracts the color, shape, and pattern features of the safety helmet in the visible light image;
[0017] The head contour and temperature distribution features in thermal infrared images are extracted using an external feature extraction subnetwork.
[0018] A cross-modal feature fusion module is used to perform channel-level fusion of visible light features and thermal infrared features, and the fusion weights are dynamically adjusted.
[0019] A pre-defined state classifier determines the wearing status of the safety helmet based on the fused features, where the wearing status includes correctly worn, not worn, and abnormally worn.
[0020] In a preferred embodiment, this application can be further configured such that, in the step of collecting multimodal data of construction workers' safety helmets at the construction site, data is collected by a visible light camera and an infrared thermal imager, and a synchronous triggering mechanism is adopted to ensure image alignment accuracy under complex lighting conditions. The synchronous triggering mechanism is provided with a timestamp calibration signal by an IMU sensor embedded in the safety helmet.
[0021] In a preferred embodiment, this application can be further configured such that, in the step of collecting multimodal data on the safety helmets of construction workers, an RFID reader is pre-installed to work with a passive RFID tag inside the safety helmet. The tag information includes the personnel's identity, job type, and safety helmet expiration date, which is used to dynamically update the personnel database and perform real-time identity verification.
[0022] In a preferred embodiment, this application can be further configured such that, in the step of fusing visible light features and thermal infrared features at the channel level through a cross-modal feature fusion module and dynamically adjusting the fusion weights, the cross-modal feature fusion module employs an attention mechanism to automatically adjust the fusion ratio of visible light features and thermal infrared features according to the ambient light intensity and thermal radiation level, so as to improve the recognition accuracy at different operating times.
[0023] In a preferred example, this application can be further configured such that, in the step of a pre-set state classifier determining the wearing status of a safety helmet based on fused features, the training dataset of the state classifier includes samples from different seasons and different types of work, and uses data augmentation techniques to simulate extreme environmental conditions such as rain, fog, and strong light to improve the generalization ability of the model.
[0024] In a preferred example, this application can be further configured as follows: in the step of the pre-set early warning response model performing a comprehensive analysis of RFID tag information and status classification results based on a self-learning algorithm, and determining whether to trigger the corresponding early warning mechanism, the early warning mechanism includes: when the same person is detected as not wearing the tag three times in a row, automatically suspending the operation of construction equipment in the area where the person is located, and sending an emergency notification to the on-site safety officer.
[0025] Secondly, the above-mentioned inventive objective of this application is achieved through the following technical solutions:
[0026] A multimodal image recognition-based construction site safety helmet status monitoring device, the device comprising: a construction site three-dimensional model construction unit, used to acquire image data of the construction site and construct an overall three-dimensional model of the construction site based on the image data;
[0027] The point coordinate marking unit is used to obtain the construction plan to extract the construction point features and to mark the point coordinates of the overall three-dimensional model;
[0028] A multimodal data acquisition unit is used to acquire multimodal data of construction workers' safety helmets at the construction site, and to perform spatiotemporal alignment and preprocessing on the multimodal data to obtain a multimodal frame sequence in a unified coordinate system. The multimodal data of the construction site safety helmets includes visible light images, thermal infrared images and RFID tag information of the construction workers.
[0029] The multimodal image dataset generation unit is used to preprocess the acquired visible light and thermal infrared images and generate a multimodal image dataset.
[0030] The image-identity dataset construction unit is used to perform spatiotemporal calibration of the multimodal image dataset using the RFID tag information, and to associate and map the calibrated multimodal image data with personnel identities to establish an image-identity dataset.
[0031] The safety helmet status information recognition unit is used to pre-set a multimodal fusion model to extract and fuse features from calibrated image data based on a self-learning algorithm in order to identify the status information of the safety helmet.
[0032] The early warning response unit is used to pre-set an early warning response model to comprehensively analyze RFID tag information and status classification results based on a self-learning algorithm, and to determine whether to trigger the corresponding early warning mechanism.
[0033] Thirdly, the above-mentioned objectives of this application are achieved through the following technical solutions:
[0034] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-described method for monitoring the status of a construction site safety helmet using multimodal image recognition.
[0035] Fourthly, the above-mentioned objectives of this application are achieved through the following technical solutions:
[0036] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for monitoring the status of construction safety helmets using multimodal image recognition. Attached Figure Description
[0037] Figure 1 This is a flowchart of a multimodal image recognition method for monitoring the status of construction site safety helmets according to an embodiment of this application;
[0038] Figure 2 This is a schematic diagram of a construction site safety helmet status monitoring device based on multimodal image recognition, according to one embodiment of this application.
[0039] Figure 3 This is a schematic diagram of an electronic device according to an embodiment of this application.
[0040] Icon labels:
[0041] 1. Three-dimensional model construction unit for construction site; 2. Point coordinate marking unit; 3. Multimodal data acquisition unit; 4. Multimodal image dataset generation unit; 5. Image-identity dataset construction unit; 6. Safety helmet status information recognition unit; 7. Early warning response unit. Detailed Implementation
[0042] The present application will be further described in detail below with reference to the accompanying drawings.
[0043] In one embodiment, such as Figure 1 As shown, this application discloses a method for monitoring the condition of construction site safety helmets using multimodal image recognition, specifically including the following steps:
[0044] S10: Acquire image data of the construction site and construct an overall three-dimensional model of the construction site based on the image data;
[0045] Specifically, in this embodiment, multiple lidar and UAV mapping systems are deployed at the power construction site to scan and acquire point cloud data and high-definition texture maps of the construction site. Data processing is performed using the PCL point cloud library, including filtering, downsampling, and surface reconstruction, to finally construct a detailed three-dimensional model of the construction site. This model can accurately present the terrain, building outlines, and placement of various construction equipment at the construction site, providing a spatial positioning basis for subsequent safety helmet status monitoring.
[0046] S20: Obtain the construction plan to extract the features of the construction points, and mark the coordinates of the points on the overall three-dimensional model;
[0047] Specifically, the process involves acquiring power construction plan documents and using natural language processing technology to extract key feature information of construction sites, such as transformer installation areas and tower erection areas. These key construction sites are then mapped onto the overall 3D model constructed by S10, and geometric algorithms are used to calculate and mark the coordinates of each construction site, assigning each site a unique 3D coordinate identifier.
[0048] S30: Collect multimodal data of safety helmets worn by construction workers at the site, and perform spatiotemporal alignment and preprocessing on the multimodal data to obtain a multimodal frame sequence in a unified coordinate system;
[0049] The multimodal data of construction site safety helmets includes visible light images, thermal infrared images, and RFID tag information of construction workers.
[0050] Specifically, in step S30, images are acquired using a visible light camera and an infrared thermal imager, and a synchronous triggering mechanism is employed to ensure image alignment accuracy under complex lighting conditions. The synchronous triggering mechanism is provided with a timestamp calibration signal by an IMU sensor embedded in the safety helmet.
[0051] Furthermore, an RFID reader is pre-installed to work with the passive RFID tag inside the safety helmet. The tag information includes personnel identity, job type, and safety helmet expiration date, which is used to dynamically update the personnel database and perform real-time identity verification.
[0052] S40: Preprocess the acquired visible light and thermal infrared images and generate a multimodal image dataset;
[0053] Specifically, preprocessing operations are performed on the acquired visible light and thermal infrared images. This includes grayscale conversion of visible light images to reduce computational complexity and temperature normalization of thermal infrared images to unify the temperature range. Based on this, data augmentation techniques such as rotation, scaling, and translation are used to expand the sample size, constructing a multimodal image dataset containing various construction scenarios and personnel states, providing rich and diverse data samples for model training.
[0054] S50: Use the RFID tag information to perform spatiotemporal calibration on the multimodal image dataset, and associate the calibrated multimodal image data with personnel identity to establish an image-identity dataset;
[0055] Specifically, the RFID tag information collected by S30 is used to perform spatiotemporal calibration on the multimodal image dataset generated by S40. The personnel identification ID in the tag information is matched with the personnel features in the image data. An association algorithm is used to accurately map the calibrated multimodal image data with personnel identification, and an image-identity dataset containing personnel identification information is established to provide data support for subsequent helmet status monitoring based on personnel identification.
[0056] S60: The pre-set multimodal fusion model extracts and fuses features from calibrated image data based on a self-learning algorithm to identify the status information of the safety helmet;
[0057] Specifically, a pre-set multimodal fusion model is used. This model, based on a self-learning algorithm, first extracts features from the calibrated image data, including edge and texture features from visible light images and temperature distribution features from thermal infrared images. The extracted features from different modalities are then fused. In this embodiment, a deep learning algorithm, such as a convolutional neural network (CNN), is used to learn and train the fused features, ultimately achieving accurate identification of the helmet's status, including correct wearing, not wearing, and improper wearing.
[0058] S70: The pre-set early warning response model comprehensively analyzes RFID tag information and status classification results based on a self-learning algorithm, and determines whether to trigger the corresponding early warning mechanism;
[0059] Specifically, in the step of the pre-set early warning response model comprehensively analyzing RFID tag information and status classification results based on a self-learning algorithm and determining whether to trigger the corresponding early warning mechanism, the early warning mechanism includes: when the same person is detected not wearing the RFID tag three times in a row, the operation of the construction equipment in the area where the person is located is automatically suspended, and an emergency notification is sent to the on-site safety officer.
[0060] Furthermore, a pre-set early warning response model, also based on a self-learning algorithm, is implemented. This model first comprehensively analyzes the RFID tag information in S50 and the safety helmet status classification results in S60. Different early warning thresholds are set based on the construction worker's identity information and the risk level of their construction site. When situations such as not wearing a safety helmet or wearing it improperly occur and reach the early warning threshold, the corresponding early warning mechanism is triggered promptly, such as sending alarm information to site management personnel or issuing audible and visual alarms on-site, to ensure the safety of construction workers.
[0061] In summary, for steps S10-S70, by fusing multiple modal data such as visible light images, thermal infrared images, and RFID tag information for construction site safety helmet status monitoring, accurate and real-time safety helmet status identification and early warning are achieved. Utilizing the 3D model of the construction site and the point features extracted from the construction plan, precise spatial positioning and construction background information are provided for safety helmet status monitoring. Simultaneously, the multimodal fusion model based on a self-learning algorithm continuously optimizes its ability to identify safety helmet status, adapting to different construction scenarios and personnel changes. The introduction of RFID tags enhances the identification and management of construction personnel, improving the accuracy and relevance of early warnings. Overall, this method effectively improves the safety management level of power construction sites, reduces the risks faced by construction personnel due to improper safety helmet wearing, thereby improving the safety of power construction and providing strong protection for power construction safety.
[0062] In step S60: The pre-configured multimodal fusion model extracts and fuses features from calibrated image data based on a self-learning algorithm to identify the helmet's status information. This step includes the following steps:
[0063] S61: Extract the color, shape, and pattern features of the safety helmet in the visible light image through the visible light feature extraction sub-network;
[0064] S62: Extract head contour and temperature distribution features from thermal infrared images through an external feature extraction subnetwork;
[0065] S63: Visible light features and thermal infrared features are fused at the channel level through a cross-modal feature fusion module, and the fusion weights are dynamically adjusted;
[0066] Specifically, in the step of fusing visible light features and thermal infrared features at the channel level through a cross-modal feature fusion module and dynamically adjusting the fusion weights, the cross-modal feature fusion module adopts an attention mechanism to automatically adjust the fusion ratio of visible light features and thermal infrared features according to the ambient light intensity and thermal radiation level, so as to improve the recognition accuracy at different operating times.
[0067] S64: The pre-set state classifier determines the wearing status of the safety helmet based on the fused features, where the wearing status includes correctly worn, not worn, and abnormally worn;
[0068] Specifically, in the step where the pre-set state classifier determines the wearing status of the safety helmet based on the fused features, the training dataset of the state classifier includes samples from different seasons and different types of work, and uses data augmentation techniques to simulate extreme environmental conditions such as rain, fog, and strong light to improve the generalization ability of the model.
[0069] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0070] In one embodiment, a multimodal image recognition-based construction site safety helmet status monitoring device is provided, which corresponds one-to-one with the multimodal image recognition-based construction site safety helmet status monitoring method described in the above embodiments. For example... Figure 2 As shown, the multimodal image recognition construction site safety helmet status monitoring device includes: a construction site three-dimensional model construction unit 1, used to acquire image data of the construction site and construct an overall three-dimensional model of the construction site based on the image data;
[0071] Point coordinate marking unit 2 is used to obtain the construction plan to extract the construction point features and to mark the point coordinates of the overall three-dimensional model;
[0072] The multimodal data acquisition unit 3 is used to acquire multimodal data of construction workers' safety helmets at the construction site, and to perform spatiotemporal alignment and preprocessing on the multimodal data to obtain a multimodal frame sequence in a unified coordinate system. The multimodal data of the construction site safety helmets includes visible light images, thermal infrared images and RFID tag information of the construction workers.
[0073] The multimodal image dataset generation unit 4 is used to preprocess the acquired visible light images and thermal infrared images and generate a multimodal image dataset.
[0074] Image-identity dataset construction unit 5 is used to perform spatiotemporal calibration of the multimodal image dataset using the RFID tag information, and to associate and map the calibrated multimodal image data with personnel identities to establish an image-identity dataset;
[0075] The safety helmet status information recognition unit 6 is used to pre-set a multimodal fusion model to extract and fuse features from calibrated image data based on a self-learning algorithm in order to identify the status information of the safety helmet;
[0076] The early warning response unit 7 is used to pre-set an early warning response model to comprehensively analyze RFID tag information and status classification results based on a self-learning algorithm, and to determine whether to trigger the corresponding early warning mechanism.
[0077] Specific limitations regarding the multimodal image recognition-based construction helmet condition monitoring device can be found in the above-mentioned limitations on the multimodal image recognition-based construction helmet condition monitoring method, and will not be repeated here. Each module in the aforementioned multimodal image recognition-based construction helmet condition monitoring device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the electronic device, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module.
[0078] In one embodiment, an electronic device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the electronic device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the database. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a multimodal image recognition method for monitoring the status of construction site safety helmets.
[0079] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0080] Acquire image data of the construction site and construct an overall three-dimensional model of the construction site based on the image data;
[0081] The construction plan is obtained to extract the features of the construction points, and the coordinates of the points are marked on the overall three-dimensional model.
[0082] Multimodal data of construction workers' safety helmets are collected, and the multimodal data is spatiotemporally aligned and preprocessed to obtain a multimodal frame sequence in a unified coordinate system. The multimodal data of the safety helmets includes visible light images, thermal infrared images and RFID tag information of the construction workers.
[0083] The acquired visible light and thermal infrared images are preprocessed to generate a multimodal image dataset;
[0084] The RFID tag information is used to perform spatiotemporal calibration on the multimodal image dataset, and the calibrated multimodal image data is associated and mapped with personnel identities to establish an image-identity dataset;
[0085] The pre-set multimodal fusion model uses a self-learning algorithm to extract and fuse features from calibrated image data in order to identify the status information of the safety helmet;
[0086] The pre-set early warning response model uses a self-learning algorithm to comprehensively analyze RFID tag information and status classification results, and determines whether to trigger the corresponding early warning mechanism.
[0087] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0088] Acquire image data of the construction site and construct an overall three-dimensional model of the construction site based on the image data;
[0089] The construction plan is obtained to extract the features of the construction points, and the coordinates of the points are marked on the overall three-dimensional model.
[0090] Multimodal data of construction workers' safety helmets are collected, and the multimodal data is spatiotemporally aligned and preprocessed to obtain a multimodal frame sequence in a unified coordinate system. The multimodal data of the safety helmets includes visible light images, thermal infrared images and RFID tag information of the construction workers.
[0091] The acquired visible light and thermal infrared images are preprocessed to generate a multimodal image dataset;
[0092] The RFID tag information is used to perform spatiotemporal calibration on the multimodal image dataset, and the calibrated multimodal image data is associated and mapped with personnel identities to establish an image-identity dataset;
[0093] The pre-set multimodal fusion model uses a self-learning algorithm to extract and fuse features from calibrated image data in order to identify the status information of the safety helmet;
[0094] The pre-set early warning response model uses a self-learning algorithm to comprehensively analyze RFID tag information and status classification results, and determines whether to trigger the corresponding early warning mechanism.
[0095] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0096] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0097] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method of site safety hat condition monitoring using multi-modal image recognition, characterised in that, The method includes the steps of: acquiring image data of the construction site and constructing an overall three-dimensional model of the construction site based on the image data; The construction plan is obtained to extract the features of the construction points, and the coordinates of the points are marked on the overall three-dimensional model. Multimodal data of construction workers' safety helmets are collected, and the multimodal data is spatiotemporally aligned and preprocessed to obtain a multimodal frame sequence in a unified coordinate system. The multimodal data of the safety helmets includes visible light images, thermal infrared images and RFID tag information of the construction workers. The acquired visible light and thermal infrared images are preprocessed to generate a multimodal image dataset; The RFID tag information is used to perform spatiotemporal calibration on the multimodal image dataset, and the calibrated multimodal image data is associated and mapped with personnel identities to establish an image-identity dataset; The pre-configured multimodal fusion model uses a self-learning algorithm to extract and fuse features from calibrated image data to identify the status information of the safety helmet. Specifically, the visible light feature extraction subnetwork extracts the color, shape, and pattern features of the safety helmet in the visible light image; The head contour and temperature distribution features in thermal infrared images are extracted using an external feature extraction subnetwork. The cross-modal feature fusion module performs channel-level fusion of visible light features and thermal infrared features and dynamically adjusts the fusion weights. The cross-modal feature fusion module adopts an attention mechanism to automatically adjust the fusion ratio of visible light features and thermal infrared features according to the ambient light intensity and thermal radiation level, so as to improve the recognition accuracy at different working times. A pre-defined state classifier determines the wearing status of the safety helmet based on the fused features, where the wearing status includes correctly worn, not worn, and abnormally worn; The pre-set early warning response model uses a self-learning algorithm to comprehensively analyze RFID tag information and status classification results, and determines whether to trigger the corresponding early warning mechanism.
2. The method for monitoring the status of construction site safety helmets using multimodal image recognition according to claim 1, characterized in that, In the step of collecting multimodal data of construction workers' safety helmets, data is collected using a visible light camera and an infrared thermal imager, and a synchronous triggering mechanism is adopted to ensure image alignment accuracy under complex lighting conditions. The synchronous triggering mechanism is provided with a timestamp calibration signal by an IMU sensor embedded in the safety helmet. It is equipped with an RFID reader that works with the passive RFID tag inside the safety helmet. The tag information includes personnel identity, job type and safety helmet expiration date, which is used to dynamically update the personnel database and perform real-time identity verification.
3. The method for monitoring the status of construction site safety helmets using multimodal image recognition according to claim 2, characterized in that, In the step where the pre-set state classifier determines the wearing status of the safety helmet based on the fused features, the training dataset of the state classifier contains samples from different seasons and different types of work, and uses data augmentation techniques to simulate extreme environmental conditions such as rain, fog, and strong light to improve the generalization ability of the model.
4. The method for monitoring the status of construction site safety helmets using multimodal image recognition according to claim 1, characterized in that, In the step of the pre-set early warning response model comprehensively analyzing RFID tag information and status classification results based on a self-learning algorithm, and determining whether to trigger the corresponding early warning mechanism, the early warning mechanism includes: when the same person is detected as not wearing the RFID tag three times in a row, automatically suspending the operation of construction equipment in the area where the person is located, and sending an emergency notification to the on-site safety officer.
5. A multimodal image recognition-based construction site safety helmet status monitoring device, applied to the multimodal image recognition-based construction site safety helmet status monitoring method according to any one of claims 1-4, characterized in that, The device includes: a construction site three-dimensional model building unit (1), used to acquire image data of the construction site and build an overall three-dimensional model of the construction site based on the image data; Point coordinate marking unit (2) is used to obtain the construction plan to extract the construction point features and mark the point coordinates of the overall three-dimensional model; The multimodal data acquisition unit (3) is used to acquire multimodal data of construction workers' safety helmets at the construction site, and to perform spatiotemporal alignment and preprocessing on the multimodal data to obtain a multimodal frame sequence under a unified coordinate system. The multimodal data of the construction site safety helmets includes visible light images, thermal infrared images and safety helmet RFID tag information of construction workers. The multimodal image dataset generation unit (4) is used to preprocess the acquired visible light images and thermal infrared images and generate a multimodal image dataset. Image-identity dataset construction unit (5) is used to perform spatiotemporal calibration of the multimodal image dataset using the RFID tag information, and to associate and map the calibrated multimodal image data with personnel identity in order to establish an image-identity dataset; The safety helmet status information recognition unit (6) is used to pre-set a multimodal fusion model to extract and fuse features from calibrated image data based on a self-learning algorithm in order to identify the status information of the safety helmet; The early warning response unit (7) is used to pre-set an early warning response model to comprehensively analyze RFID tag information and status classification results based on a self-learning algorithm, and to determine whether to trigger the corresponding early warning mechanism.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of a construction site safety helmet status monitoring method based on multimodal image recognition as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of a construction site safety helmet status monitoring method based on multimodal image recognition as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Power operator behavior identification early warning system and method based on video analysis
CN120220241A
Improved YOLOv11s safety helmet wearing detection model and optimization method thereof
CN120356237A