A Real-time Safety Monitoring Method for Construction Workers Based on Multimodal Data Integration
A multi-modal data integration method for construction site safety monitoring improves risk identification accuracy by using physiological, audio, and location data, addressing the limitations of image-based systems in complex construction environments.
Patent Information
- Application Number
- CN202211182762.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-27
AI Technical Summary
In the prior art, the accuracy of identifying safety risks at construction sites is insufficient, especially when there are many workers and the environment is complex, and the accuracy of identifying safety risks through monitoring images is insufficient.
The multimodal data integration method is used to obtain the worker's physiological index data, voice signal data and geographical location data, and combine the target image to predict the security risk level through the prediction model. The prediction model includes the first encoding module, the second encoding module, the fusion module and the prediction module, and the feature extraction and fusion of the multimodal data are used for risk assessment.
It improves the accuracy of safety risks identification on construction sites, avoids the inaccurate identification of human behaviors caused by the complex environment of the construction site and the large number of workers, and improves the effectiveness of safety monitoring.
Smart Images

Figure CN115661737B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of construction site safety monitoring, and particularly relates to a real-time safety monitoring method for construction workers based on multi-modal data integration. Background Art
[0002] Construction sites are characterized by a complex environment, a large number of workers, and multiple types of work. Relying solely on manual observation of the images captured by surveillance cameras to determine whether there are safety accidents or potential hazards is often not timely enough. In the prior art, there is a method of performing safety monitoring by identifying human behaviors in the images captured by surveillance cameras based on intelligence. However, due to the large number of workers in construction sites and the variable nature of building materials, there are often many occlusions in the images, and the method of only identifying safety risks from surveillance images is not accurate enough.
[0003] Therefore, the prior art still needs to be improved and enhanced. Summary of the Invention
[0004] In view of the above-mentioned defects of the prior art, the present invention provides a real-time safety monitoring method for construction workers based on multi-modal data integration, aiming to solve the problem of insufficient accuracy in identifying safety risks in construction sites in the prior art.
[0005] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0006] In a first aspect of the present invention, there is provided a real-time safety monitoring method for construction workers based on multi-modal data integration, the method comprising:
[0007] Obtaining physiological index data, voice signal data, and geographical location data of each worker;
[0008] Obtaining a target image, wherein the target image includes a plurality of workers;
[0009] Inputting the physiological index data, the voice signal data, the geographical location data, and the target image into a trained prediction model to obtain a safety risk level output by the prediction model.
[0010] In the real-time safety monitoring method for construction workers based on multi-modal data integration, wherein the obtaining of the geographical location data of each worker includes:
[0011] Obtaining first position data fed back by wearable devices corresponding to each worker;
[0012] Obtaining second position data fed back by mobile devices corresponding to each worker;
[0013] Obtaining the geographical location data of each worker according to the first position data and the second position data respectively corresponding to each worker.
[0014] The real-time safety monitoring method for construction workers based on multi-modal data integration, wherein obtaining the geographical location data of each worker according to the first location data and the second location data respectively corresponding to each worker includes:
[0015] When the difference between the first location data and the second location data corresponding to a worker is within a preset range, take the average value of the first location data and the second location data as the geographical location data of the worker;
[0016] When the difference between the first location data and the second location data corresponding to a worker exceeds the preset range, discard the first location data and the second location data corresponding to the worker currently, and re-obtain the first location data and the second location data corresponding to the worker.
[0017] The real-time safety monitoring method for construction workers based on multi-modal data integration, wherein the prediction model includes a first encoding module, a second encoding module, a third encoding module, a fusion module and a prediction module, and obtaining the safety risk level output by the prediction model includes:
[0018] Input the physiological index data and the geographical location data into the first encoding module, and extract the features of the physiological index data and the geographical location data through the first encoding module and convert them into feature images;
[0019] Input the feature image and the target image into the second encoding module, and obtain the first feature vector output by the second encoding module;
[0020] Input the voice signal data into the third encoding module, and obtain the second feature vector output by the third encoding module;
[0021] Input the first feature vector and the second feature vector into the fusion module for fusion to obtain a target feature vector;
[0022] Input the target feature vector into the prediction module, and obtain the safety risk level output by the prediction module.
[0023] The real-time safety monitoring method for construction workers based on multi-modal data integration, wherein inputting the feature image and the target image into the second encoding module and obtaining the first feature vector output by the second encoding module includes:
[0024] For the image to be processed input into the second encoding module, perform the following operations:
[0025] Divide the image to be processed into multiple image patches, extract initial features for each image patch, and for each image patch, fuse the corresponding initial features and the position information of the image patch in the image to be processed to obtain the block features corresponding to each image patch respectively;
[0026] Perform an attention mechanism on the block features corresponding to each image patch respectively, and obtain the first feature vector according to the result of the attention mechanism.
[0027] The real-time safety monitoring method for construction workers based on multi-modal data integration, wherein the training process of the prediction model is as follows:
[0028] Construct an initial prediction model, and train the initial prediction model according to multiple groups of labeled data to obtain the prediction model;
[0029] Wherein, each group of labeled data includes sample physiological index data, sample voice signal data, sample geographical location data, sample target images, and safety risk level annotation results.
[0030] The real-time safety monitoring method for construction workers based on multi-modal data integration, wherein training the initial prediction model according to multiple groups of labeled data to obtain the prediction model includes:
[0031] Select target labeled data;
[0032] Input the sample physiological index data, sample voice signal data, sample geographical location data, and sample target images in the target labeled data into the initial prediction model, and obtain the safety risk level prediction result output by the initial prediction model;
[0033] Obtain the first loss according to the safety risk level prediction result and the safety risk level annotation result in the target labeled data;
[0034] Input the sample first feature vector and sample second feature vector obtained during the process of obtaining the safety risk level prediction result into the reconstruction module, and obtain the reconstructed data output by the reconstruction module;
[0035] Obtain the second loss based on the differences between the sample physiological index data, the sample voice signal data, the sample geographical location data, and the sample target images and the reconstructed data;
[0036] Obtain the training loss corresponding to the target labeled data according to the first loss and the second loss, and update the parameters of the initial prediction model according to the training loss corresponding to the target labeled data;
[0037] Repeat the step of selecting the target labeled data until the parameters of the initial prediction model converge, and use the model with converged parameters as the prediction model.
[0038] In a second aspect of the present invention, there is provided a real-time safety monitoring device for construction workers based on multi-modal data integration, comprising:
[0039] A first data acquisition module for acquiring physiological index data, voice signal data, and geographical location data of each worker;
[0040] A second data acquisition module for acquiring a target image, where the target image includes multiple workers;
[0041] A prediction module for inputting the physiological index data, the voice signal data, the geographical location data, and the target image into a trained prediction model to obtain a safety risk level output by the prediction model.
[0042] In a third aspect of the present invention, there is provided a terminal, which includes a processor and a computer-readable storage medium communicatively connected to the processor. The computer-readable storage medium is adapted to store multiple instructions, and the processor is adapted to call the instructions in the computer-readable storage medium to execute the steps of implementing the real-time safety monitoring method for construction workers based on multi-modal data integration described in any one of the above.
[0043] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the real-time safety monitoring method for construction workers based on multi-modal data integration described in any one of the above.
[0044] Compared with the prior art, the present invention provides a real-time safety monitoring method for construction workers based on multi-modal data integration. By collecting physiological index data, voice signal data, geographical location data of workers, and an image including multiple workers, the multi-modal data is input into a neural network model for predicting the safety risk level, avoiding inaccurate recognition of human behaviors caused by complex facilities and numerous workers in the construction site images. The present invention can improve the accuracy of safety risk recognition in the construction site. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a flowchart of an embodiment of the real-time safety monitoring method for construction workers based on multi-modal data integration provided by the present invention;
[0046] Figure 2 It is a schematic structural diagram of an embodiment of the real-time safety monitoring device for construction workers based on multi-modal data integration provided by the present invention;
[0047] Figure 3 Schematic diagram of the principle of an embodiment of the terminal provided by the present invention. Detailed implementation manners
[0048] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] The real-time safety monitoring method for construction workers based on multi-modal data integration provided by the present invention can be applied to a terminal with computing capabilities. The terminal can execute the real-time safety monitoring method for construction workers based on multi-modal data integration provided by the present invention for power grid peak shaving scheduling. The terminal can be but is not limited to various computers, mobile terminals, smart home appliances, wearable devices, etc.
[0050] Embodiment 1
[0051] As Figure 1 shown, in an embodiment of the real-time safety monitoring method for construction workers based on multi-modal data integration, the method includes the steps of:
[0052] S100. Obtain the physiological index data, voice signal data, and geographical location data of each worker;
[0053] S200. Obtain a target image, where the target image includes multiple workers.
[0054] In this embodiment, in order to prevent the reduction in the accuracy of human behavior recognition and subsequent judgment of safety risks when only using the monitoring images of the construction site due to the large number of workers and complex building materials in the construction site, multi-modal data is collected for identifying safety risks. Specifically, the physiological index data of the workers may include the heart rate, blood pressure, body temperature, blood oxygen saturation, etc. of the workers, and the collection of the physiological index data can be achieved through wearable devices such as smart bracelets and smart watches. The voice signal data can be collected through wearable devices or mobile devices such as handheld communicators and mobile phones. The geographical location data can be collected through wearable devices or mobile devices, and the target image is obtained by shooting with a camera set in the construction site.
[0055] The physiological index data, the voice signal data, and the geographical location data may be data collected within a preset time period before the current moment, and the target image may be multiple images taken within the preset time period before the current moment.
[0056] In a possible implementation manner, in order to solve the problem of insufficient positioning accuracy in a complex engineering site, two or more devices are used to obtain the geographical location data of each worker, including:
[0057] Obtain the first position data of the wearable device feedback corresponding to each worker;
[0058] Obtain the second position data of the mobile device feedback corresponding to each worker;
[0059] Obtain the geographical location data of each worker according to the first position data and the second position data respectively corresponding to each worker.
[0060] Further, the obtaining the geographical location data of each worker according to the first position data and the second position data respectively corresponding to each worker includes:
[0061] When the difference between the first position data and the second position data corresponding to a worker is within a preset range, take the average value of the first position data and the second position data as the geographical location data of the worker;
[0062] When the difference between the first position data and the second position data corresponding to a worker exceeds the preset range, discard the first position data and the second position data corresponding to the current worker, and re-obtain the first position data and the second position data corresponding to the worker.
[0063] Taking the wearable device as a smart bracelet and the mobile device as a handheld terminal as an example, cross-compare the position signals collected by the two. If the difference between the two exceeds the preset range, re-collect the position signals. If the difference between the two is within the preset range, take the average value of the two for subsequent processing.
[0064] Please refer to again Figure 1 , the method provided in this embodiment further includes the step of:
[0065] S300. Input the physiological index data, the voice signal data, the geographical location data, and the target image into the trained prediction model, and obtain the safety risk level output by the prediction model.
[0066] Specifically, the prediction model includes a first encoding module, a second encoding module, a third encoding module, a fusion module, and a prediction module. The obtaining the safety risk level output by the prediction model includes:
[0067] S310. Input the physiological index data and the geographical location data into the first encoding module, and extract the features of the physiological index data and the geographical location data through the first encoding module and convert them into feature images;
[0068] S320. Input the feature image and the target image into the second encoding module, and obtain the first feature vector output by the second encoding module;
[0069] S330. Input the voice signal data into the third encoding module, and obtain the second feature vector output by the third encoding module;
[0070] S340. Input the first feature vector and the second feature vector into the fusion module for fusion to obtain a target feature vector;
[0071] S350. Input the target feature vector into the prediction module, and obtain the security risk level output by the prediction module.
[0072] The first encoding module is used to extract the features of non-image and non-sound data and convert them into a 2D feature image. The first encoding module may include a filtering unit and a convolutional unit. After performing Kalman filtering on the input data through the filtering unit, the convolutional unit is used to extract features and convert them into a 2D image to obtain the feature image. For example, the convolutional unit extracts a feature matrix, and each value in the matrix is used as a pixel value in the image to obtain the feature image. The second encoding module is used for image processing. Specifically, inputting the feature image and the target image into the second encoding module to obtain the first feature vector output by the second encoding module includes:
[0073] For the image to be processed input into the second encoding module, perform the following operations:
[0074] Divide the image to be processed into multiple image blocks, extract initial features for each image block, and for each image block, fuse the corresponding initial features and the position information of the image block in the image to be processed to obtain the block features corresponding to each image block respectively;
[0075] Perform an attention mechanism on the block features corresponding to each image block respectively, and obtain the first feature vector according to the attention mechanism result.
[0076] Specifically, the first feature vector includes a first image feature vector and a second image feature. The second encoding module includes two processing modules: a first processing module and a second processing module. The first processing module is used to process the feature image obtained after feature extraction and conversion of the data to obtain the first image feature vector, and the second processing module is used to process the image captured by the imaging device to obtain the second image feature vector. The structures of the two processing models are the same, that is, the processing methods for the input images are the same. Taking the target image as an example below, the process of the second encoding module processing it is introduced:
[0077] Input the target image into the second processing module in the second encoding module. First, divide the target image into multiple image blocks. For each image block, initially extract features, which can be achieved through a convolutional layer. Then, for each image block, fuse the initial features corresponding to the image block and the position information of the image block in the target image to obtain the block feature corresponding to the image block. After that, for the block feature corresponding to each image block, execute the attention mechanism. It should be noted that the attention mechanism can be executed multiple times. After each execution of the attention mechanism, the block feature corresponding to each image block will be updated. Fuse (for example, directly concatenate) the block features corresponding to each image block after the last calculation of the attention mechanism to obtain the second image feature vector.
[0078] The process of obtaining the first image feature vector is the same as the process of obtaining the first image feature vector, except that the feature image is input into the first processing module in the second encoding module.
[0079] The training process of the prediction model is described below. The training process of the prediction model is as follows:
[0080] Construct an initial prediction model and train the initial prediction model according to multiple sets of labeled data to obtain the prediction model.
[0081] Each set of labeled data includes sample physiological index data, sample voice signal data, sample geographical location data, sample target images, and safety risk level annotation results.
[0082] Furthermore, in order to reduce the workload of data annotation, in this embodiment, first use open-source data to train the modules in the prediction model, and then use the labeled data in the construction site scenario for fine-tuning training, which can significantly reduce the requirement for the amount of labeled data in the construction site scenario.
[0083] The construction of the initial prediction model includes:
[0084] Use open-source data to train the second encoding module, the third encoding module, and the prediction module.
[0085] Based on the second encoding module, the third encoding module, and the prediction module trained with open-source data, construct the initial prediction model.
[0086] The second encoding module and the third encoding module can be pre-trained using open-source data with labels. Specifically, the open-source data used for pre-training the second encoding module and the third encoding module is data with classification labels. During the training process, only one of the first processing module and the second processing module in the second encoding module is trained, and the trained parameters are shared after training. Taking the training of the first processing module in the second encoding module as an example, during the training process, the sample data is input into the first processing module, classified based on the feature vectors output by the first processing module, and the parameters of the first processing module are updated according to the difference between the classification result and the classification label corresponding to the sample data, so that the first processing module has the function of initially extracting classification features, and then fine-tuning training is performed according to the corresponding construction site scene annotation data.
[0087] Training the initial prediction model according to multiple groups of annotation data to obtain the prediction model includes:
[0088] Selecting target annotation data;
[0089] Inputting the sample physiological index data, sample voice signal data, sample geographical location data, and sample target image in the target annotation data into the initial prediction model to obtain the safety risk level prediction result output by the initial prediction model;
[0090] Obtaining a first loss according to the safety risk level prediction result and the safety risk level annotation result in the target annotation data;
[0091] Inputting the sample first feature vector and sample second feature vector obtained during the process of obtaining the safety risk level prediction result into the reconstruction module to obtain the reconstructed data output by the reconstruction module;
[0092] Obtaining a second loss based on the difference between the sample physiological index data, the sample voice signal data, the sample geographical location data, and the sample target image and the reconstructed data;
[0093] Obtaining the training loss corresponding to the target annotation data according to the first loss and the second loss, and updating the parameters of the initial prediction model according to the training loss corresponding to the target annotation data;
[0094] Re-executing the step of selecting target annotation data until the parameters of the initial prediction model converge, and taking the model with converged parameters as the prediction model.
[0095] After selecting the target annotation data from the multiple sets of annotation data, the sample physiological index data, sample voice signal data, sample geographical location data, and sample target image in the target annotation data are input into the initial prediction model, and are processed in the manner of steps S310 to S350 in the foregoing text. That is, the sample physiological index data and the sample geographical location data are input into the first encoding module in the initial prediction model, and the first encoding module extracts the features of the physiological index data and the geographical location data and converts them into sample feature images. The sample feature image and the sample target image are input into the second encoding module to obtain a sample first feature vector output by the second encoding module. The sample voice signal data is input into the third encoding module to obtain a sample second feature vector output by the third encoding module. The sample first feature vector and the sample second feature vector are input into the fusion module for fusion to obtain a sample target feature vector. The sample target feature vector is input into the prediction module to obtain a safety risk level prediction result output by the prediction module.
[0096] In general model training, only the difference between the output of the model and the annotation result is used to obtain the loss to update the model parameters. In this embodiment, in order to improve the model training efficiency, a reconstruction module is set to train together with the initial prediction model. Specifically, in addition to obtaining the first loss according to the safety risk level prediction result and the safety risk level annotation result in the target annotation data, the sample first feature vector and the sample second feature vector are also input into the reconstruction module.
[0097] The sample first feature vector includes a sample first image feature vector and sample second image features. The sample first image feature vector is input into the reconstruction module, and the reconstruction module outputs a reconstructed first image. The sample second image features are input into the reconstruction module, and the reconstruction module outputs a reconstructed second image. The sample second feature vector is input into the reconstruction module, and the reconstruction module outputs a reconstructed voice signal. Specifically, different processing units can be respectively set in the reconstruction module to process the features of different modality data. A first sub-loss is obtained according to the difference between the reconstructed first image and the sample feature image. A second sub-loss is obtained according to the difference between the reconstructed second image and the sample target image. A third sub-loss is obtained according to the difference between the reconstructed voice signal and the sample voice signal data. The first sub-loss, the second sub-loss, and the third sub-loss are summed to obtain the second loss.
[0098] Sum the first loss and the second loss (which can be a weighted sum) to obtain the training loss corresponding to the target annotation data, and update the parameters of the initial prediction model based on the training loss corresponding to the target annotation data.
[0099] Further, in order to handle complex scenarios on construction sites, in this embodiment, one type of data other than the sample target image in the sample annotation data is randomly set to 0, that is, the sample physiological index data, the sample voice signal data, or the sample geographical location data in the target annotation data is set to 0. Specifically, updating the parameters of the initial prediction model according to the training loss corresponding to the target annotation data includes:
[0100] When the target annotation data is the selected data, set the sample physiological index data, the sample voice signal data, or the sample geographical location data in the target annotation data to 0 to obtain processed annotation data;
[0101] Obtain the processed target feature vector corresponding to the processed annotation data, and determine the third loss based on the difference between the processed target feature vector and the sample target feature vector corresponding to the target annotation data;
[0102] Update the parameters of the initial prediction model according to the training loss corresponding to the target annotation data and the third loss.
[0103] Among them, the selected data is randomly selected from the multiple sets of annotation data.
[0104] Updating the parameters of the initial prediction model according to the training loss corresponding to the target annotation data and the third loss is to sum the training loss corresponding to the target annotation data and the third loss, and update the parameters of the initial prediction model with the minimum sum as the optimization goal.
[0105] After inputting the processed annotation data into the initial prediction model, obtain the data output by the fusion module in the initial prediction model as the processed target feature vector. Updating the parameters of the initial prediction model with the minimum loss as the optimization goal can make the feature vectors for classifying risk levels extracted when there is data missing as close as possible to the feature vectors for classifying risk levels extracted when there is no data missing, so that in the case of missing a certain type of data in a complex engineering construction scenario, a certain accuracy of prediction can still be achieved, thus ensuring robustness and practicality.
[0106] In summary, this embodiment provides a real-time safety monitoring method for construction workers based on multi-modal data integration. By collecting physiological index data, voice signal data, geographical location data of workers, and images including multiple workers, the multi-modal data is input into a neural network model for predicting the safety risk level, avoiding inaccurate human behavior recognition caused by complex facilities and numerous workers in construction site images. The present invention can improve the accuracy of safety risk identification in construction sites.
[0107] It should be understood that although the steps in the flowcharts given in the accompanying drawings of the present invention are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0108] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0109] Embodiment 2
[0110] Based on the above embodiments, the present invention also correspondingly provides a real-time safety monitoring device for construction workers based on multi-modal data integration, as Figure 2 shown. The real-time safety monitoring device for construction workers based on multi-modal data integration includes:
[0111] A first data acquisition module, configured to acquire physiological index data, voice signal data, and geographical location data of each worker, specifically as described in Embodiment 1;
[0112] A second data acquisition module, configured to acquire a target image, where the target image includes multiple workers, specifically as described in Embodiment 1;
[0113] A prediction module, configured to input the physiological index data, the voice signal data, the geographical location data, and the target image into a trained prediction model, and obtain the safety risk level output by the prediction model, specifically as described in Embodiment 1.
[0114] Embodiment 3
[0115] Based on the above embodiments, the present invention also correspondingly provides a terminal, as Figure 3 shown. The terminal includes a processor 10 and a memory 20. Figure 3 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0116] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a real-time safety monitoring program 30 for construction workers based on multi-modal data integration is stored on the memory 20, and this real-time safety monitoring program 30 for construction workers based on multi-modal data integration can be executed by the processor 10, thereby implementing the real-time safety monitoring method for construction workers based on multi-modal data integration in this application.
[0117] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other chips, which are used to run the program code stored in the memory 20 or process data, such as executing the real-time safety monitoring method for construction workers based on multi-modal data integration, etc.
[0118] In one embodiment, when the processor 10 executes the real-time safety monitoring program 30 for construction workers based on multi-modal data integration in the memory 20, the following steps are implemented:
[0119] Obtain the physiological index data, voice signal data, and geographical location data of each worker;
[0120] Obtain a target image, where the target image includes multiple workers;
[0121] Input the physiological index data, the voice signal data, the geographical location data, and the target image into a trained prediction model to obtain the safety risk level output by the prediction model.
[0122] Embodiment 4
[0123] The present invention also provides a computer-readable storage medium, in which one or more programs are stored, and the one or more programs can be executed by one or more processors to implement the steps of the real-time safety monitoring method for construction workers based on multi-modal data integration as described above.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time safety monitoring method for construction workers based on multi-modal data integration, characterized in that, The method includes: Obtaining physiological index data, voice signal data, and geographical location data of each worker; Obtaining a target image, where the target image includes multiple workers; Inputting the physiological index data, the voice signal data, the geographical location data, and the target image into a trained prediction model to obtain the safety risk level output by the prediction model; The prediction model includes a first encoding module, a second encoding module, a third encoding module, a fusion module, and a prediction module. Obtaining the safety risk level output by the prediction model includes: Inputting the physiological index data and the geographical location data into the first encoding module, and extracting the features of the physiological index data and the geographical location data through the first encoding module and converting them into feature images; Inputting the feature image and the target image into the second encoding module to obtain a first feature vector output by the second encoding module; Inputting the voice signal data into the third encoding module to obtain a second feature vector output by the third encoding module; Inputting the first feature vector and the second feature vector into the fusion module for fusion to obtain a target feature vector; Inputting the target feature vector into the prediction module to obtain the safety risk level output by the prediction module; Inputting the feature image and the target image into the second encoding module to obtain a first feature vector output by the second encoding module includes: For the image to be processed input into the second encoding module, perform the following operations: Dividing the image to be processed into multiple image blocks, extracting initial features for each image block, and for each image block, fusing the corresponding initial features and the position information of the image block in the image to be processed to obtain block features corresponding to each image block respectively; Performing an attention mechanism on the block features corresponding to each image block respectively, and obtaining the first feature vector according to the attention mechanism result.
2. The real-time safety monitoring method for construction workers based on multi-modal data integration according to claim 1, characterized in that Obtaining the geographical location data of each worker includes: Obtaining first position data fed back by wearable devices corresponding to each worker; Obtaining second position data fed back by mobile devices corresponding to each worker; Obtaining the geographical location data of each worker according to the first position data and the second position data corresponding to each worker respectively.
3. The real-time safety monitoring method for construction workers based on multi-modal data integration according to claim 2, wherein Obtaining the geographical location data of each worker according to the first position data and the second position data corresponding to each worker respectively includes: When the difference between the first position data and the second position data corresponding to a worker is within a preset range, taking the average value of the first position data and the second position data as the geographical location data of the worker; When the difference between the first position data and the second position data corresponding to a worker exceeds the preset range, discarding the first position data and the second position data corresponding to the current worker, and re-obtaining the first position data and the second position data corresponding to the worker.
4. The real-time safety monitoring method for construction workers based on multi-modal data integration according to claim 1, characterized in that The training process of the prediction model is: Constructing an initial prediction model, and training the initial prediction model according to multiple groups of labeled data to obtain the prediction model; Among them, each group of labeled data includes sample physiological index data, sample voice signal data, sample geographical location data, sample target images, and safety risk level annotation results.
5. The real-time safety monitoring method for construction workers based on multi-modal data integration according to claim 4, characterized in that, Training the initial prediction model according to multiple groups of labeled data to obtain the prediction model includes: Selecting target labeled data; Inputting the sample physiological index data, sample voice signal data, sample geographical location data, and sample target images in the target labeled data into the initial prediction model, and obtaining the safety risk level prediction result output by the initial prediction model; Obtaining a first loss according to the safety risk level prediction result and the safety risk level annotation result in the target labeled data; Inputting the sample first feature vector and sample second feature vector obtained during the process of obtaining the safety risk level prediction result into a reconstruction module, and obtaining the reconstructed data output by the reconstruction module; Obtaining a second loss based on the difference between the sample physiological index data, the sample voice signal data, the sample geographical location data, and the sample target images and the reconstructed data; Obtaining the training loss corresponding to the target labeled data according to the first loss and the second loss, and updating the parameters of the initial prediction model according to the training loss corresponding to the target labeled data; Re-executing the step of selecting target labeled data until the parameters of the initial prediction model converge, and taking the model with converged parameters as the prediction model.
6. A real-time safety monitoring device for construction workers based on multi-modal data integration, characterized in that, The real-time safety monitoring device for construction workers based on multi-modal data integration is applied to the real-time safety monitoring method for construction workers based on multi-modal data integration according to any one of claims 1-5, and includes: A first data acquisition module, configured to acquire the physiological index data, voice signal data, and geographical location data of each worker; A second data acquisition module, configured to acquire a target image, where the target image includes multiple workers; A prediction module, configured to input the physiological index data, the voice signal data, the geographical location data, and the target image into the trained prediction model, and obtain the safety risk level output by the prediction model.
7. A terminal, characterized in that, The terminal includes: a processor and a computer-readable storage medium communicatively connected to the processor. The computer-readable storage medium is adapted to store multiple instructions, and the processor is adapted to call the instructions in the computer-readable storage medium to execute the steps of implementing the real-time safety monitoring method for construction workers based on multi-modal data integration according to any one of claims 1-5 above.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the real-time safety monitoring method for construction workers based on multi-modal data integration according to any one of claims 1-5.
Citation Information
Patent Citations
Face-based driving risk prediction method and device, equipment and storage medium
CN111414874A
Method and system for dynamically displaying safety activity elements of construction site
CN115049975A