Intelligent Construction Site Safety Management and Control Method and System Based on Multi-Source Data Analysis
Through the intelligent construction site safety control method of multi-source data analysis, combined with site information, weather information and image information, BERT and improved SlowFast model are used for feature extraction and fusion, which solves the shortcomings of traditional monitoring methods and achieves comprehensive, accurate, and real-time monitoring and management of construction site workers' behavior.
Patent Information
- Application Number
- CN202510580593.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Traditional construction site safety monitoring methods have incomplete and discontinuous monitoring, and they cannot detect and warn of safety hazards in a timely manner, resulting in low accuracy and reliability of safety monitoring.
The intelligent construction site safety control method based on multi-source data analysis is adopted, and by obtaining site information, weather information and image information, the BERT model and the improved SlowFast model are used to extract features, and feature fusion and behavior recognition are combined with Pearson's correlation coefficient and MLP network to achieve comprehensive, accurate and real-time monitoring of workers' behavior.
It improves the accuracy and reliability of construction site safety monitoring, can better capture the characteristics of workers' behavior, reduce misjudgment, and achieve comprehensive, accurate, and real-time monitoring and management of construction site workers' behavior.
Smart Images

Figure CN120086699B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent construction site management and control, and particularly relates to an intelligent construction site safety management and control method and system based on multi-source data analysis. Background Art
[0002] The internal environment of a construction site is complex and changeable. When construction workers work inside the construction site, they not only need to operate in accordance with relevant regulations, but also need to pay attention to the safety status of the area where they are located.
[0003] At present, construction site safety management and control can reduce the probability of accidents to a certain extent. However, there are many deficiencies in traditional construction site safety monitoring means. Manual inspections are difficult to cover comprehensively and are easily affected by the quality and fatigue of personnel, while the warning effect of fixed safety notices is limited. This results in incomplete and discontinuous monitoring of construction workers, there are safety supervision loopholes, potential safety hazards cannot be discovered and warned in time, and the accuracy and reliability of construction site safety monitoring are reduced. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an intelligent construction site safety management and control method and system based on multi-source data analysis, aiming to solve the technical problems proposed in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions.
[0006] An embodiment of the present invention provides an intelligent construction site safety management and control method based on multi-source data analysis. The management and control method includes the following steps:
[0007] S1. Obtain multi-source information of the construction site, where the multi-source information includes site information, weather information, and image information;
[0008] S2. Perform text preprocessing on the site information and weather information respectively, and use the BERT model to extract features from the preprocessed site information text and weather information text respectively to obtain site features and weather features, and splice and process the site features and weather features to obtain text features;
[0009] S3. Use an improved SlowFast model to extract visual features of the image information; the improved SlowFast model includes a fast branch and a slow branch. A double-layer attention module is introduced into the fast branch, and a feature enhancement module is introduced into the slow branch. The weights of the fast branch and the slow branch are adjusted based on a dynamic weight mechanism, and the outputs of the two branches are feature-fused according to the adjusted weights to obtain visual features;
[0010] S4. Calculate the Pearson correlation coefficient to construct a cross-modal relationship matrix between the text and visual features, and weighted-fuse the text features and visual features according to the cross-modal relationship matrix to obtain the final fusion features;
[0011] S5. Use the fused features as the input of the MLP network. Among them, introduce multi-stage hierarchical fusion residuals in the MLP network, extract features through residual MLP blocks in multiple stages, fuse shallow and deep features at each stage, and map the fused features to the corresponding behavior categories to obtain the preliminary behavior categories of the workers in the image.
[0012] S6. Compare the preliminary behavior category results with the site features and weather features, and adjust the preliminary behavior categories according to the predetermined correction rules to obtain the final behavior categories.
[0013] S7. Based on the final behavior categories, conduct safety control over the behaviors of the construction workers.
[0014] Furthermore, in step S3, the steps of introducing a double-layer attention module in the fast branch include:
[0015] S311. Linearly transform the input features to generate the query matrix, key matrix, and value matrix in the attention mechanism.
[0016] S312. Construct a directed graph and establish the attention relationship between different regions, calculate the average values of the query matrix and key matrix in each region to generate the regional query matrix and regional key matrix, and generate the adjacency matrix by taking the dot product of the regional query matrix and regional key matrix. The adjacency matrix is used to measure the correlation between different regions.
[0017] S313. Prune the adjacency matrix, including the adjacency matrix of the top k highly correlated regions, to obtain the routing index matrix.
[0018] S314. Based on the attention mechanism concentrated on k routing regions, aggregate the key matrix and value matrix tensors of all routing regions to generate the aggregated key matrix and value matrix.
[0019] S315. Conduct an attention operation on the aggregated key matrix and value matrix, introduce the local context enhancement term LEC to derive the result tensor, and obtain the output features of the fast branch.
[0020] Furthermore, in step S3, the feature enhancement module is a multi-branch structure, and the output feature maps of multiple branches are spliced; the spliced feature map and the input feature map are added through residual connection and output through the ReLU activation function to obtain the output features of the slow branch.
[0021] Furthermore, in the step of adjusting the weights of the fast branch and the slow branch based on the dynamic weight mechanism in step S3, the weight calculation formula is expressed as: ;
[0022] Among them, It represents the result of global average pooling of the feature sequence of the slow branch; It represents the result of global average pooling of the feature sequence of the fast branch; conv1, conv2, and conv3 all represent convolution operations; σ represents the sigmoid activation function, which is used to limit the output result within the range of 0 to 1, and the output result is used to represent the magnitude of the weight.
[0023] Further, in step S3, the step of fusing the outputs of the two branches according to the adjusted weights to obtain visual features includes:
[0024] Obtain the output features of the fast branch and the output features of the slow branch, denoted as and , where, represents the output features of the slow branch, represents the output features of the fast branch;
[0025] Perform global average pooling on the feature sequences of the slow branch and the fast branch respectively to obtain the pooling results and ;
[0026] Perform convolution operations on the two pooling results respectively and take the difference to calculate the motion feature difference F between the fast and slow branches, expressed as: ;
[0027] where, conv1 and conv2 both represent convolution operations, and σ represents the activation function;
[0028] Perform a convolution operation on the motion feature difference F and use the sigmoid activation function to generate the feature weight , expressed as: ;
[0029] where, conv3 represents the convolution operation, and σ represents the sigmoid activation function;
[0030] Perform a dot product operation on the feature weight and the feature of the slow branch to generate an enhanced feature map , where, ;
[0031] Fuse the enhanced feature map with the feature of the fast branch as the subsequent input of the slow branch to achieve the feature fusion of the fast and slow branches and obtain visual features.
[0032] Further, in step S4, the Pearson correlation coefficient is expressed as: ;
[0033] In the formula, is and covariance; is standard deviation; is standard deviation; is mean value of is mean value of ; E represents expected value.
[0034] Furthermore, in step S4, the step of constructing a cross-modal relationship matrix between text features and visual features includes:
[0035] Substitute text features and visual features into the variables of the Pearson correlation coefficient respectively to construct a relationship fusion matrix, and obtain the relationship fusion matrix of text features and visual features , where is the number of text features and visual features; the dimension of the matrix is ; the matrix reflects the mutual relationship between text features and visual features;
[0036] Substitute visual features and text features into the variables of the Pearson correlation coefficient to construct a relationship fusion matrix, and use the Pearson correlation coefficient to obtain the relationship matrix of visual features and text features .
[0037] Furthermore, in step S4, the step of weighted-fusing text features and visual features according to the cross-modal relationship matrix to obtain the final fused features includes:
[0038] After obtaining the relationship matrices and , weight the visual features and text features respectively; where:
[0039] Weight the visual features through the text-to-visual weighted relationship matrix , expressed as: , represents the visual features generated after being weighted by the text-to-visual relationship matrix;
[0040] Weight the text features through the visual-to-text weighted relationship matrix , expressed as: , represents the text features generated after being weighted by the visual-to-text weighted relationship matrix;
[0041] and are the enhanced text features and visual features respectively;
[0042] Finally, through the method of weighted average, the enhanced text features and visual features are fused together, expressed as: ; where is the final fused feature vector, and are hyperparameters that control the contribution degrees of text features and visual features in the final fused feature.
[0043] Furthermore, in step S5, multi-stage hierarchical fusion residuals are introduced into the MLP network, including:
[0044] The MLP network is divided into multiple stages, each stage contains several residual MLP blocks, and the structure of each residual MLP block is expressed as: ;
[0045] where represents the output of the th layer, and MLP(·) represents a multi-layer perceptron;
[0046] At the end of each stage, the output features of the current stage are fused with the output features of the previous stage, expressed as: ;
[0047] where represents the fused feature of the s-th stage, and concat(·, ·) represents the feature concatenation operation;
[0048] Skip connections are added between different stages of the network to fuse shallow features and deep features, expressed as: ;
[0049] where represents the output of the j-th skip connection, represents the transformation of features;
[0050] The fused feature of the last stage is input into the classification layer, expressed as: ;
[0051] where O represents the final classification output, and W and b represent the weights and biases of the classification layer respectively.
[0052] Another embodiment of the present invention provides a smart construction site safety management and control system based on multi-source data analysis, and this management and control system includes the following modules:
[0053] A data acquisition module, which is used to acquire multi-source information of the construction site, and the multi-source information includes site information, weather information, and image information;
[0054] The text feature extraction module is used to perform text preprocessing on the site information and weather information respectively, and use the BERT model to extract features from the preprocessed site information text and weather information text respectively to obtain site features and weather features, and splice and process the site features and weather features to obtain text features;
[0055] The visual feature extraction module is used to extract the visual features of the image information by using an improved SlowFast model; the improved SlowFast model includes a fast branch and a slow branch. A double-layer attention module is introduced into the fast branch, and a feature enhancement module is introduced into the slow branch. The weights of the fast branch and the slow branch are adjusted based on a dynamic weight mechanism, and the outputs of the two branches are feature-fused according to the adjusted weights to obtain visual features;
[0056] The feature fusion module is used to construct a cross-modal relationship matrix between the text and visual features by calculating the Pearson correlation coefficient, and weighted-fuse the text features and visual features according to the cross-modal relationship matrix to obtain the final fused features;
[0057] The behavior classification module is used to use the fused features as the input of the MLP network. Among them, multi-stage hierarchical fusion residuals are introduced into the MLP network, features are extracted through residual MLP blocks in multiple stages, and shallow and deep features are fused at each stage, and the fused features are mapped to the corresponding behavior categories to obtain the preliminary behavior categories of the workers in the image;
[0058] The behavior correction module is used to compare the preliminary behavior category results with the site features and weather features, and adjust the preliminary behavior categories according to the predetermined correction rules to obtain the final behavior categories;
[0059] The behavior control module is used to perform safety control on the behaviors of the construction workers based on the final behavior categories.
[0060] Compared with the prior art, the beneficial effects of the intelligent construction site safety control method and system based on multi-source data analysis of the present invention are:
[0061] First, by integrating site information, weather information, and image information, the present invention can comprehensively understand the actual situation of the construction site. The BERT model is used to extract text features, enabling subsequent behavior recognition to consider the influence of the site and weather. By introducing a double attention module in the fast branch of the improved SlowFast model, it can better capture the motion information in the image sequence and improve the recognition accuracy of workers' behaviors. By introducing a feature enhancement module in the slow branch, it can enhance the ability to extract key features of workers using safety equipment. Based on the dynamic weight mechanism to adjust the weights of the fast and slow branches, the model can adaptively balance the contributions of the two branches according to different inputs, thereby better integrating motion information and appearance information and improving the model's ability to capture different behavior features.
[0062] Second, by calculating the Pearson correlation coefficient to construct a cross-modal relationship matrix, the present invention can quantify the correlation between text features and visual features, providing a basis for cross-modal feature fusion. By weighted fusing text features and visual features according to the relationship matrix, deep fusion of multi-modal information can be achieved, making full use of site and weather information to assist behavior recognition and improving the accuracy and robustness of behavior recognition.
[0063] Third, taking the fused features as the input of the MLP network, through the improved MLP network, using the multi-stage hierarchical fusion residual MLP, the adaptability of the network model to different feature patterns is enhanced through feature fusion and residual connection, realizing the preliminary classification of workers' behaviors. Comparing the preliminary behavior category results with site features and weather features and adjusting according to the predetermined correction rules can further improve the accuracy of behavior recognition. Site features and weather features provide context information for the occurrence of behaviors, helping to verify and correct the preliminary classification results and reducing misjudgments.
[0064] In summary, the control method of the present invention can realize comprehensive, accurate, and real-time monitoring and management of the behaviors of construction site workers through the acquisition and fusion of multi-source information and the combination of safety control measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.
[0066] Figure 1 It is the implementation flowchart of the intelligent construction site safety control method based on multi-source data analysis of the present invention;
[0067] Figure 2 It is a sub-flowchart of the intelligent construction site safety control method based on multi-source data analysis of the present invention;
[0068] Figure 3 This is another sub - flow chart of the intelligent construction site safety control method based on multi - source data analysis of the present invention;
[0069] Figure 4 This is yet another sub - flow chart of the intelligent construction site safety control method based on multi - source data analysis of the present invention;
[0070] Figure 5 This is the structural block diagram of the intelligent construction site safety control system based on multi - source data analysis of the present invention;
[0071] Figure 6 This is the structural block diagram of a computer device provided by the present invention. Detailed implementation manners
[0072] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0073] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.
[0074] Please refer to Figure 1 , in an embodiment of the present invention, an intelligent construction site safety control method based on multi - source data analysis is provided. The control method includes the following steps:
[0075] S1. Obtain multi - source information of the construction site. The multi - source information includes site information, weather information and image information;
[0076] By integrating site information, weather information and image information, the present invention can understand the actual situation of the construction site more comprehensively. Among them, the site information mainly includes the layout of the construction site, including the specific positions and shapes of buildings, construction areas, material stacking areas and passages, etc. Through the site information, the functions of different areas can be clearly divided, such as high - altitude operation areas, foundation pit operation areas, welding operation areas and ordinary construction areas, etc., and at the same time, the safety levels and special requirements of each area are marked. The collection of weather information is achieved through two main channels: one is to call a high - precision weather forecast API to obtain weather prediction data for the next few hours to several days in the area where the construction site is located, including meteorological elements such as temperature, humidity, rainfall, wind force, wind direction and air pressure; the other is to install reliable weather monitoring equipment, such as a weather station, at the construction site to real - time monitor the actual weather conditions at the construction site.
[0077] Furthermore, image information is a key data source for behavior recognition. High-definition and high-frame-rate cameras are installed in various key areas and key parts of the construction site to ensure a reasonable layout of the cameras, which can comprehensively cover the main operation areas, passages, and dangerous areas of the construction site and avoid monitoring blind spots.
[0078] S2. Perform text preprocessing on the site information and weather information respectively, and use the BERT model to extract features from the preprocessed site information text and weather information text respectively to obtain site features and weather features, and perform splicing processing on the site features and weather features to obtain text features;
[0079] The present invention uses the BERT model to extract text features, enabling subsequent behavior recognition to take into account the influence of the site and weather. First, text preprocessing is performed on the site information and weather information respectively to ensure the standardization and consistency of the data. The preprocessing process includes operations such as removing noise in the text, unifying the text format, and word segmentation. Then, the BERT model is used to extract features from the preprocessed site information text and weather information text respectively. The BERT model captures rich semantic information and context relationships from the text. Specifically, the preprocessed site information text and weather information text are respectively input into the pre-trained BERT model, and the BERT model will encode each word or sentence according to its internal multi-layer bidirectional Transformer structure to generate corresponding feature vectors, thereby obtaining site features and weather features. Finally, the site features and weather features are spliced to form a comprehensive text feature vector.
[0080] S3. Use an improved SlowFast model to extract visual features of the image information. The improved SlowFast model includes a fast branch and a slow branch. A double-layer attention module is introduced into the fast branch, and a feature enhancement module is introduced into the slow branch. The weights of the fast branch and the slow branch are adjusted based on the dynamic weight mechanism, and the outputs of the two branches are feature-fused according to the adjusted weights to obtain visual features;
[0081] The present invention combines the improved SlowFast model. By introducing a double-layer attention module into the fast branch, it can better capture the motion information in the image sequence and improve the recognition accuracy of workers' behaviors. By introducing a feature enhancement module into the slow branch, it can enhance the ability to extract key features of workers using safety equipment. The present invention adjusts the weights of the fast branch and the slow branch based on the dynamic weight mechanism, enabling the model to adaptively balance the contributions of the two branches according to different inputs, thereby better fusing motion information and appearance information and improving the model's ability to capture different behavior features.
[0082] S4. By calculating the Pearson correlation coefficient, construct a cross-modal relationship matrix between text and visual features, and weighted-fuse the text features and visual features according to the cross-modal relationship matrix to obtain the final fused features;
[0083] Through calculating the Pearson correlation coefficient to construct a cross-modal relationship matrix, the present invention can quantify the correlation between text features and visual features, provide a basis for cross-modal feature fusion. By weighted-fusing the text features and visual features according to the relationship matrix, deep fusion of multi-modal information can be realized, making full use of site and weather information to assist in behavior recognition, and improving the accuracy and robustness of behavior recognition;
[0084] S5. Use the fused features as the input of the MLP network. Among them, introduce multi-stage hierarchical fusion residuals into the MLP network, extract features through multiple stages of residual MLP blocks, and fuse shallow and deep features at each stage, and map the fused features to the corresponding behavior categories to obtain the preliminary behavior categories of the workers in the image;
[0085] S6. Compare the preliminary behavior category results with the site features and weather features, and adjust the preliminary behavior categories according to the predetermined correction rules to obtain the final behavior categories;
[0086] S7. Based on the final behavior categories, conduct safety control over the behaviors of construction site workers;
[0087] The present invention uses the fused features as the input of the MLP network. Through the improved MLP network, by using multi-stage hierarchical fusion residuals MLP, through feature fusion and residual connection, the adaptability of the network model to different feature patterns is enhanced, and preliminary classification of workers' behaviors is realized;
[0088] Furthermore, comparing the preliminary behavior category results with the site features and weather features and making adjustments according to the predetermined correction rules can further improve the accuracy of behavior recognition;
[0089] The site features and weather features provide context information for the occurrence of behaviors, which helps to verify and correct the preliminary classification results and reduce misjudgments.
[0090] In the embodiments of the present invention, by combining the position of the worker with the site information, when controlling the behavior of the worker, the specific area where the worker is currently located can be determined. For example, whether the worker is in the high-altitude operation area, near the foundation pit, in the tower crane operation area or the ordinary construction passage, etc. Different areas correspond to different types of possible behaviors and different safety requirements, which provides key context information for spatial perception and behavior positioning; in addition, according to the functional division of the site, the behavior of the worker in certain areas can be reasonably predicted. For example, in the material stacking area, it is normal for the worker to perform the behavior of carrying materials, but if the worker performs behaviors related to material processing such as welding or cutting in this area, there may be certain potential safety hazards.
[0091] Therefore, in the embodiments of the present invention, after the site information is fused with multi-source information such as weather information and image information, more comprehensive context can be provided. For example, in a high-altitude operation area with strong wind weather, when the worker shows behaviors such as body tilt, combining the site information (high-altitude operation area) and weather information (strong wind), it can be more accurately judged that this is a normal reaction of the worker to maintain balance in the strong wind environment, rather than a dangerous and illegal behavior. This multi-source information fusion enables behavior recognition to fully consider site factors, thereby improving the accuracy of recognition.
[0092] In step S6 of the present invention, by comparing the preliminary behavior category result with the site characteristics and weather characteristics, based on the comparison result, the behavior category result is revised; here it is considered that the construction site is a complex environment with various uncertainties; the weather may also change suddenly, and the behavior pattern of the worker will change greatly, such as the walking speed slows down and the operation actions may be more cautious, etc.; therefore, the subsequent behavior correction step of the present invention can further correct the behavior recognition result by considering more factors such as the trend and change speed of the weather change.
[0093] Please refer to Figure 2 , in step S3 of the embodiments of the present disclosure, the step of introducing a double attention module into the fast branch includes:
[0094] S311. Linearly transform the input features to generate the query matrix, key matrix and value matrix in the attention mechanism;
[0095] S312. Construct a directed graph and establish the attention relationship between different regions, calculate the average values of the query matrix and the key matrix in each region to generate the region query matrix and the region key matrix, and dot-product the region query matrix and the region key matrix to generate the adjacency matrix, and the adjacency matrix is used to measure the correlation between different regions;
[0096] S313. Prune the adjacency matrix, including the adjacency matrix of the top k regions with high correlation, to obtain the routing index matrix;
[0097] S314. Aggregate the key matrix and value matrix tensors of all routing regions based on an attention mechanism concentrated on k routing regions to generate an aggregated key matrix and value matrix;
[0098] S315. Perform an attention operation on the aggregated key matrix and value matrix, introduce a local context enhancement term LEC to derive a result tensor, and obtain the output features of the fast branch.
[0099] Further, in step S3, the feature enhancement module is a multi-branch structure, and the output feature maps of multiple branches are concatenated; the concatenated feature map and the input feature map are added through a residual connection and output through a ReLU activation function to obtain the output features of the slow branch; in one implementation of the multi-branch structure, a three-branch structure is adopted, specifically including:
[0100] The first branch with two convolutional layers. One convolutional layer uses a 1×1 convolutional kernel with a stride of stride to adjust the number of channels of the input feature map to 2×inter_planes of the original; the other convolutional layer uses a 3×3 convolutional kernel with a stride of 1 and a padding of 1 to extract local features. The first branch can increase the number of channels of the features and extract richer local features without changing the size of the feature map.
[0101] The second branch with four convolutional layers. One convolutional layer uses a 1×1 convolutional kernel with a stride of 1 to adjust the number of channels of the input feature map to inter_planes; another convolutional layer uses a 1×3 convolutional kernel with a stride of stride and a padding of (0,1) to expand the receptive field in the height direction; the last convolutional layer uses a 3×1 convolutional kernel with a stride of stride and a padding of (1,0) to expand the receptive field in the width direction; the fourth convolutional layer uses a 3×3 convolutional kernel with a stride of 1 and a padding of 5 and a dilation rate of 5 to further extract dilated convolutional features, expand the receptive field and capture more context information.
[0102] The third branch with three convolutional layers. One convolutional layer uses a 1×1 convolutional kernel with a stride of stride to adjust the number of channels of the input feature map to 2×inter_planes of the original; another convolutional layer uses a 3×1 convolutional kernel with a stride of stride and a padding of (1,0), and the last convolutional layer uses a 1×3 convolutional kernel with a stride of stride and a padding of (0,1) to extract features from different convolutional kernel dimension orders and enrich the diversity of features.
[0103] Further, the output feature maps of the three branches are concatenated; the number of channels is adjusted to out_planes through a convolutional layer, and finally added to the input feature map through a residual connection and output through a ReLU activation function to obtain the output feature of the slow branch.
[0104] In the embodiment of the present invention, in the step of adjusting the weights of the fast branch and the slow branch based on the dynamic weight mechanism provided in step S3, the weight calculation formula is expressed as: ; where represents the result of global average pooling of the feature sequence of the slow branch; represents the result of global average pooling of the feature sequence of the fast branch; conv1, conv2, and conv3 all represent convolutional operations; σ represents the sigmoid activation function, which is used to limit the output result within the range of 0 to 1, and the output result is used to represent the size of the weight.
[0105] As Figure 3 shown, in step S3, the step of performing feature fusion on the outputs of the two branches according to the adjusted weights to obtain visual features includes:
[0106] S321. Obtain the output feature of the fast branch and the output feature of the slow branch, and denote them as and , where represents the output feature of the slow branch, represents the output feature of the fast branch;
[0107] S323. Perform global average pooling on the feature sequences of the slow branch and the fast branch respectively to obtain the pooling results and ;
[0108] S323. Perform convolutional operations on the two pooling results respectively and take the difference to calculate the motion feature difference F between the fast and slow branches, which is expressed as: ; where conv1 and conv2 both represent convolutional operations, and σ represents the activation function;
[0109] S324. Perform a convolutional operation on the motion feature difference F and use the sigmoid activation function to generate the feature weight , which is expressed as: ; where conv3 represents the convolutional operation, and σ represents the sigmoid activation function;
[0110] S325. Perform a dot product operation on the feature weight and the feature of the slow branch to generate the enhanced feature map , where ;
[0111] S326. Fuse the enhanced feature map with the features of the fast branch as the subsequent input of the slow branch to achieve the feature fusion of the fast and slow branches and obtain visual features.
[0112] Furthermore, in step S4, the Pearson correlation coefficient is expressed as: ;
[0113] In the formula, is and 's covariance; is 's standard deviation; is 's standard deviation; is 's mean value, is 's mean value; E represents the expected value.
[0114] Furthermore, in step S4, the steps of constructing the cross-modal relationship matrix between the text and visual features include: substituting the text features and visual features into the variables of the Pearson correlation coefficient respectively to construct the relationship fusion matrix and obtain the text feature and visual feature relationship fusion matrix , where 's value is the number of text features and visual features; the dimension of the matrix is ; the matrix reflects the mutual relationship between the text features and visual features; substituting the visual features and text features into the variables of the Pearson correlation coefficient to construct the relationship fusion matrix, and using the Pearson correlation coefficient, the visual feature and text feature relationship matrix can be obtained.
[0115] In step S4 of the present invention, the steps of weighted fusing the text features and visual features according to the cross-modal relationship matrix to obtain the final fused features include: after obtaining the relationship matrix and the relationship matrix , weighting the visual features and text features respectively; among them: weighting the visual features through the text-to-visual weighted relationship matrix , which is expressed as: , represents the visual features generated after being weighted by the text-to-visual relationship matrix; weighting the text features through the visual-to-text weighted relationship matrix , which is expressed as: , represents the text features generated after being weighted by the visual-to-text weighted relationship matrix; among them, and They are the enhanced text features and visual features respectively; finally, through the method of weighted average, the enhanced text features and visual features are fused together, expressed as: ; where is the final fused feature vector, and are hyperparameters that control the contribution degrees of text features and visual features in the final fused features.
[0116] Furthermore, please refer to Figure 4 In step S5, multi-stage hierarchical fusion residuals are introduced into the MLP network, including:
[0117] S411. Divide the MLP network into multiple stages, each stage contains several residual MLP blocks, and the structure of each residual MLP block is expressed as: ; where represents the output of the layer, and MLP(·) represents a multi-layer perceptron;
[0118] S412. At the end of each stage, fuse the output features of the current stage with the output features of the previous stage, expressed as: ; where represents the fused feature of the s-th stage, and concat(·, ·) represents a feature concatenation operation;
[0119] S413. Add skip connections between different stages of the network to fuse shallow features and deep features, expressed as: ; where represents the output of the j-th skip connection, represents the transformation of features;
[0120] S414. Input the fused feature of the last stage into the classification layer, expressed as: ; where O represents the final classification output, and W and b represent the weight and bias of the classification layer respectively.
[0121] In summary, the control method of the present invention can achieve comprehensive, accurate, and real-time monitoring and management of the behaviors of construction workers through the acquisition and fusion of multi-source information and the combination of safety control measures.
[0122] Please refer to Figure 5 In another embodiment of the present invention, a smart construction site safety control system based on multi-source data analysis is provided. The control system includes the following modules:
[0123] A data acquisition module 81, which is used to acquire multi-source information of the construction site, and the multi-source information includes site information, weather information, and image information;
[0124] The text feature extraction module 82 is used to perform text preprocessing on the site information and weather information respectively, and use the BERT model to extract features from the preprocessed site information text and weather information text respectively to obtain site features and weather features, and splice and process the site features and weather features to obtain text features;
[0125] The visual feature extraction module 83 is used to extract the visual features of the image information by using an improved SlowFast model; the improved SlowFast model includes a fast branch and a slow branch. A double-layer attention module is introduced into the fast branch, and a feature enhancement module is introduced into the slow branch. The weights of the fast branch and the slow branch are adjusted based on the dynamic weight mechanism, and the outputs of the two branches are feature-fused according to the adjusted weights to obtain visual features;
[0126] The feature fusion module 84 is used to calculate the Pearson correlation coefficient to construct a cross-modal relationship matrix between the text and visual features, and weighted-fuse the text features and visual features according to the cross-modal relationship matrix to obtain the final fused features;
[0127] The behavior classification module 85 is used to take the fused features as the input of the MLP network. Among them, multi-stage hierarchical fusion residuals are introduced into the MLP network, features are extracted through multiple-stage residual MLP blocks, and shallow and deep features are fused at each stage, and the fused features are mapped to the corresponding behavior categories to obtain the preliminary behavior categories of the workers in the image;
[0128] The behavior correction module 86 is used to compare the preliminary behavior category results with the site features and weather features, and adjust the preliminary behavior categories according to the predetermined correction rules to obtain the final behavior categories;
[0129] The behavior control module 87 is used to perform safety control on the behaviors of the construction site workers based on the final behavior categories.
[0130] As Figure 6 shown, the computer device includes a processor, a memory, a network interface, an input device, and a display screen connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and can also store a computer program. When the computer program is executed by the processor, the processor can implement the intelligent construction site safety control method based on multi-source data analysis.
[0131] The internal memory can also store a computer program. When the computer program is executed by the processor, the processor can execute the intelligent construction site safety control method based on multi-source data analysis.
[0132] Those skilled in the art can understand, Figure 6The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0133] In one embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the processor is caused to execute the intelligent construction site safety control method based on multi-source data analysis provided in the above embodiment.
[0134] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to the memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0136] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0137] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
[0138] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A smart construction site safety control method based on multi-source data analysis, characterized in that, The control method includes the following steps: S1. Obtain multi-source information of the construction site, where the multi-source information includes site information, weather information, and image information; S2. Perform text preprocessing on the site information and weather information respectively, and use the BERT model to extract features from the preprocessed site information text and weather information text respectively to obtain site features and weather features. Concatenate the site features and weather features to obtain text features; S3. Use an improved SlowFast model to extract visual features of the image information; the improved SlowFast model includes a fast branch and a slow branch. A double-layer attention module is introduced into the fast branch, and a feature enhancement module is introduced into the slow branch. The feature enhancement module is a multi-branch structure, and the output feature maps of multiple branches are concatenated; the concatenated feature map and the input feature map are added through residual connection, and the output of the slow branch is obtained after passing through the ReLU activation function. Adjust the weights of the fast branch and the slow branch based on the dynamic weight mechanism, and fuse the outputs of the two branches according to the adjusted weights to obtain visual features; S4. Calculate the Pearson correlation coefficient to construct a cross-modal relationship matrix between the text and visual features, and weighted-fuse the text features and visual features according to the cross-modal relationship matrix to obtain the final fused features; S5. Use the fused features as the input of the MLP network. Among them, multi-stage hierarchical fusion residuals are introduced into the MLP network, features are extracted through residual MLP blocks in multiple stages, and shallow and deep features are fused in each stage. Map the fused features to the corresponding behavior categories to obtain the preliminary behavior categories of the workers in the image; In the step of introducing multi-stage hierarchical fusion residuals into the MLP network, the MLP network is divided into multiple stages, each stage contains several residual MLP blocks. At the end of each stage, the output features of the current stage are fused with the output features of the previous stage, and skip connections are added between different stages of the network to fuse shallow features and deep features. The fused features of the last stage are input into the classification layer; S6. Compare the preliminary behavior category results with the site features and weather features, and adjust the preliminary behavior category according to the predetermined correction rules to obtain the final behavior category; S7. Based on the final behavior category, perform safety control on the behaviors of the construction site workers.
2. The intelligent construction site safety control method based on multi-source data analysis according to claim 1, wherein, In step S3, the step of introducing a double-layer attention module into the fast branch includes: S311. Perform a linear transformation on the input features to generate a query matrix, a key matrix, and a value matrix in the attention mechanism; S312. Construct a directed graph and establish an attention relationship between different regions. Calculate the average values of the query matrix and the key matrix in each region to generate a region query matrix and a region key matrix. Dot-product the region query matrix and the region key matrix to generate an adjacency matrix, and the adjacency matrix is used to measure the correlation between different regions; S313. Perform pruning processing on the adjacency matrix, including the adjacency matrix of the top k highly correlated regions, to obtain a routing index matrix; S314. Aggregate the key matrix and value matrix tensors of all routing regions based on an attention mechanism concentrated on k routing regions to generate an aggregated key matrix and value matrix; S315. Perform an attention operation on the aggregated key matrix and value matrix, introduce a local context enhancement term LEC to derive a result tensor, and obtain the output features of the fast branch.
3. The intelligent construction site safety control method based on multi-source data analysis according to claim 2, wherein In step S3, in the step of adjusting the weights of the fast branch and the slow branch based on a dynamic weight mechanism, the weight calculation formula is expressed as: ; Among them, represents the result of global average pooling of the feature sequence of the slow branch; represents the result of global average pooling of the feature sequence of the fast branch; conv1, conv2, and conv3 all represent convolution operations; σ represents the sigmoid activation function, which is used to limit the output result within the range of 0 to 1, and the output result is used to represent the magnitude of the weight.
4. The intelligent construction site safety control method based on multi-source data analysis according to claim 3, characterized in that, In step S3, the step of performing feature fusion on the outputs of the two branches according to the adjusted weights to obtain visual features includes: Obtain the output features of the fast branch and the output features of the slow branch, denoted as and , where represents the output features of the slow branch, represents the output features of the fast branch; Perform global average pooling on the feature sequences of the slow branch and the fast branch respectively to obtain the pooling results and ; Perform convolution operations on the two pooling results respectively and take the difference to calculate the motion feature difference F between the fast and slow branches, expressed as: ; where conv1 and conv2 both represent convolution operations, and σ represents the activation function; Perform a convolution operation on the motion feature difference F and use the sigmoid activation function to generate feature weights , which is expressed as: ; where conv3 represents the convolution operation and σ represents the sigmoid activation function; Multiply the feature weights with the features of the slow branch to perform a dot product operation to generate an enhanced feature map , where ; Fuse the enhanced feature map with the features of the fast branch as the subsequent input of the slow branch, to achieve the feature fusion of the fast and slow branches and obtain visual features.
5. The intelligent construction site safety control method based on multi-source data analysis according to claim 4, wherein, In step S4, the Pearson correlation coefficient is expressed as: ; In the formula, is and 's covariance; is 's standard deviation; is 's standard deviation; is 's mean value, is 's mean value; E represents the expected value.
6. The intelligent construction site safety control method based on multi-source data analysis according to claim 5, wherein, In step S4, the step of constructing a cross-modal relationship matrix between text and visual features includes: Substitute the text features and visual features into the variables of the Pearson correlation coefficient respectively to construct a relationship fusion matrix, and obtain the relationship fusion matrix of text features and visual features , where The value of is the number of text features and visual features; the dimension of the matrix is ; the matrix reflects the mutual relationship between text features and visual features; Substitute the visual features and text features into the variables of the Pearson correlation coefficient to construct a relationship fusion matrix. Using the Pearson correlation coefficient, a relationship matrix between visual features and text features can be obtained. 。 7. The intelligent construction site safety control method based on multi-source data analysis according to claim 6, characterized in that In step S4, the step of weighted fusing text features and visual features according to the cross-modal relationship matrix to obtain the final fused features includes: After obtaining the relationship matrix and the relationship matrix weight the visual features and text features respectively; where: The weighted relationship matrix from text to vision is used to weight the visual features, expressed as: , represents the visual features generated after being weighted by the relationship matrix from text to vision; Weighted relationship matrix from vision to text to weight text features, expressed as: , represents the text features generated after being weighted by the weighted relationship matrix from vision to text; and are text features and visual features respectively; Finally, through the method of weighted average, the weighted text features and visual features are fused together, expressed as: ; where is the final fusion feature vector, and are hyperparameters that control the contribution degrees of text features and visual features in the final fusion feature.
8. The intelligent construction site safety control method based on multi-source data analysis according to claim 7, characterized in that In step S5, introducing multi-stage hierarchical fusion residuals into the MLP network includes: The structure of each residual MLP block is expressed as: ; Among them, represents the output of the layer, represents a multi-layer perceptron; At the end of each stage, the output features of the current stage are fused with the output features of the previous stage, expressed as: ; where represents the fused features of the s-th stage, represents the feature concatenation operation; Adding skip connections between different stages of the network to fuse shallow features and deep features, which is expressed as: ; where represents the output of the j-th skip connection, represents transforming the features; The fusion features of the last stage are input into the classification layer, expressed as: ; where O represents the final classification output, and W and b represent the weights and biases of the classification layer, respectively.
9. A control system for implementing the intelligent construction site safety control method based on multi-source data analysis according to any one of claims 1 to 8, characterized in that, The control system includes the following modules: A data acquisition module for acquiring multi-source information of the construction site, where the multi-source information includes site information, weather information, and image information; A text feature extraction module for respectively performing text preprocessing on the site information and weather information, and using a BERT model to respectively extract features from the preprocessed site information text and weather information text to obtain site features and weather features, and performing splicing processing on the site features and weather features to obtain text features; A visual feature extraction module for extracting visual features of image information using an improved SlowFast model; The improved SlowFast model includes a fast branch and a slow branch. A double-layer attention module is introduced into the fast branch, and a feature enhancement module is introduced into the slow branch. The weights of the fast branch and the slow branch are adjusted based on a dynamic weight mechanism, and the outputs of the two branches are fused according to the adjusted weights to obtain visual features; A feature fusion module for constructing a cross-modal relationship matrix between text and visual features by calculating the Pearson correlation coefficient, and weighted fusing text features and visual features according to the cross-modal relationship matrix to obtain the final fused features; A behavior classification module for using the fused features as the input of the MLP network. Among them, multi-stage hierarchical fusion residuals are introduced into the MLP network, features are extracted through residual MLP blocks in multiple stages, and shallow and deep features are fused at each stage, and the fused features are mapped to the corresponding behavior categories to obtain the preliminary behavior categories of the workers in the image; A behavior correction module for comparing the preliminary behavior category results with the site features and weather features, and adjusting the preliminary behavior categories according to a predetermined correction rule to obtain the final behavior categories; A behavior control module for performing safety control on the behaviors of construction site workers based on the final behavior categories.
Citation Information
Patent Citations
Behavior recognition method and system based on improved Slowfast
CN116844236A
Agricultural product classification method fusing double-flow attention integration and cross-modal fusion
CN119919932A