A Method for Dynamically Perceiving Regional Targets and Environments for Construction Site Video Surveillance
Through the distributed camera deployment and multi-modal data fusion technology of the construction site video surveillance system, the problems of construction site monitoring coverage, data fusion accuracy, target detection accuracy and risk assessment efficiency are solved, and the intelligent upgrade of construction site video surveillance is achieved, improving the real-time and stability of the system.
Patent Information
- Application Number
- CN202510000504.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The existing construction site video surveillance methods have shortcomings in monitoring coverage, data fusion accuracy, target detection accuracy, risk assessment efficiency and computing resource utilization efficiency, and cannot meet the needs of intelligent management of modern construction sites.
By optimizing the deployment of distributed cameras on the construction site, combining the Kafka Broker cluster multi-threaded video data and multi-type environmental sensor data, using Mahjong distance for data fusion, and multi-modal data analysis through sensor neural network and RNN construction site situation awareness model, real-time risk assessment and alarm of personnel behavior and environmental conditions within the construction site are achieved.
It improves the coverage rate of construction site monitoring, realizes the dynamic integration of environmental multimodal data, improves the accuracy of target detection and behavioral analysis, realizes accurate analysis of comprehensive situations and dynamic risk assessment, ensures the real-time and stability of the system, and supports the application and expansion of complex construction site scenarios.
Smart Images

Figure CN119397206B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent monitoring, and specifically relates to a method for dynamically perceiving regional targets and environments for construction site video monitoring. Background Art
[0002] With the continuous advancement of the urbanization process and the rapid development of infrastructure construction, the number and scale of construction sites have increased year by year. As a dynamic and high-risk working environment, construction sites have frequent personnel flow, complex operation tasks, and changing environmental factors, thus facing many safety hazards and management problems. To improve the safety and management efficiency of construction site operations, traditional video monitoring methods have been widely used, but due to their technical limitations, they have not been able to meet the actual needs of modern construction site intelligent management.
[0003] Existing construction site video monitoring methods mainly rely on fixed-angle cameras for area coverage and analyze the monitoring images through manual or simple algorithms. However, these methods have the following deficiencies:
[0004] (1) Insufficient optimization and deployment of the monitoring area
[0005] Traditional camera deployment methods are usually based on experience or simple geometric analysis, and it is difficult to reasonably optimize complex construction site scenarios. For example, there may be blind spots in the monitoring area, or key areas cannot be effectively covered due to building occlusion, thus reducing the comprehensiveness and reliability of the monitoring system.
[0006] (2) Poor ability to fuse multi-modal data
[0007] The construction site environment not only includes personnel activities, but also involves many environmental parameters (such as temperature and humidity, gas concentration, light intensity, etc.), and this information is crucial for situation awareness and risk assessment. However, traditional monitoring methods lack the ability to perform real-time fusion analysis of multiple types of data and cannot achieve dynamic association between video data and environmental parameters, thus limiting the comprehensive perception of the overall situation.
[0008] (3) Insufficient accuracy in target detection and behavior analysis
[0009] In construction site video monitoring, accurately detecting and identifying multiple targets (such as personnel) and their behavior characteristics is the basis for situation awareness. However, due to the complex background of the construction site scene, severe occlusion of multiple targets, and rapid changes in behavior, traditional target detection algorithms have significant deficiencies in accuracy and robustness and are difficult to meet actual needs.
[0010] (4) Lack of dynamic risk assessment ability
[0011] The main purpose of the construction site monitoring system is to identify and warn of potential safety risks. However, existing methods often rely on a single data source or simple rules when evaluating dynamic risk events, unable to achieve a comprehensive analysis of personnel behavior and environmental conditions, which easily leads to missed alarms or false alarms.
[0012] (5) Insufficient computing resource utilization efficiency
[0013] In large-scale construction site monitoring scenarios, the volume of video data is huge and sensor information is complex. Traditional systems have problems such as unreasonable computing resource allocation and low multi-threaded concurrency efficiency during the processing, resulting in limitations in real-time performance and response speed.
[0014] In response to the above problems, there is an urgent need for an intelligent method for construction site video monitoring. Summary of the Invention
[0015] To overcome the deficiencies of the prior art, the present invention provides a method for dynamically perceiving regional targets and the environment for construction site video monitoring, which is used to solve the deficiencies in monitoring coverage, data fusion accuracy, target detection accuracy, risk assessment efficiency, and resource utilization efficiency of existing construction site video monitoring methods, thereby significantly improving the intelligent level of construction site video monitoring.
[0016] To solve the above problems, the technical solutions adopted by the present invention are as follows:
[0017] A method for dynamically perceiving regional targets and the environment for construction site video monitoring, comprising the following steps:
[0018] Perform an initial deployment of distributed cameras on the construction site, and optimize the deployment by obtaining unit discretization segmentation, distances and angles between units;
[0019] Based on the optimized distributed cameras, collect video data through a Kafka Broker cluster in a multi-threaded manner. At the same time, collect environmental parameters through multi-type environmental sensors and fuse them based on Mahalanobis distance to obtain fused environmental parameters;
[0020] Perform multi-target detection on the collected video data and analyze its behavioral characteristics to obtain video target data;
[0021] Extract environmental parameter feature representations from the fused environmental parameters based on a sensor neural network, and perform multi-modal dynamic fusion with the video target data based on learnable parameters. Analyze the comprehensive situation through an RNN construction site situation awareness model to obtain real-time comprehensive situation data;
[0022] Based on the real-time comprehensive situation data, use Bayesian network analysis to perform real-time risk assessment on the personnel behavior and environmental conditions within the construction site, identify risk events, and automatically generate warning information;
[0023] Among them, when performing multi-object detection, it includes: obtaining the target positions of key frames by constructing a construction site target detector; capturing spatial semantic information and dynamic information through different paths by constructing a construction site feature extraction model to generate a fused feature map, and completing behavior classification in combination with the target positions of key frames.
[0024] As a preferred embodiment of the present invention, when performing deployment optimization, it includes:
[0025] Performing unit discretization segmentation on the construction site to obtain a number of site units;
[0026] Obtaining the site units where the distributed cameras are set and obtaining the site units corresponding to the directions of the distributed cameras ;
[0027] Obtaining the site units from the key areas that must be video-monitored and obtaining the site units according to the distances between the site units and the included angle between the site units ;
[0028] If the included angle is less than a preset value, it is considered that the distributed camera can monitor and cover the site unit ;
[0029] By obtaining the included angle between each site unit and the site unit corresponding to the direction of the distributed camera , the monitoring coverage area of the distributed camera is determined.
[0030] As a preferred embodiment of the present invention, after determining the monitoring coverage area of the distributed camera, it includes:
[0031] Determining the first edge site unit and the second edge site unit of the building that blocks the monitoring ; Obtaining the angle formed by the first edge site unit
[0032] , the second edge site unit and the site unit ; Determining a site unit in the monitoring coverage area, and determining the judgment parameter according to the angles formed between it and the first edge site unit
[0033] , the second edge site unit , the site unit , the distances between the site units and the angle ; to determine whether the site unit is blocked by a building by taking the value of ;
[0034] According to the monitoring coverage area and the judgment parameter , determine the site units in the key area that cannot be covered by monitoring or are blocked by buildings, and then adjust the distributed cameras so that all site units in the key area can be covered by monitoring and are not blocked by buildings.
[0035] As a preferred embodiment of the present invention, when collecting video data, it includes:
[0036] Receive video data for parsing and encapsulation, serialize it, and then send the video message data to the Kafka message queue for storage;
[0037] Create multiple receivers with the same groupId for the Kafka data in the same Topic, pull the Kafka data in different Partition partitions based on the thread pool, store it in the form of data blocks, and aggregate the Kafka data pulled by multiple receivers;
[0038] After deserializing the aggregated Kafka data, convert the obtained Json message into Json data format, and parse and convert the Json data into Mat data type that can be used for image processing.
[0039] As a preferred embodiment of the present invention, when sending to the Kafka message queue for storage, it includes:
[0040] Send the video message data to the corresponding Topic according to the distributed camera id, and store the video message data in the corresponding Partition partition according to the acquisition angle corresponding to the video message data;
[0041] Each Topic contains multiple Partition partitions. Configure a Kafka Broker cluster node for each Partition partition in each Topic to form a Kafka Broker cluster. The number of threads in the thread pool is the same as the number of Partition partitions of its corresponding Topic.
[0042] As a preferred embodiment of the present invention, when performing fusion based on Mahalanobis distance, it includes:
[0043] Obtain the data detected by different types of environmental sensors and of the Mahalanobis distance;
[0044] Take the Mahalanobis distance as the trust function to obtain the data and between functions;
[0045] Using the said function, generate the credibility of the data and use the said credibility to obtain the information entropy of the data so as to further combine the data and the Mahalanobis distance with a number of data and a number of data the difference between the average Mahalanobis distances to generate the data a new credibility, and perform normalization to obtain the basic probability assignment value of the data ;
[0046] Use the basic probability assignment value to eliminate abnormal data, and perform data fusion on the remaining data through the combination rule according to the basic probability assignment value of the remaining data to obtain the said fused environment parameters.
[0047] As a preferred embodiment of the present invention, when obtaining the target position of the key frame, it includes:
[0048] Divide the video data into multiple time periods, and extract key frames from each time period;
[0049] Construct an improved as a construction site target detector, detect the target position by inputting the key frame into the construction site target detector, and input module;
[0050] When capturing spatial semantic information and dynamic information to generate a fused feature map through different paths, it includes:
[0051] Construct an improved model as a construction site feature extraction model, and input the collected video data;
[0052] Through path running at low temporal resolution and high spatial resolution to capture the spatial semantic information in the video data, and output the first feature map;
[0053] Through path running at high temporal resolution and low spatial resolution to capture the dynamic information of the rapidly changing behaviors in the video data, and using a smaller convolution width, output the second feature map;
[0054] Perform scale transformation on the second feature map and then fuse it with the first feature map to obtain the fused feature map, and input the fused feature map into module;
[0055] When completing behavior classification in combination with the target position of the key frame, it includes:
[0056] Through The module projects the target box generated based on the target position information onto the fusion feature map, obtains the corresponding feature matrix, and generates the fusion feature map after unified sizing. After pooling, fully connected layers, and behavior classification, video target data is obtained;
[0057] Among them, the improvement When detecting the head, use The module, in the neck network, Replace the module with a lightweight Module.
[0058] As a preferred embodiment of the present invention, when constructing the improved Model, it includes:
[0059] Introduce the attention module for the time dimension into the Path and The residual module of the path, and use the contribution of the weights to improve the attention mechanism;
[0060] Among them, when using the contribution of the weights to improve the attention mechanism, it includes: obtaining the variance and weights of each time dimension of the video data processed each time;
[0061] After performing time attention extraction based on the weights of the time dimension and the variance of the time dimension, output the first feature map and the second feature map, as shown in Formula 13:
[0062] (13);
[0063] In the formula, Is the first feature map or the second feature map, Is the activation function, Is the combination of the weights and variances of each time dimension, Is the variance of each time dimension, Is the weight of each time dimension, Is the video data processed each time.
[0064] As a preferred embodiment of the present invention, when extracting the environmental parameter feature representation, it includes:
[0065] Construct a sensor neural network based on three-dimensional convolution and input the preprocessed fused environmental parameters;
[0066] After being processed by the first Convolutional layer, successively pass through Activation function and The max pooling layer, and then successively pass through the second convolutional layer and the third convolutional layer for processing, while reducing the spatio-temporal resolution of the features, normalizing through the batch normalization layer, and outputting the environmental parameter feature representation through the fourth convolutional layer;
[0067] Among them, by using cavity convolutions with multiple dilation rates in the second convolutional layer and the third convolutional layer to extract the features of large-scale targets;
[0068] During the feature extraction process, a time series attention module is added to adaptively adjust the feature weights of each fused environmental parameter and highlight the feature representation of the environmental parameter. Residual connections and a feature pyramid structure are adopted to integrate features at different levels and provide multi-scale feature representations.
[0069] As a preferred embodiment of the present invention, when analyzing the comprehensive situation through the RNN construction site situation awareness model, it includes:
[0070] For each site unit in the construction site, and in combination with the personnel behavior rules and environmental parameter standards corresponding to each site unit, various anticipated behavior risk events and environmental risk events are scanned one by one to form a behavior-environment risk event rule dictionary;
[0071] Deep learning is performed on the behavior-environment risk event rule dictionary and the behavior-environment normal event rule dictionary through the RNN network model to establish a non-linear mapping relationship between the video-environment fusion features and the behavior risk events and environmental risk events, and an RNN construction site situation awareness model is obtained;
[0072] The video-environment fusion features obtained in real time are cross-identified using a feature error correction model to obtain an identified feature set; the identified feature set is input into the RNN construction site situation awareness model for identification to obtain real-time comprehensive situation data;
[0073] Among them, when performing deep learning, it includes: adding an algorithm to randomly discard some neurons in the model, and updating the parameters by introducing an algorithm and an exponential decay algorithm for the learning rate.
[0074] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0075] (1) Improve the monitoring coverage rate and deployment rationality of the construction site
[0076] By discretizing the construction site into units and optimizing the deployment of cameras based on the distance and angle relationships between site units, distributed cameras can effectively cover key areas and reduce monitoring blind spots.
[0077] (2)Realize the dynamic fusion of multi-modal environmental data
[0078] Using a multi-sensor data fusion algorithm based on Mahalanobis distance to eliminate abnormal environmental data, ensure data credibility, and generate high-precision fused environmental parameters through combination rules.
[0079] Deep feature extraction of the fused environmental parameters in the sensor neural network, through the time series attention module and multi-scale feature integration, provides comprehensive and accurate environmental dynamic perception ability.
[0080] (3)Improve the accuracy and robustness of object detection and behavior analysis
[0081] Adopt an improved construction site object detector and construction site feature extraction model, and capture the spatial semantic information and behavior dynamic information of construction site objects through the fusion of high and low time resolutions and spatial resolutions of multiple paths.
[0082] Introduce lightweight model design, which improves the computational efficiency while maintaining the detection accuracy, and is especially suitable for multi-object detection and behavior analysis in complex construction site environments.
[0083] (4)Realize the accurate analysis of the comprehensive situation and dynamic risk assessment
[0084] Based on the RNN construction site situation awareness model, fuse video target data and environmental parameter features, and establish a non-linear mapping relationship between behavior risk events and environmental risk events.
[0085] Through the construction of a behavior-environment risk event rule dictionary and deep learning optimization, it is possible to realize the comprehensive situation analysis of various complex construction site risk scenarios and identify potential risk events.
[0086] (5)Automated real-time risk warning and alarm information generation
[0087] Through real-time comprehensive situation data and Bayesian network analysis, the behavior and environmental conditions of personnel in the construction site are evaluated in real time to identify medium and high-risk events.
[0088] (6)Efficiency and stability of data collection and processing
[0089] Based on the multi-threaded design of the Kafka Broker cluster, efficient collection, storage and processing of massive video data and environmental parameters are realized, ensuring the real-time performance and stability of the system.
[0090] By means of a thread pool and a distributed storage strategy, the performance bottleneck problem of traditional monitoring systems in large-scale data processing is solved.
[0091] (7)Support the application and extension of complex construction site scenarios
[0092] Through the three-dimensional convolutional structure of the sensor neural network and the time series attention module, it can adapt to a variety of complex construction site scenarios (such as large-scale construction sites, dynamic construction environments).
[0093] The system design is highly modular, facilitating expansion according to actual application requirements, such as adding new sensor types or accessing more video data sources.
[0094] (8)Algorithm optimization and improvement of computing resource utilization efficiency
[0095] Introduce an attention mechanism, lightweight module, and dilated convolution in object detection and feature extraction, significantly reducing the computational overhead and improving the processing efficiency.
[0096] Add a neuron random dropout algorithm and a learning rate exponential decay mechanism to the deep learning model, significantly enhancing the generalization ability and stability of the model.
[0097] In summary, the present invention can significantly improve the intelligent level of construction site video monitoring, achieving a comprehensive improvement in monitoring coverage, data fusion accuracy, object detection accuracy, and risk assessment efficiency, providing reliable technical support for construction site safety management, and having important application value and practical significance.
[0098] The following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings
[0099] Figure 1 is a step diagram of the method for regional target and environmental dynamic perception for construction site video monitoring provided by the present invention;
[0100] Figure 2 is a real-time risk assessment flowchart provided by the present invention. Specific Embodiments
[0101] The method for regional target and environmental dynamic perception for construction site video monitoring provided by the present invention, as Figure 1 shown, includes the following steps:
[0102] Step S1: Initially deploy distributed cameras at the construction site, and optimize the deployment by obtaining unit discretization segmentation, distances and angles between units;
[0103] Step S2: Collect video data in multiple threads based on the optimized distributed cameras deployed on the Kafka Broker cluster. At the same time, collect environmental parameters through multiple types of environmental sensors and fuse them based on the Mahalanobis distance to obtain fused environmental parameters;
[0104] Step S3: Perform multi-object detection on the collected video data through the construction site target detector combined with the construction site feature extraction model, and analyze its behavioral characteristics to obtain video target data;
[0105] Step S4: Extract environmental parameter feature representations from the fused environmental parameters based on the sensor neural network, and perform multi-modal dynamic fusion with the video target data based on learnable parameters. Analyze the comprehensive situation through the RNN construction site situation awareness model to obtain real-time comprehensive situation data;
[0106] Step S5: Based on the real-time comprehensive situation data, conduct real-time risk assessment on the personnel behavior and environmental conditions within the construction site, identify medium-risk events and high-risk events, automatically generate warning information, and link relevant control systems to execute emergency responses.
[0107] Specifically, the multiple types of environmental sensors include: temperature sensors, humidity sensors, light intensity sensors, gas sensors, noise sensors, particulate matter and dust sensors.
[0108] The environmental parameters include: temperature, humidity, light intensity, gas concentration, and noise.
[0109] In the above step S1, when performing deployment optimization through unit discretization segmentation, obtaining the distance and angle between units, it includes:
[0110] Perform unit discretization segmentation on the construction site to obtain several site units;
[0111] Obtain the site units where the distributed cameras are set , and obtain the site units corresponding to the directions of the distributed cameras ;
[0112] Obtain the site units from the key areas that must be video-monitored , and based on the distance between the site units, obtain the site units , site units The included angle between them is shown in Formula 1:
[0113] (1);
[0114] In the formula, is the included angle, is the site unit and site unit The distance between them, is the distance between the site unit and the site unit ; is the distance between the site unit and the site unit ; is the distance between the site unit and the site unit ;
[0115] If the included angle is less than the preset value, it is considered that the distributed camera can monitor and cover the site unit ;
[0116] By obtaining the included angle between each site unit and the site unit corresponding to the direction of the distributed camera the monitoring coverage area of the distributed camera is determined
[0117] Furthermore, after determining the monitoring coverage area of the distributed camera, it includes
[0118] determining the first edge site unit of the building blocking the monitoring and the second edge site unit ;
[0119] obtaining the angle formed by the first edge site unit , the second edge site unit and the site unit ;
[0120] determining a site unit within the monitoring coverage area, and obtaining the angle formed by the first edge site unit , the site unit and the site unit ;
[0121] obtaining the angle formed by the second edge site unit , the site unit and the site unit ;
[0122] obtaining the maximum angle of the angle and judging whether the site unit is blocked by the building, as shown in formula 2
[0123] (2);
[0124] In the formula is the judgment parameter is the maximum angle , is to take , the larger value among is the distance between the site unit and the site unit ; is the distance between the first edge site unit and the site unit ; is the distance from the second edge site unit to the site unit ;
[0125] If the judgment parameter is 1, it indicates that the site unit is within the monitoring coverage area but is blocked by a building;
[0126] According to the monitoring coverage area and the judgment parameter , determine the site units in the key area that cannot be monitored and covered by the distributed cameras or are blocked by buildings;
[0127] By adjusting the installation positions of the distributed cameras and / or adding distributed cameras, all site units in the key area can be monitored and covered by the distributed cameras and are not blocked by buildings.
[0128] In the above step S2, when collecting video data based on the Kafka Broker cluster in multiple threads, it includes:
[0129] Receive video data from the distributed cameras, perform parsing and encapsulation, and after serialization, send the video message data to the Kafka message queue for storage;
[0130] Create a KafkaInputDStream object as a subclass of ReceiverInputDStream, and use the KafkaInputDStream object to create multiple KafkaReceiver receivers with the same groupId for Kafka data in the same Topic to pull Kafka data from different Partition partitions, and use the union() method to aggregate the Kafka data pulled by multiple KafkaReceiver receivers;
[0131] Deserialize the aggregated Kafka data, convert the obtained Json message into Json data format, and use the OpenCV vision library to parse and convert the Json data into a Mat data type that can be used for image processing;
[0132] Among them, when sending to the Kafka message queue for storage, it includes:
[0133] By obtaining the Key value of the video message data, determining the distributed camera id to which the video message data belongs, and sending the video message data to the corresponding Topic according to the distributed camera id, and storing the video message data in the corresponding Partition partition according to the acquisition angle corresponding to the video message data;
[0134] Each Topic contains multiple Partition partitions. By creating a Kafka Producer object, a Kafka Broker cluster node is configured for each Partition partition in each Topic to form a Kafka Broker cluster.
[0135] Furthermore, when pulling Kafka data from different Partition partitions, it includes:
[0136] Call the onStart() method to start receiving Kafka data, and establish a connection with the Kafka Broker cluster to obtain a consumerConnector connection object;
[0137] Use the consumerConnector connection object to create a TopicMessageStreams to store the mapping relationship table List[KafkaStream] of the Topic and its Partition partitions;
[0138] Create a thread pool executorPool for each Topic, and perform message consumption through the thread pool executorPool;
[0139] Create a MessageHandler message processor for each Partition partition data stream, and configure a processor thread for the MessageHandler message processor in the thread pool executorPool;
[0140] After the processor thread starts, the KafkaReceiver receiver will pull Kafka data and pass it to the ReceiverSupervisor for processing;
[0141] The ReceiverSupervisor will store the received Kafka data in the form of data blocks through the BlockManager;
[0142] The ReceiverSupervisor sends a message with the encapsulated data block as a parameter to the ReceiverTracker, which acts as a scheduler, through the AddBlock method to notify it of the new data block;
[0143] The ReceiverTracker assigns the received data blocks to the batch as input data for the resilient distributed dataset at every preset time;
[0144] Among them, each element in the mapping relationship table List[KafkaStream] represents a data stream of a Partition of this Topic, and the number of threads in the thread pool executorPool is the same as the number of Partitions of the corresponding Topic.
[0145] In the above step S2, when performing fusion based on the Mahalanobis distance, it includes:
[0146] Obtain the data detected by different types of environmental sensors and the Mahalanobis distance, as shown in Equation 3:
[0147] (3);
[0148] In the formula, is the Mahalanobis distance between the data and , is the transpose, is the covariance matrix of the multi-dimensional random variable;
[0149] Take the Mahalanobis distance as the belief function to obtain the and between the data function, as shown in Equation 4:
[0150] (4);
[0151] In the formula, is the and between the data function;
[0152] Use the function to generate the credibility of the data , and use the credibility to obtain the information entropy of the data , thereby further generating the new credibility of the data , as shown in Equation 5, Equation 6, and Equation 7:
[0153] (5);
[0154] In the formula, is the credibility of the data , is the number of data and data .
[0155] (6);
[0156] In the formula, is the information entropy of the data .
[0157] (7);
[0158] In the formula, is the new credibility of the data , is the difference between the Mahalanobis distance of the data and and the average Mahalanobis distance of data and data , as shown in formula 8:
[0159] (8);
[0160] In the formula, is the Mahalanobis distance of the data and , is the average Mahalanobis distance of data and data
[0161] Let the new credibility of the data be normalized to obtain the basic probability assignment value of the data , as shown in formula 9:
[0162] (9);
[0163] In the formula, is the basic probability assignment value of the data ;
[0164] Using the basic probability assignment value, abnormal data is removed from the data , and data fusion is performed on the remaining data according to the combination rule of Dempster-Shafer evidence theory to obtain the fused environmental parameters.
[0165] In the above step S3, when performing multi-object detection on the collected video data, it includes:
[0166] Dividing the collected video data into multiple time periods according to a preset time interval, and according to a preset sampling number of frames and step size Extracting key frames from each time period, as shown in formula 10:
[0167] (10);
[0168] In the formula, is the key frame;
[0169] Construct an improved as a construction site target detector, detect the target position by inputting the key frame into the construction site target detector, and input the detected target position information into module;
[0170] Construct an improved model as a construction site feature extraction model, input the collected video data into the construction site feature extraction model, and run through the path in the construction site feature extraction model with low temporal resolution and high spatial resolution to capture the spatial semantic information in the video data and output the first feature map;
[0171] Run through the path in the construction site feature extraction model with high temporal resolution and low spatial resolution to capture the dynamic information of the rapidly changing behavior in the video data, and adopt a smaller convolution width to reduce the channel capacity to reduce the computational amount, and output the second feature map;
[0172] After performing scale transformation on the second feature map through convolution, fuse it with the first feature map to obtain a fused feature map, and input the fused feature map into module;
[0173] Through the module, project the target box generated based on the target position information onto the fused feature map to obtain the corresponding feature matrix, and expand each along the time dimension to obtain , generate a fused feature map with a unified size, and after being processed by the pooling layer and the fully connected layer, use classifier for behavior classification to obtain video target data;
[0174] Among them, the improved adopts a module at the detection head, and replaces the module with a lightweight module in the neck network;
[0175] Through The module improves the model's utilization of multi-scale features. By means of lightweight The module improves the neck network structure, reduces the number of model parameters and computational complexity, and reduces the model inference time;
[0176] Path and The path both include: a convolutional layer, a pooling layer, a first residual module, a second residual module, a third residual module, and a fourth residual module connected in sequence;
[0177] Improve The length of the video data processed by the model each time is frames, where The path samples at a stride of frames for processing, The path samples at a stride of frames for processing. Sample frames for processing.
[0178] Furthermore, when constructing the improved model as a construction site feature extraction model, it includes:
[0179] Introduce the attention module for the time dimension into the path and the residual modules of the path, and use the contribution of weights to improve the attention mechanism;
[0180] Among them, when using the contribution of weights to improve the attention mechanism, it includes: obtaining the variance of each time dimension of the video data processed each time, as shown in Equation 11:
[0181] (11);
[0182] In the formula, is the variance of each time dimension, is the video data processed each time, is the mean of the video data processed each time, is the standard deviation of the video data processed each time, and are respectively the scale and displacement trainable affine transformation parameters, is a constant used to prevent the variance from being 0, is the attention module for the time dimension;
[0183] Obtain the weight of each time dimension, as shown in Equation 12:
[0184] (12);
[0185] In the formula, is the weight of each time dimension, is the scale factor of each time dimension, used to represent the change intensity of time; is the number of time dimensions;
[0186] After time attention extraction is performed according to the weight of the time dimension and the variance of the time dimension, a first feature map and a second feature map are output, as shown in Formula 13:
[0187] (13);
[0188] In the formula, is the first feature map or the second feature map, is the activation function, is the combination of the weight and variance of each time dimension.
[0189] In the above step S4, when extracting the environmental parameter feature representation, it includes:
[0190] Construct a sensor neural network based on three-dimensional convolution, and the sensor neural network includes: a first convolution layer, max pooling layer, a second convolution layer, a third convolution layer, a batch normalization layer, and a fourth convolution layer;
[0191] After preprocessing the fused environmental parameters, input them into the sensor neural network. After being processed by 32 convolution kernels in the first convolution layer, sequentially pass through the activation function and the max pooling layer, and then sequentially pass through 64 convolution kernels in the second convolution layer, 128 convolution kernels in the third convolution layer for processing, while reducing the spatio-temporal resolution of the features, normalizing through the batch normalization layer, and outputting the environmental parameter feature representation of 256 channels by using convolution kernels in the fourth convolution layer;
[0192] Among them, features of large-scale targets are extracted by using cavity convolution with multiple dilation rates in the intermediate layer of the sensor neural network;
[0193] During the feature extraction process, a time series attention module is added. The time series attention module is used to obtain the correlation between adjacent fused environmental parameters, adaptively adjust the feature weights of each fused environmental parameter, and highlight the feature representation of the environmental parameters. The residual connection and feature pyramid structure are adopted to alleviate the vanishing gradient of the deep network. At the same time, the feature pyramid structure is used to integrate features at different levels and provide multi-scale feature representations.
[0194] In step S4 above, when performing multi-modal dynamic fusion, it includes:
[0195] Through learnable parameters dynamically adjust the weights of the video target data and the feature representation of the environmental parameters, and perform multi-modal fusion according to the dynamically adjusted weights to obtain the video-environment fusion features, as shown in Equation 13:
[0196] (13);
[0197] In the formula, is the video-environment fusion feature, is the feature representation of the environmental parameters, is the video target data;
[0198] Among them, the learnable parameter , whose value is dynamically generated according to the statistics of the input features, so as to adaptively adjust the fusion strategy according to different construction site scenarios.
[0199] In step S4 above, when analyzing the comprehensive situation through the RNN construction site situation awareness model, it includes:
[0200] For each site unit in the construction site, and in combination with the personnel behavior rules and environmental parameter standards corresponding to each site unit, scan various anticipated behavior risk events and environmental risk events one by one to form a behavior-environment risk event rule dictionary;
[0201] Provide a behavior-environment regular event rule dictionary, and perform deep learning on the behavior-environment risk event rule dictionary and the behavior-environment regular event rule dictionary through the RNN network model to establish a non-linear mapping relationship between the video-environment fusion feature and the behavior risk event and the environmental risk event, and obtain the RNN construction site situation awareness model, as shown in Equation 14:
[0202] (14);
[0203] In the formula, is the video-environment fusion feature, is the number of video-environment fusion features, is the behavior risk event, is the number of behavioral risk events, is an environmental risk event, is the number of environmental risk events, is the non - linear mapping relationship between video - environment fusion features and behavioral risk events, is the non - linear mapping relationship between video - environment fusion features and environmental risk events;
[0204] The video - environment fusion features obtained in real - time are used by a feature error - correction model for cross - identification to obtain an identification feature set;
[0205] The identification feature set is input into an RNN construction site situation awareness model for identification to obtain real - time comprehensive situation data;
[0206] Among them, when performing deep learning, it includes:
[0207] The video - environment fusion features, various behavioral risk events, and various environmental risk events are encoded by numbering, and through normalization for data normalization processing, and an algorithm is added to the model. By randomly discarding some neurons, the structure of the neural network is changed to reduce the dependence between neurons. By introducing an algorithm and an exponential decay algorithm of the learning rate to update the parameters, so as to accelerate the convergence speed of the model.
[0208] In the above step S5, as Figure 2 shown, when performing real - time risk assessment, it includes:
[0209] The real - time comprehensive situation data is analyzed for correlation using a Bayesian network and a behavior - environment risk event rule dictionary to determine whether there are strongly correlated behavior events and / or environmental events, and whether there are moderately correlated behavior events and / or environmental events;
[0210] If there are, the strongly correlated behavior events are identified as high - risk behavior events, and the strongly correlated environmental events are identified as high - risk environmental events;
[0211] The moderately correlated behavior events are identified as medium - risk behavior events, and the moderately correlated environmental events are identified as medium - risk environmental events.
[0212] The above - mentioned implementation manners are only the preferred implementation manners of the present invention, and cannot be used to limit the scope of protection of the present invention. Any non - substantial changes and substitutions made by those skilled in the art on the basis of the present invention belong to the scope of protection required by the present invention.
Claims
1. A method for dynamically perceiving regional targets and environment for construction site video monitoring, characterized in that, It includes the following steps: Perform an initial deployment of distributed cameras at the construction site, and optimize the deployment by unit discretization segmentation, obtaining the distance and angle between units; Based on the optimized distributed cameras, collect video data through the Kafka Broker cluster in a multi-threaded manner. At the same time, collect environmental parameters through multi-type environmental sensors and fuse them based on the Mahalanobis distance to obtain fused environmental parameters; Perform multi-object detection on the collected video data, and analyze its behavioral characteristics to obtain video target data; Extract environmental parameter feature representations from the fused environmental parameters based on the sensor neural network, and perform multi-modal dynamic fusion with the video target data based on learnable parameters. Analyze the comprehensive situation through the RNN construction site situation awareness model to obtain real-time comprehensive situation data; Based on the real-time comprehensive situation data, use Bayesian network analysis to perform real-time risk assessment on the personnel behavior and environmental conditions within the construction site, identify risk events, and automatically generate warning information; Among them, when performing multi-object detection, it includes: obtaining the target position of the key frame by constructing a construction site target detector; constructing a construction site feature extraction model to capture spatial semantic information and dynamic information through different paths to generate a fused feature map, and combining the target position of the key frame to complete behavior classification; When obtaining the target position of the key frame, it includes: Divide the video data into multiple time periods, and extract key frames from each time period; Build improvement As a construction site target detector, the target position is detected by inputting key frames into the construction site target detector, and module; When capturing spatial semantic information and dynamic information through different paths to generate a fused feature map, it includes: Build improvement Construct a model as a construction site feature extraction model and input the collected video data; Via The path runs with low temporal resolution and high spatial resolution, captures the spatial semantic information in the video data, and outputs the first feature map; Through The path runs with high temporal resolution and low spatial resolution, captures the dynamic information of the rapidly changing behavior in the video data, and adopts a smaller convolution width to output a second feature map; After performing a scale transformation on the second feature map and fusing it with the first feature map, a fused feature map is obtained, and the fused feature map is input into module; When combining the target position of the key frame to complete behavior classification, it includes: Through The module projects the target box generated based on the target location information onto the fused feature map, obtains the corresponding feature matrix, generates the fused feature map after unified sizing, and through pooling, fully connected layers, and behavior classification, obtains the video target data; Among them, the improvement adopts a module at the detection head, and in the neck network, the module is replaced with a lightweight module.
2. The method for regional target and environmental dynamic perception for construction site video monitoring according to claim 1, characterized in that, When performing deployment optimization, it includes: Perform unit discretization segmentation on the construction site to obtain several site units; Obtain the site unit set by the distributed camera and obtain the site unit corresponding to the direction of the distributed camera ; Obtain site units from key areas that must be video-monitored and, based on the distances between the site units, obtain the angle between site unit and site unit If the included angle is less than a preset value, it is considered that the distributed camera can monitor and cover the site unit for monitoring coverage; By obtaining the site units corresponding to the directions of the distributed cameras for each site unit and the included angles therebetween, the monitoring coverage area of the distributed cameras is determined.
3. The method for dynamically perceiving regional targets and environment for construction site video monitoring according to claim 2, characterized in that, After determining the monitoring coverage area of the distributed cameras, it includes: Determine a first edge site unit of a building for occlusion monitoring and a second edge site unit ; Obtain the first edge site unit , the second edge site unit and the site unit to form an angle ; Determine a site unit within the monitoring coverage area , based on the angles formed with the first edge site unit , the second edge site unit , the site unit , the distances between the site units, and the angle , determine the value of the judgment parameter , so as to determine whether the site unit is blocked by a building; According to the monitoring coverage area and judgment parameters , determine the site units within the key area that cannot be covered by monitoring or are blocked by buildings, so as to adjust the distributed cameras so that all site units within the key area can be covered by monitoring and are not blocked by buildings.
4. The method for dynamically perceiving regional targets and environment for construction site video monitoring according to claim 3, characterized in that, When collecting video data, it includes: Receive the video data for parsing and encapsulation, serialize it, and then send the video message data to the Kafka message queue for storage; Create multiple receivers with the same groupId for the Kafka data in the same Topic, pull the Kafka data in different Partition partitions based on the thread pool, and store it in the form of data blocks. Aggregate the Kafka data pulled by multiple receivers; After deserializing the aggregated Kafka data, convert the obtained Json message into the Json data format, and parse and convert the Json data into the Mat data type that can be used for image processing.
5. The method for dynamically perceiving regional targets and environment for construction site video monitoring according to claim 4, characterized in that When sending it to the Kafka message queue for storage, it includes: Send the video message data to the corresponding Topic according to the distributed camera id, and store the video message data in the corresponding Partition partition according to the acquisition angle corresponding to the video message data; Each Topic contains multiple Partition partitions. Configure a Kafka Broker cluster node for each Partition partition in each Topic to form a Kafka Broker cluster. The number of threads in the thread pool is the same as the number of Partition partitions of its corresponding Topic.
6. The method for region target and environmental dynamic perception for construction site video monitoring according to any one of claims 1-5, characterized in that, When performing fusion based on Mahalanobis distance, it includes: Obtain data detected by different types of environmental sensors and Mahalanobis distance; Taking the Mahalanobis distance as a belief function to obtain data and between function; Using the said function, generate the credibility of the data , and use the credibility to obtain the information entropy of the data , so as to further combine the Mahalanobis distance between the data and with the difference between the average Mahalanobis distance of data and data to generate the new credibility of the data , and perform normalization to obtain the basic probability assignment value of the data ; Eliminate abnormal data using the basic probability assignment value, and perform data fusion on the remaining data through the combination rule based on the basic probability assignment value to obtain the fused environmental parameters.
7. The method for dynamically perceiving regional targets and environment for construction site video monitoring according to claim 1, characterized in that, In constructing an improved model, it includes: Introduce the attention module for the time dimension into the path sum the residual module of the path, and use the contribution of the weights to improve the attention mechanism; Among them, when improving the attention mechanism by using the contribution of weights, it includes: obtaining the variance and weights of each time dimension of the video data processed each time; After performing temporal attention extraction based on the weights of the time dimensions and the variances of the time dimensions, the first feature map and the second feature map are output, as shown in Equation 13: (13); Wherein, is the first feature map or the second feature map, is the activation function, is the combination of the weight and variance of each time dimension, is the variance of each time dimension, is the weight of each time dimension, is the video data processed each time.
8. The method for regional target and environmental dynamic perception for construction site video monitoring according to claim 7, characterized in that When extracting the feature representation of environmental parameters, it includes: Constructing a sensor neural network based on three-dimensional convolution and inputting the preprocessed fused environmental parameters; After the first convolution layer processing, it successively passes through the activation function and the max pooling layer, and then successively passes through the second convolution layer and the third convolution layer for processing, while reducing the spatio-temporal resolution of the features, normalizing through the batch normalization layer, and outputting the environmental parameter feature representation through the fourth convolution layer; Among them, by using at the second convolutional layer and the third convolutional layer, dilated convolutions with multiple dilation rates are used to extract features of large-scale targets; Adding a time series attention module during the feature extraction process, adaptively adjusting the feature weights of each fused environmental parameter and highlighting the feature representation of the environmental parameters, adopting residual connection and feature pyramid structure to integrate features at different levels and providing multi-scale feature representations.
9. The method for dynamically perceiving regional targets and environment for construction site video monitoring according to claim 8, wherein, When analyzing the comprehensive situation through the RNN construction site situation awareness model, it includes: For each site unit in the construction site, and combining the personnel behavior rules and environmental parameter standards corresponding to each site unit, scanning various anticipated behavioral risk events and environmental risk events one by one to form a behavior-environment risk event rule dictionary; Deep learning the behavior-environment risk event rule dictionary and the behavior-environment normal event rule dictionary through the RNN network model to establish a non-linear mapping relationship between the video-environment fusion features and the behavioral risk events and environmental risk events, and obtaining the RNN construction site situation awareness model; The video-environment fusion features obtained in real time are cross-identified using a feature error correction model to obtain an identified feature set; the identified feature set is input into an RNN construction site situation awareness model for identification to obtain real-time comprehensive situation data; Among them, when performing deep learning, it includes: adding an algorithm to randomly discard some neurons, and updating the parameters by introducing an algorithm and an exponential decay algorithm of the learning rate.
Citation Information
Patent Citations
Equipment fault diagnosis method and device based on multi-sensor data fusion
CN111931806A
Risk identification and intelligent pre-control system and method for coal mine driving face
CN114673558A
Underground mine abnormal video action understanding method combining transfer learning and regional intrusion
CN117423157A