A vision-based deep learning instance segmentation and tracking method for off-site construction
By using deep learning mask R-CNN algorithm and Kalman filter optimization worker tracking method in off-site construction environment, the problem of poor tracking performance in the off-site construction environment is solved, and higher tracking accuracy and robustness are achieved, and construction site safety and detection efficiency are improved.
Patent Information
- Application Number
- CN202210565509.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-23
AI Technical Summary
The existing vision-based worker tracking method is difficult to obtain robust tracking performance in off-site construction environments, especially when dealing with occlusion, scale changes, background clutter and sudden movements, which are not robust and difficult to meet the needs of long-term monitoring and analysis.
The mask R-CNN algorithm based on deep learning is used for instance segmentation, combined with Kalman filtering and fuzzy inference, optimize the correlation steps of worker tracking, and use Hungarian algorithm to perform instance assignment to improve the robustness and accuracy of tracking.
It improves the tracking accuracy and robustness of off-site construction workers, and makes construction site management stronger tracking, thereby improving the visual inspection efficiency and on-site safety of non-site construction workers.
Smart Images

Figure CN114897937B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and deep learning, and in particular relates to a vision-based deep learning instance segmentation and tracking method for off-site construction, including a vision-based instance segmentation method and a mask-based optimization method for an association step. Background Art
[0002] Instance segmentation is to detect targets in images to obtain target areas of different categories, subdivide target areas of the same category to obtain specific target candidate areas, and segment each candidate area to obtain the segmentation result of the target image.
[0003] Instance segmentation is widely used in various computer vision processing tasks, such as autonomous driving, medical diagnosis, and public security management. However, current instance segmentation technology has low accuracy and weak robustness, making it difficult to meet the needs of long-term monitoring and analysis.
[0004] Vision-based worker tracking refers to extracting worker trajectories from video recordings, which is a basic step in vision-based construction worker monitoring. Many studies in the construction industry have developed vision-based tracking methods to facilitate monitoring of construction workers. It is worth noting that most of these studies focus on tracking workers in on-site construction, while only a few scholars have devoted themselves to developing methods for tracking off-site construction workers. Chu et al. adopted the MDNet method to track workers in off-site construction, which learned multiple domain features from Convolutional Neural Networks to achieve robust tracking performance in construction scenarios. However, their method was developed for tracking a single target and cannot be directly used for multiple worker tracking.
[0005] In off-site construction, vision-based tracking conditions are complex due to the relatively limited workspace, large number of workers, and frequent posture changes. Existing vision-based tracking methods have difficulty in achieving robust worker tracking performance in off-site construction environments for two reasons. First, existing methods use target detection to identify worker objects, which is difficult to achieve reliable performance when dealing with tracking challenges such as occlusion, scale changes, background clutter, and sudden movements; second, existing methods associate worker objects across frames based on bounding box information. Considering that workers have similar visual features when wearing personal protective equipment (Personal Protective Equipment), horizontal association of bounding boxes is prone to errors in off-site construction. Summary of the invention
[0006] In view of the above-mentioned defects of the prior art, the purpose of the present invention is to provide a vision-based off-site construction deep learning instance segmentation and tracking method, which solves the above-mentioned technical problems through a vision-based instance segmentation method and a mask-based optimization method for the association step.
[0007] A technical solution adopted by the present invention is:
[0008] A vision-based off-site construction deep learning instance segmentation tracking method includes an instance segmentation method, which includes the following steps:
[0009] Step S1: Instance segmentation: The instance segmentation module detects workers from the input video and obtains the worker's bounding box and segmentation mask;
[0010] Step S2: Instance association: The instance association module constructs a working instance association matrix on every two consecutive frames of the input video;
[0011] Step S3: Instance assignment: The instance assignment module generates tracking results in instance assignment using the Hungarian algorithm.
[0012] In the step S1, the instance segmentation module uses the deep learning-based Mask R-CNN algorithm to mask the R-CNN; the Mask R-CNN method includes three modules, namely, a feature extractor module, a region proposal network module (RPN) and an extended classifier network module (Extended Classifier Network);
[0013] By using the Masked R-CNN algorithm to describe the body parts of the worker, a more accurate mask is obtained to describe the occluded worker compared to the target detection and segmentation method.
[0014] The specific steps of the Mask R-CNN deep learning algorithm are as follows:
[0015] The input image is processed by a ResNet101 neural network to extract feature maps in the feature extraction module; the feature map of the N×N spatial window is slid using the RPN module; in this module, 12 anchor boxes are initialized as regions of interest (ROI) for each sliding window, and these anchor boxes are defined by three aspect ratios (1:1, 1:2, 2:1) and four scale ratios (322, 642, 1282, 5122); each region of interest is processed by a three-layer convolutional network, and a two-layer fully connected box classification and box regression RPN module generates 300 ROIs according to their likelihood of being objects.
[0016] The Mask R-CNN deep learning algorithm uses the region of interest alignment technology to extract fixed shape features from the feature map of each region of interest; in the ECLS module, these fixed shape features are processed by three neural networks for box classification, box regression and mask regression respectively; among them, the box classification network generates the confidence of each ROI belonging to any predefined class; the regression network predicts the pixel coordinates of each object, and the mask regression network predicts the segmentation mask of each object at the pixel level.
[0017] The instance association module in step S2 adopts a mask-based approach to optimize the association step of worker tracking.
[0018] Combining Kalman filter prediction and fuzzy reasoning, a new instance association method is obtained.
[0019] Taking the bounding box and mask of the worker generated by the instance segmentation module in step S1 as input, the segmentation results of every two consecutive frames of workers are associated, and an association matrix is generated for the instance assignment module.
[0020] In the instance association module of step S2, when an image is input, once a new worker instance is segmented in a frame, the Kalman filter tracker will be initialized to track this instance by using the bounding box information of this worker, where a unique ID number is assigned to the Tracklet; the Kalman filter uses a series of observation data that changes over time and produces an estimate of the next time step; the state of each object is simulated as follows:
[0021] STATE = [Cx, Cy, u, v]
[0022] Where Cx and Cy represent the horizontal and vertical coordinates of the center point of the object's bounding box, respectively; u and v represent the object's velocities in the horizontal and vertical coordinates; in other words, the Kalman filter is only used to track the center point of the artifact instance, not the bounding box.
[0023] The specific steps of step S2 associating the modules are as follows:
[0024] The Kalman filter uses the bounding box information of the previous frame to predict the center point position of the worker in the current frame; the motion vector is calculated as the center point movement of the same worker instance between the current frame and the previous frame; the tracking mask on the current frame is obtained by adding the detected mask to the motion vector; the association matrix on the current frame is calculated as the Mask Intersection-Over-Union of the tracking mask and the segmentation mask on the current frame.
[0025] The beneficial technical effects of the present invention are:
[0026] 1. Introducing instance segmentation into off-site construction worker tracking makes construction site management more traceable, thereby improving construction site safety;
[0027] 2. A new dataset of construction worker images is proposed for training instance segmentation methods;
[0028] 3. A new mask-based instance association method is proposed to improve the robustness of worker tracking;
[0029] 4. Improve the visual inspection efficiency and on-site safety of off-site construction workers as a whole. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flow chart of an embodiment of the vision-based off-site construction deep learning instance segmentation and tracking method provided by the present invention;
[0031] Figure 2 It is a flowchart of an embodiment of the Masked R-CNN method provided by the present invention;
[0032] Figure 3 It is a schematic diagram of the detection results of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0033] The embodiments of the present invention are described in detail below. The following embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation methods and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0034] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments without conflict.
[0035] Example:
[0036] The present invention provides a vision-based off-site construction deep learning instance segmentation and tracking method, which includes a vision-based instance segmentation method and a mask-based optimization method for the association step.
[0037] Figure 1 is a flow chart of an example segmentation method provided by an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0038] Step S1, instance segmentation module, uses the deep learning algorithm Mask R-CNN to detect workers from all frames of the input video; the segmentation result of Mask R-CNN contains two kinds of information about the workers, namely the bounding box and the mask, where the segmentation mask can provide richer worker instance information at the pixel level;
[0039] Step S2, instance association module, constructs a working instance association matrix on every two consecutive frames; first, the Kalman filter uses the bounding box information of the previous frame to predict the center point position of the worker in the current frame; then, the motion vector is calculated as the center point movement of the same working instance between the current frame and the previous frame; by adding the detected mask to the motion vector, the tracking mask on the current frame can be obtained; finally, the association matrix on the current frame is calculated as the mask intersection and union of the tracking mask and the segmentation mask on the current frame;
[0040] Step S3, the instance allocation module uses the Hungarian algorithm to generate tracking results.
[0041] Figure 2 The Mask R-CNN algorithm based on vision and combined with deep learning is provided in an embodiment of the present invention. Figure 2 As shown, the algorithm includes the following:
[0042] Three main modules, namely feature extractor, region proposal network and extended classifier network; First, the input image is processed by ResNet101 neural network to extract feature map in feature extraction module; Then, the feature map of N×N spatial window is slid using RPN module; In this module, 12 anchor boxes are initialized as regions of interest for each sliding window, which are defined by three aspect ratios (1:1, 1:2, 2:1) and four scale ratios (322, 642, 1282, 5122); Each region of interest is processed by a three-layer convolutional network, Two layers of fully connected box classification and box regression are used; finally, the RPN module generates 300 ROIs based on their likelihood of being objects; in addition, the region of interest alignment technique is used to extract fixed shape features from the feature map of each region of interest; in the ECLS module, these fixed shape features are processed by three neural networks for box classification, box regression and mask regression respectively; among them, the box classification network produces the confidence of each ROI belonging to any predefined class; the regression network predicts the pixel coordinates of each object, and the mask regression network predicts the segmentation mask of each object at the pixel level.
[0043] In the step S2 of the present invention, in the instance association module, the following optimization method is adopted:
[0044] In the input image, once a new worker instance is segmented in a frame, the Kalman filter tracker will be initialized to track this instance by using the bounding box information of this worker, where a unique ID number is assigned to the tracklet; the Kalman filter uses a series of observations over time and produces an estimate of the next time step, which has been used for vision-based tracking; in this study, the state of each object is simulated as follows:
[0045] STATE = [Cx, Cy, u, v]
[0046] Where Cx and Cy represent the horizontal and vertical coordinates of the center point of the object's bounding box, respectively; u and v represent the object's velocities in the horizontal and vertical coordinates; in other words, the Kalman filter is only used to track the center point of the artifact instance, not the bounding box.
[0047] The present invention focuses on improving the efficiency of visual detection by optimizing the method based on the principle of computer vision, which includes the following steps: Step S1: Instance segmentation: The instance segmentation module detects the worker from the input video and obtains the bounding box and segmentation mask of the worker; Step S2: Instance association: The instance association module constructs a work instance association matrix on every two consecutive frames of the input video; Step S3: Instance assignment: The instance assignment module generates tracking results in the instance assignment using the Hungarian algorithm. The present invention is used to track the body contours of off-site construction workers on the construction site, introduces instance segmentation into off-site construction worker tracking, makes the site management more traceable, and thus improves the visual detection efficiency and on-site safety of off-site construction workers.
[0048] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A vision-based deep learning instance segmentation and tracking method for off-site construction, It is characterized in that The following steps are involved: Step S1: Instance segmentation: The instance segmentation module detects workers from the input video and obtains the worker's bounding box and segmentation mask; In the step S1, the instance segmentation module uses a mask R-CNN algorithm based on deep learning to mask R-CNN; the mask R-CNN algorithm includes three modules: a feature extractor module, a region proposal network module, and an extended classifier network module; The specific steps of the Mask R-CNN algorithm are as follows: The input image is processed by a ResNet101 neural network to extract feature maps in the feature extraction module; the feature map of the N×N spatial window is slid using the RPN module; in this module, 12 anchor boxes are initialized as regions of interest for each sliding window, which are defined by three aspect ratios (1:1, 1:2, 2:1) and four scale ratios (322, 642, 1282, 5122); each region of interest is processed by a three-layer convolutional network, a two-layer fully connected box classification and box regression RPN module to generate 300 ROIs according to their likelihood of being an object; Step S2: Instance association: The instance association module constructs a worker instance association matrix on every two consecutive frames of the input video; In the instance association module of step S2, when an image is input, once a new worker instance is segmented in a frame, the Kalman filter tracker will be initialized to track the worker instance by using the bounding box information of the worker instance, where a unique ID number is assigned to the Tracklet; the Kalman filter uses a series of observation data that changes over time and generates an estimate of the next time step; the state of each worker instance is simulated as follows: STATE = [Cx, Cy, u, v] Where Cx and Cy represent the horizontal and vertical coordinates of the center point of the object's bounding box, respectively; u and v represent the object's velocities in the horizontal and vertical coordinates; the Kalman filter is used to track the center point of the worker instance instead of tracking the bounding box; The specific steps of step S2 associating the modules are as follows: The Kalman filter uses the bounding box information of the previous frame to predict the center point position of the worker instance in the current frame; the motion vector is calculated as the center point movement of the same worker instance between the current frame and the previous frame; the tracking mask on the current frame is obtained by adding the detected mask to the motion vector; the association matrix on the current continuous frame is calculated as the mask intersection and union of the tracking mask and the segmentation mask on the current frame; Step S3: Instance assignment: The instance assignment module generates tracking results in instance assignment using the Hungarian algorithm; The vision-based off-site construction deep learning instance segmentation tracking method introduces instance segmentation into off-site construction worker tracking, making construction site management more traceable, thereby improving the visual inspection efficiency and on-site safety of off-site construction workers.
2. According to claim 1, the vision-based deep learning instance segmentation and tracking method for off-site construction, It is characterized in that The Mask R-CNN algorithm uses masks to describe the body parts of the worker, obtaining a more accurate mask to describe the occluded worker than the target detection and segmentation method.
3. The Mask R-CNN deep learning algorithm according to claim 1, It is characterized in that The Mask R-CNN algorithm uses the region of interest alignment technology to extract fixed shape features from the feature map of each region of interest; in the ECLS module, these fixed shape features are processed by three neural networks for box classification, box regression and mask regression respectively; among which, the box classification network generates the confidence of each ROI belonging to any predefined class; the regression network predicts the pixel coordinates of each object, and the mask regression network predicts the segmentation mask of each object at the pixel level.
4. The vision-based off-site construction deep learning instance segmentation and tracking method according to claim 1, It is characterized in that The instance association module in step S2 uses a mask-based approach to optimize the association step of worker tracking.
5. The instance association module of step S2 according to claim 4, It is characterized in that In the S2 instance association module in the step, Kalman filter prediction and fuzzy reasoning are combined to obtain an instance association method.
6. The instance association module of step S2 according to claim 5, It is characterized in that Taking the bounding box and segmentation mask of the worker generated by the instance segmentation module in step S1 as input, the segmentation results of every two consecutive frames of workers are associated, and an association matrix is generated for the instance assignment module.
Citation Information
Patent Citations
Near-drowning behavior detection method fusing UWB indoor location with video target detection and tracking technology
CN109102678A
Multi-target tracking method based on Mask R-CNN and apparent feature fusion
CN113506317A