Construction personnel multi-target tracking method and system based on detection

Through the improved YOLOV8 model and the collaborative pooling attention mechanism, combined with deformable convolution and BYTE association algorithm, the problem of insufficient accuracy and speed in the existing construction personnel tracking technology is solved, and fast and reliable multi-objective tracking of construction personnel is achieved.

CN119919447APending Publication Date: 2025-05-02国网甘肃省电力公司金昌供电公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411695587.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing construction personnel tracking technology has problems such as insensitive targets and complex backgrounds, and the accuracy and speed need to be improved. The traditional human eye supervision is low, insufficient accuracy and high cost.

Method used

The multi-objective tracking method of construction workers is adopted to improve the accuracy and speed of object detection and tracking by building an improved YOLOV8 model and a collaborative pooling attention mechanism, combining deformable convolution and BYTE association algorithm.

Benefits of technology

Fast and reliable multi-objective tracking for construction workers is achieved, the accuracy and speed of the model is improved, the computational complexity is reduced, the real-time detection is enhanced, and the problem of target loss and mismatch is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919447A_ABST
    Figure CN119919447A_ABST
Patent Text Reader

Abstract

The invention discloses a construction personnel multi-target tracking method and system based on detection, and the method comprises the steps: constructing a target detector, training the target detector through employing a training data set, obtaining a trained target detector, and enabling the target detector to comprise an improved YOLOV8 model; inputting a video to be tracked into the trained target detector frame by frame to obtain a target detection result of each frame of image; and a target tracker is constructed based on a BYTE association algorithm, the input end of the target tracker is connected with the output end of the target detector, and the target tracker is used for processing the target detection result output by the target detector to obtain a target tracking result of the constructor. Based on a collaborative pooling attention mechanism, the structure of the neck network of the target detector is improved, deep information of a construction site image can be better captured, the positioning capability of a small-size target is improved, and multi-target tracking of construction personnel can be rapidly and reliably realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target tracking, and in particular relates to a detection-based construction worker multi-target tracking method and system. Background Art

[0002] Positioning and tracking construction workers in construction scenarios is of great significance for real-time determination of the location and movement trajectory of construction workers, thereby timely discovering safety hazards and improving the level of supervision. In actual construction scenarios, the tracking of construction workers often relies on human eyes for supervision. This traditional method has problems such as low efficiency, insufficient accuracy and high cost, and is also a waste of manpower.

[0003] With the continuous development of deep learning technology, target tracking technology has gradually been applied to construction worker tracking tasks. However, existing methods are often insensitive to small targets and complex backgrounds, and their accuracy and speed need to be improved. A fast and reliable method for multi-target tracking of construction workers is needed. Summary of the invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a construction worker multi-target tracking method and system based on detection. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0005] The present invention provides a detection-based construction worker multi-target tracking method, comprising:

[0006] S1: constructing a target detector, training the target detector using a training data set, and obtaining a trained target detector, wherein the target detector includes an improved YOLOV8 model;

[0007] S2: Inputting the original video to be tracked into the trained target detector frame by frame to obtain the target detection result of each frame image;

[0008] S3: constructing a target tracker based on the BYTE association algorithm, wherein the input end of the target tracker is connected to the output end of the target detector, so as to process the target detection result of the target detector and obtain the target tracking result of the construction personnel.

[0009] In one embodiment of the present invention, the S1 includes:

[0010] S1.1: construct an image dataset of construction workers at a construction site and process the image dataset to obtain a training dataset;

[0011] S1.2: Constructing a target detector, the target detector includes a backbone network, a neck network and a head network connected in sequence, wherein the backbone network is used to extract basic features of different scales of the input image, the neck network is used to process and fuse the basic features of different scales to obtain feature maps of different scales after processing, and the head network is used to generate the final target detection result according to the feature maps of different scales after processing;

[0012] S1.3: Constructing a loss function corresponding to the target detector based on the bulldozer distance;

[0013] S1.4: Train the target detector using the training data set to obtain a trained target detector.

[0014] In one embodiment of the present invention, the backbone network includes a first convolution module Conv1, a first C-DCN unit, a second C-DCN unit, a third C-DCN unit, a fourth C-DCN unit and an SPPF module connected in sequence, wherein:

[0015] The first C-DCN unit, the second C-DCN unit, the third C-DCN unit and the SPPF module are all connected to the neck network, and are used to extract basic features of different scales of the input image and transmit them to the neck network.

[0016] In one embodiment of the present invention, the first C-DCN unit, the second C-DCN unit, the third C-DCN unit and the fourth C-DCN unit have the same structure, and are all formed by splicing a convolution module and a C2f-DCN module, wherein the C2f-DCN module is obtained by replacing the ordinary convolution in the bottleneck structure of the C2f module with a deformable convolution, and the calculation formula of the deformable convolution is:

[0017]

[0018] Where K is the total number of sampling points, k represents the kth sampling point, and w k is the feature map weight of the kth sampling point, m k represents the modulation scalar of the kth sampling point, p 0 is the center position of the convolution kernel, p k is the kth position of the predefined grid sample in the regular convolution, Δp k is the sampling offset of the deformable convolution, m k With Δp k are all learnable parameters.

[0019] In one embodiment of the present invention, the neck network includes a first upsampling module Upsample1, a first concatenation module Concat1, a first C-SPA unit, a second upsampling module Upsample2, a second concatenation module Concat2, a second C-SPA unit, a third upsampling module Upsample3, a third concatenation module Concat3, a third C-SPA unit, a second convolution module Conv2, a fourth concatenation module Concat4, a fourth C-SPA unit, a third convolution module Conv3, a fifth concatenation module Concat5, a fifth C-SPA unit, a fourth convolution module Conv4, a sixth concatenation module Concat6 and a sixth C-SPA unit, which are sequentially connected, wherein:

[0020] The output end of the first C-DCN unit is connected to the input end of the third splicing module Concat3, the output end of the second C-DCN unit is connected to the input end of the second splicing module Concat2, the output end of the third C-DCN unit is connected to the input end of the first splicing module Concat1, and the output end of the SPPF module is respectively connected to the input end of the first up-sampling module Upsample1 and the input end of the sixth splicing module Concat6;

[0021] The output end of the first C-SPA unit is connected to the input end of the fifth splicing module Concat5, and the output end of the second C-SPA unit is connected to the input end of the fourth splicing module Concat4;

[0022] Output ends of the third C-SPA unit, the fourth C-SPA unit, the fifth C-SPA unit and the sixth C-SPA unit are all connected to the head network, so as to respectively obtain processed feature maps of different scales and transmit them to the head network.

[0023] In one embodiment of the present invention, the first C-SPA unit, the second C-SPA unit, the third C-SPA unit, the fourth C-SPA unit, the fifth C-SPA unit and the sixth C-SPA unit have the same structure, and are all formed by splicing a C2f module and a SPA module, wherein:

[0024] The SPA module is a collaborative pooling attention mechanism. The SPA module is used to decompose the input data in the horizontal direction through global average pooling and global maximum pooling to obtain a horizontal average pooling tensor With horizontal max pooling tensor At the same time, the input data is decomposed in the vertical direction through global average pooling and global maximum pooling to obtain the vertical average pooling tensor With vertical max pooling tensor Concatenate to obtain horizontal and vertical concatenated tensors:

[0025]

[0026] Among them, X W is the horizontal concatenation tensor, X H is the vertical concatenation tensor;

[0027] Then we get the collaborative pooling tensor X′:

[0028] X′=δ(F(concat(X W ,X H )))

[0029] Among them, δ is a nonlinear activation operation and F is a convolution operation.

[0030] The output Y of the SPA module is expressed as:

[0031] Y=σ(F((X′) split ))*X

[0032] Among them, σ is the sigmoid activation function, and the split operation is used to divide the tensor into a width tensor and a height tensor.

[0033] In one embodiment of the present invention, the head network includes a first detection head Detect1, a second detection head Detect2, a third detection head Detect3 and a fourth detection head Detect4, wherein:

[0034] The output end of the third C-SPA unit is connected to the input end of the first detection head Detect1, the output end of the fourth C-SPA unit is connected to the input end of the second detection head Detect2, the output end of the fifth C-SPA unit is connected to the input end of the third detection head Detect3, and the output end of the sixth C-SPA unit is connected to the input end of the fourth detection head Detect4.

[0035] In one embodiment of the present invention, the loss function is:

[0036] L loc =IOU γ *WD Loss

[0037] Among them, IOU is the intersection-over-union ratio between the predicted bounding box and the true bounding box, γ is a hyperparameter used to balance the focal distance of unbalanced samples, and WDLoss is the bulldozer loss, which is calculated as follows:

[0038]

[0039] Among them, the hyperparameter α is the adjustment factor, C is a constant related to the data set, and W 2 is the bulldozer distance between the predicted bounding box and the true bounding box.

[0040] Another aspect of the present invention provides a detection-based construction worker multi-target tracking system, which is used to execute the construction worker multi-target tracking method described in any one of the above embodiments. The system includes a target detector and a target tracker connected to each other, wherein:

[0041] The target detector includes a trained improved YOLOV8 model, which is used to obtain the detection result of the target in each frame image of the video to be tracked;

[0042] The target tracker is based on the BYTE association algorithm and is used to process the detection results of the target detector to obtain the target tracking results of the construction personnel.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] 1. The detection-based multi-target tracking method and system for construction workers provided by the present invention improves the structure of the neck network of the target detector based on the collaborative pooling attention mechanism to better capture and utilize the deep information of the network, thereby improving the model accuracy; by additionally designing a small target detection head for the target detector, the detection capability of small-sized targets is improved; the Gaussian distribution distance is used to measure the similarity between frames, the loss function is improved, and the positioning capability of small-sized targets is further improved, thereby realizing fast and reliable multi-target tracking of construction workers.

[0045] 2. The detection-based multi-target tracking method and system for construction workers provided by the present invention designs a lightweight convolution module based on deformable convolution to improve the model structure of the target detector, reduce the computational complexity, improve the speed of the model, and enhance the real-time performance of the model detection.

[0046] 3. The present invention constructs a target tracker based on the BYTE association algorithm and connects it to the target detector, which improves the tracking efficiency and stability and avoids problems such as target loss and mismatch.

[0047] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flow chart of a method for multi-target tracking of construction personnel based on detection provided by an embodiment of the present invention;

[0049] Figure 2 is a schematic diagram of the structure of a target detector provided by an embodiment of the present invention;

[0050] Figure 3 is a schematic structural diagram of a C-DCN unit provided by an embodiment of the present invention;

[0051] Figure 4 It is a detection-based multi-target tracking framework provided by an embodiment of the present invention;

[0052] Figure 5 It is a visual display of three consecutive frames of tracking results obtained by using the method of an embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the following is a detailed description of a detection-based construction personnel multi-target tracking method and system proposed in accordance with the present invention in combination with the accompanying drawings and specific implementation methods.

[0054] The above and other technical contents, features and effects of the present invention are clearly presented in the following detailed description of the specific implementation modes in conjunction with the accompanying drawings. Through the description of the specific implementation modes, the technical means and effects adopted by the present invention to achieve the predetermined purpose can be more deeply and specifically understood. However, the attached drawings are only for reference and explanation purposes and are not used to limit the technical solutions of the present invention.

[0055] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants are intended to cover non-exclusive inclusion, so that an article or device including a series of elements includes not only those elements, but also other elements that are not explicitly listed. In the absence of more restrictions, the elements defined by the statement "including one..." do not exclude the existence of other identical elements in the article or device including the elements.

[0056] Embodiment 1

[0057] See also Figure 1 , Figure 1 1 is a flow chart of a method for tracking multiple targets of construction personnel based on detection provided by an embodiment of the present invention. The method for tracking multiple targets of construction personnel includes:

[0058] S1: Build a target detector and train the target detector using the training data set to obtain the trained target detector.

[0059] Step S1 of this embodiment includes:

[0060] S1.1: Construct an image dataset of construction workers at the construction site and process the image dataset to obtain a training dataset.

[0061] Specifically, step S1.1 further includes:

[0062] S1.11 collects construction site images and builds a construction worker image dataset at the construction site in combination with publicly available construction site image data.

[0063] S1.12: Use LabelImg software to label the construction workers in each image in the construction worker image dataset, and define the label as "person" to represent the target construction worker.

[0064] S1.13: Perform data augmentation on each image using methods such as rotation, cropping, exposure, and noise addition to increase the number of samples in the construction worker image dataset and improve the generalization ability of subsequent models.

[0065] S1.14: Divide the processed construction worker image dataset into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0066] In this embodiment, there are 7792 images in the construction personnel image dataset. After the above processing, the number of samples in the training set, validation set and test set is shown in Table 1.

[0067] Table 1. Number of samples in the training set, validation set, and test set

[0068] Subdataset Training set Validation set Test Set Sample size 5450 1561 781

[0069] S1.2: Build an object detector.

[0070] The target detector of this embodiment includes an improved YOLOV8 model, which includes a backbone network, a neck network and a head network connected in sequence, wherein the backbone network is used to extract basic features of different scales of the input image, the neck network is used to process and fuse the basic features of different scales to obtain feature maps of different scales after processing, and the head network is used to generate the final target detection result according to the feature maps of different scales after processing.

[0071] In this embodiment, the backbone network, neck network and head network of the original YOLOv8 model are designed, a small target detection head is added, and an improved model integrating deformable convolution and collaborative pooling attention is constructed for construction personnel detection; the improved YOLOv8 model is obtained by improving the backbone network, neck network and head network of the YOLOv8 model.

[0072] See also Figure 2 , Figure 2 is a schematic diagram of the structure of a target detector provided by an embodiment of the present invention. The target detector is an improved YOLOV8 model, including a backbone network 1, a neck network 2 and a head network 3 connected in sequence.

[0073] The backbone network 1 of this embodiment includes a first convolution module Conv1, a first C-DCN unit 101, a second C-DCN unit 102, a third C-DCN unit 103, a fourth C-DCN unit 104 and an SPPF module connected in sequence, wherein the first C-DCN unit 101, the second C-DCN unit 102, the third C-DCN unit 103 and the SPPF module are respectively connected to the neck network, for respectively extracting basic features of different scales of the input image and transmitting them to the neck network.

[0074] For further information, see Figure 3 , Figure 3 1 is a schematic diagram of the structure of a C-DCN unit provided by an embodiment of the present invention. The first C-DCN unit 101, the second C-DCN unit 102, the third C-DCN unit 103, and the fourth C-DCN unit 104 of this embodiment have the same structure, and are all spliced ​​by the convolution module conv and the C2f-DCN module, wherein the C2f-DCN module is obtained by replacing the ordinary convolution in the bottleneck structure of the C2f module with a deformable convolution, and the deformable convolution calculation formula is as follows:

[0075]

[0076] Where K is the total number of sampling points, k represents the kth sampling point, and w k is the feature map weight of the kth sampling point, m k represents the modulation scalar of the kth sampling point, p 0 is the center position of the convolution kernel, p k is the kth position of the predefined grid sample in the regular convolution, Δp k is the sampling offset of the deformable convolution. k With Δp k Both are learnable parameters and are updated during the back-propagation process of training.

[0077] The neck network 2 of this embodiment includes a first upsampling module Upsample1, a first concatenation module Concat1, a first C-SPA unit 201, a second upsampling module Upsample2, a second concatenation module Concat2, a second C-SPA unit 202, a third upsampling module Upsample3, a third concatenation module Concat3, a third C-SPA unit 203, a second convolution module Conv2, a fourth concatenation module Concat4, a fourth C-SPA unit 204, a third convolution module Conv3, a fifth concatenation module Concat5, a fifth C-SPA unit 205, a fourth convolution module Conv4, a sixth concatenation module Concat6 and a sixth C-SPA unit 206, which are sequentially connected. The output end of the first C-DCN unit is connected to the input end of the third splicing module Concat3, the output end of the second C-DCN unit is connected to the input end of the second splicing module Concat2, the output end of the third C-DCN unit is connected to the input end of the first splicing module Concat1, and the output end of the SPPF module is respectively connected to the input end of the first up-sampling module Upsample1 and the input end of the sixth splicing module Concat6; the output end of the first C-SPA unit 201 is connected to the input end of the fifth splicing module Concat5, and the output end of the second C-SPA unit 202 is connected to the input end of the fourth splicing module Concat4; the outputs of the third C-SPA unit 203, the fourth C-SPA unit 204, the fifth C-SPA unit 205 and the sixth C-SPA unit 206 are respectively connected to the head network, so as to obtain feature maps of different scales after processing and transmit them to the head network.

[0078] Furthermore, the first C-SPA unit 201, the second C-SPA unit 202, the third C-SPA unit 203, the fourth C-SPA unit 204, the fifth C-SPA unit 205 and the sixth C-SPA unit 206 in the neck network 2 of this embodiment have the same structure, and are all spliced ​​by the C2f module and the SPA module. The SPA module is a collaborative pooling attention mechanism. For a given input, the SPA module decomposes the input data in the horizontal direction through global average pooling and global maximum pooling to obtain a horizontal average pooling tensor With horizontal max pooling tensor At the same time, the input data is decomposed in the vertical direction through global average pooling and global maximum pooling to obtain the vertical average pooling tensor With vertical max pooling tensor Then concatenate them to obtain horizontal and vertical concatenation tensors:

[0079]

[0080] Among them, X W is the horizontal concatenation tensor, X H is the vertical concatenation tensor.

[0081] Then the collaborative pooling tensor X′ is obtained by the following formula:

[0082] X′=δ(F(concat(X W ,X H )))

[0083] Among them, δ is a nonlinear activation operation, and F is a convolution operation;

[0084] Finally, the output Y of the SPA module is calculated as follows:

[0085] Y=σ(F((X′) split ))*X

[0086] Among them, σ is the sigmoid activation function, and the split operation is used to divide the tensor into a width tensor and a height tensor.

[0087] Continue to see Figure 2 The head network 3 of this embodiment includes a first detection head Detect1, a second detection head Detect2, a third detection head Detect3 and a fourth detection head Detect4. The output end of the third C-SPA unit 203 is connected to the input end of the first detection head Detect1, the output end of the fourth C-SPA unit 204 is connected to the input end of the second detection head Detect2, the output end of the fifth C-SPA unit 205 is connected to the input end of the third detection head Detect3, and the output end of the sixth C-SPA unit 206 is connected to the input end of the fourth detection head Detect4.

[0088] The first detection head Detect1, the second detection head Detect2, the third detection head Detect3 and the fourth detection head Detect4 of this embodiment have the same structure, and the feature maps input to the four detection heads have different sizes in the length and width directions, which are 160×160, 80×80, 40×40 and 20×20 respectively. Among them, the first detection head Detect1 is an additional detection head for detecting small-sized targets.

[0089] S1.3: Construct the loss function corresponding to the target detector based on the bulldozer distance.

[0090] The training of the target detector requires the definition of a loss function. When defining the loss function, in order to focus on the positioning of small-sized targets, this embodiment is based on the bulldozer distance and introduces the focal distance and adjustment factor, and uses the Gaussian distribution distance to measure the similarity between the boxes (the predicted bounding box and the real bounding box) to construct an improved positioning loss L loc, the calculation formula of improved positioning loss is as follows:

[0091] L loc =IOU γ *WD Loss

[0092] Among them, IOU is the intersection-over-union ratio between the predicted bounding box and the true bounding box, the hyperparameter γ is the focal distance used to balance the unbalanced samples, and in this embodiment, the value is 0.8, and WDLoss is the bulldozer loss, which is calculated as follows:

[0093]

[0094] The hyperparameter α is the adjustment factor, which is 3 in this example. C is a constant related to the data set, which is 2 in this example. 2 To calculate the bulldozer distance between the predicted bounding box and the true bounding box, the formula is as follows:

[0095]

[0096] Among them, (x a ,y a ,w a ,h a ) and (x b ,y b ,w b ,h b ) represent the predicted bounding box and the true bounding box respectively, x a and a Respectively represent the horizontal and vertical coordinates of the center point of the predicted bounding box, w a and h a Respectively represent the width and length of the predicted bounding box, x b and b Respectively represent the horizontal and vertical coordinates of the center point of the real bounding box, w b and h b Represent the width and length of the true bounding box respectively.

[0097] S1.4: Train the target detector using the training data set to obtain a trained target detector.

[0098] After the loss function is defined, the training data set processed in step S1 is input into the input of the backbone network 1 of the target detector in batches, and the total loss is calculated through forward propagation, which is used as a guide for back propagation to optimize the internal weight parameters of the target detector. In this embodiment, 120 epochs of training are performed, the batch size is set to 24, and the initial learning rate is 10 -4 , momentum factor 0.9, and the learning rate dropped to 10 at 80 epochs-5 .

[0099] S2: Input the original video to be tracked into the trained target detector frame by frame to obtain the target detection result of each frame image.

[0100] In the actual target detection process, after the training is completed, the best weight file best.pt is obtained, which is loaded into the improved YOLOV8 model to obtain the best target detector. The original image to be detected (the image in the test set is used in this embodiment) is input into the input end of the target detector, and the bounding box, category label and confidence of the target construction worker in the image can be obtained at the output end.

[0101] In this embodiment, the video to be processed is cut into multiple frames of images at predetermined intervals, and the multiple frames of images are sequentially input into the trained target detector to obtain the detection results of the construction workers in each frame of the image, including the bounding box, category label and confidence.

[0102] S3: A target tracker is constructed based on the BYTE association algorithm. The input end of the target tracker is connected to the output end of the target detector to process the target detection result of the target detector to obtain the target tracking result of the construction personnel.

[0103] In this embodiment, step S3 specifically includes the following steps:

[0104] S3.1: For the target detection result of each frame of the video to be processed, the confidence level is greater than or equal to the upper threshold θ h The bounding box is attributed to the high-score bounding box Bh. In this embodiment, the upper threshold θ is taken h =0.6; the confidence level is lower than the upper threshold θ h and is greater than or equal to the lower threshold θ l The low score bounding box Bl of the bounding box, in this embodiment, takes the threshold θ l =0.2; the confidence level is lower than the lower threshold θ l In the initial stage of tracking, if there is a bounding box with a confidence level higher than the set threshold o and the same target is detected in two consecutive frames, a new trajectory is generated using the position of the target and added to the trajectory set T. In the subsequent tracking process, the Kalman filter is used to predict all the trajectories in the trajectory set T, and the new position of each trajectory in the trajectory set T in the current frame is predicted, and the updated new trajectory is obtained.

[0105] S3.2: Perform the first association: For the high-scoring bounding box Bh, use the Hungarian algorithm to associate and match it with the new trajectory updated after the Kalman filter prediction, and obtain the similarity between the high-scoring bounding box Bh and the updated new trajectory. The similarity threshold of 0.2 is used as the matching standard. If the obtained similarity is less than 0.2, it means that there is no match. The unmatched trajectory is stored in the set Tremain, and the unmatched bounding box is stored in the set Bremain.

[0106] S3.3: Perform a second association: Perform a second association between the low-score bounding box Bl and the unmatched trajectory Tremain based on the IOU calculation, store the unmatched trajectory in the set Tre-remain, and remove all unmatched low-score bounding boxes Bl.

[0107] S3.4: Treat the set Tre-remain as the trajectory of the temporarily lost target and store it in the subset Tlost of the trajectory set T. If there is still no match after saving 30 frames, the trajectory will be deleted.

[0108] S3.5: For the unmatched bounding box Bremain outputted by the first association, if its confidence is higher than the set threshold σ and the same target is detected in two consecutive frames, it is used to initialize and generate a new trajectory. In this embodiment, the threshold σ=0.5.

[0109] After the above steps, the action trajectory of each construction worker in the video clip to be detected is finally obtained.

[0110] After building the target tracker, such as Figure 4 As shown, the output end of the target detector is connected to the input end of the target tracker to form a detection-based multi-target tracking method for construction workers. For the video to be processed, the target detector detects it frame by frame, and the target tracker generates a tracking trajectory in real time according to the detection results to realize multi-target tracking of construction workers in the construction scene.

[0111] In order to verify the effectiveness of the multi-target tracking method for construction workers based on detection in the present invention, a video of a construction site is used as a test sample, and the multi-target tracking method constructed by the present invention is input, and compared with the existing SORT algorithm and DeepSort algorithm. MOTA, IDs, and FPS are selected as model evaluation indicators, where MOTA is the tracking accuracy, IDs is the total number of ID switches, and FPS is the number of frames transmitted per second. The test results are shown in Table 2.

[0112] Table 2. Comparison of results of different methods

[0113] method MOTA / % IDs <![CDATA[FPS / (frame * s -1 )]]> SORT 83.2 23 59 DeepSort 85.1 19 63 Method of the present invention 93.5 11 72

[0114] It can be seen from Table 2 that the method proposed in the present invention achieves higher tracking accuracy and the lowest number of ID switching times, and the tracking precision is improved; a higher FPS is achieved, and the tracking speed is improved.

[0115] See also Figure 5 , Figure 5 It is a visualization display of three consecutive frames of tracking results obtained by using the method of the embodiment of the present invention. As can be seen from the figure, the method of the present invention can accurately locate the target appearing in the video and continuously track it. It also has high performance for multiple targets, small-sized targets, etc., and can support the multi-target tracking task of construction workers in construction scenes.

[0116] Embodiment 2

[0117] Based on Example 1, this embodiment provides a detection-based multi-target tracking system for construction workers, including a target detector and a target tracker connected to each other. The target detector includes a trained improved YOLOV8 model for obtaining target detection results for each frame image in the original video to be tracked; the target tracker is based on a BYTE association algorithm, and is used to process the target detection results of the target detector to obtain target tracking results of the construction workers.

[0118] The detection-based construction worker multi-target tracking method and system provided by the present invention improves the detection capability of small-sized targets by additionally designing a small target detection head for the target detector; based on the collaborative pooling attention mechanism, the structure of the neck network is improved to better capture and utilize the deep information of the network, thereby improving the model accuracy; based on the deformable convolution, a lightweight convolution module is designed to improve the model structure, reduce the computational complexity, and increase the model speed; the Gaussian distribution distance is used to measure the similarity between frames, the loss function is improved, and the positioning capability of small-sized targets is further improved, thereby realizing fast and reliable multi-target tracking of construction workers.

[0119] The present invention designs a lightweight convolution module based on deformable convolution to improve the model structure of the target detector, reduces the computational complexity, improves the speed of the model, and enhances the real-time performance of the model detection; based on the BYTE association algorithm, a target tracker is constructed and connected to the target detector, which improves the tracking efficiency and stability and avoids problems such as target loss and mismatch.

[0120] In the several embodiments provided by the present invention, it should be understood that the system and method disclosed by the present invention can be implemented in other ways. For example, the system embodiment described above is only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0121] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0122] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A detection-based construction worker multi-target tracking method, characterized in that: include: S1: constructing a target detector, training the target detector using a training data set, and obtaining a trained target detector, wherein the target detector includes an improved YOLOV8 model; S2: Inputting the original video to be tracked into the trained target detector frame by frame to obtain the target detection result of each frame image; S3: constructing a target tracker based on the BYTE association algorithm, wherein the input end of the target tracker is connected to the output end of the target detector, so as to process the target detection result of the target detector and obtain the target tracking result of the construction personnel.

2. The detection-based construction worker multi-target tracking method according to claim 1 is characterized in that: The S1 includes: S1.1: construct an image dataset of construction workers at a construction site and process the image dataset to obtain a training dataset; S1.2: Constructing a target detector, the target detector includes a backbone network, a neck network and a head network connected in sequence, wherein the backbone network is used to extract basic features of different scales of the input image, the neck network is used to process and fuse the basic features of different scales to obtain feature maps of different scales after processing, and the head network is used to generate the final target detection result according to the feature maps of different scales after processing; S1.3: Constructing a loss function corresponding to the target detector based on the bulldozer distance; S1.4: Train the target detector using the training data set to obtain a trained target detector.

3. The detection-based construction worker multi-target tracking method according to claim 2 is characterized in that: The backbone network includes a first convolution module Conv1, a first C-DCN unit, a second C-DCN unit, a third C-DCN unit, a fourth C-DCN unit and an SPPF module connected in sequence, wherein: The first C-DCN unit, the second C-DCN unit, the third C-DCN unit and the SPPF module are all connected to the neck network, and are used to extract basic features of different scales of the input image and transmit them to the neck network.

4. The detection-based construction worker multi-target tracking method according to claim 3 is characterized in that: The first C-DCN unit, the second C-DCN unit, the third C-DCN unit and the fourth C-DCN unit have the same structure, and are all formed by splicing a convolution module and a C2f-DCN module, wherein the C2f-DCN module is obtained by replacing the ordinary convolution in the bottleneck structure of the C2f module with a deformable convolution, and the calculation formula of the deformable convolution is: Where K is the total number of sampling points, k represents the kth sampling point, and w k is the feature map weight of the kth sampling point, m k represents the modulation scalar of the kth sampling point, p0 is the center position of the convolution kernel, and p k is the kth position of the predefined grid sample in the regular convolution, Δp k is the sampling offset of the deformable convolution, m k With Δp k These are all learnable parameters.

5. The detection-based construction worker multi-target tracking method according to claim 3 is characterized in that: The neck network includes a first upsampling module Upsample1, a first concatenation module Concat1, a first C-SPA unit, a second upsampling module Upsample2, a second concatenation module Concat2, a second C-SPA unit, a third upsampling module Upsample3, a third concatenation module Concat3, a third C-SPA unit, a second convolution module Conv2, a fourth concatenation module Concat4, a fourth C-SPA unit, a third convolution module Conv3, a fifth concatenation module Concat5, a fifth C-SPA unit, a fourth convolution module Conv4, a sixth concatenation module Concat6 and a sixth C-SPA unit, which are sequentially connected, wherein: The output end of the first C-DCN unit is connected to the input end of the third splicing module Concat3, the output end of the second C-DCN unit is connected to the input end of the second splicing module Concat2, the output end of the third C-DCN unit is connected to the input end of the first splicing module Concat1, and the output end of the SPPF module is respectively connected to the input end of the first up-sampling module Upsample1 and the input end of the sixth splicing module Concat6; The output end of the first C-SPA unit is connected to the input end of the fifth splicing module Concat5, and the output end of the second C-SPA unit is connected to the input end of the fourth splicing module Concat4; Output ends of the third C-SPA unit, the fourth C-SPA unit, the fifth C-SPA unit and the sixth C-SPA unit are all connected to the head network, so as to respectively obtain processed feature maps of different scales and transmit them to the head network.

6. The detection-based construction worker multi-target tracking method according to claim 5, characterized in that: The first C-SPA unit, the second C-SPA unit, the third C-SPA unit, the fourth C-SPA unit, the fifth C-SPA unit and the sixth C-SPA unit have the same structure, and are all formed by splicing a C2f module and a SPA module, wherein: The SPA module is a collaborative pooling attention mechanism. The SPA module is used to decompose the input data in the horizontal direction through global average pooling and global maximum pooling to obtain a horizontal average pooling tensor With horizontal max pooling tensor At the same time, the input data is decomposed in the vertical direction through global average pooling and global maximum pooling to obtain the vertical average pooling tensor With vertical max pooling tensor Concatenate to obtain horizontal and vertical concatenated tensors: Among them, X W is the horizontal concatenation tensor, X H is the vertical concatenation tensor; Then we get the collaborative pooling tensor X′: X′=δ(F(concat(X W ,X H ))) Among them, δ is a nonlinear activation operation and F is a convolution operation. The output Y of the SPA module is expressed as: Y=σ(F((X′) split ))*X Among them, σ is the sigmoid activation function, and the split operation is used to divide the tensor into a width tensor and a height tensor.

7. The detection-based construction worker multi-target tracking method according to claim 5, characterized in that: The head network includes a first detection head Detect1, a second detection head Detect2, a third detection head Detect3 and a fourth detection head Detect4, wherein: The output end of the third C-SPA unit is connected to the input end of the first detection head Detect1, the output end of the fourth C-SPA unit is connected to the input end of the second detection head Detect2, the output end of the fifth C-SPA unit is connected to the input end of the third detection head Detect3, and the output end of the sixth C-SPA unit is connected to the input end of the fourth detection head Detect4.

8. The detection-based construction worker multi-target tracking method according to claim 2, characterized in that: The loss function is: L loc =IOU γ *WDLoss Among them, IOU is the intersection-over-union ratio between the predicted bounding box and the true bounding box, γ is a hyperparameter used to balance the focal distance of unbalanced samples, and WDLoss is the bulldozer loss, which is calculated as follows: Among them, the hyperparameter α is the adjustment factor, C is a constant related to the dataset, and W2 is the bulldozer distance between the predicted bounding box and the true bounding box.

9. A detection-based construction worker multi-target tracking system, characterized in that: The system is used to execute the construction worker multi-target tracking method according to any one of claims 1 to 8, wherein the system comprises a target detector and a target tracker connected to each other, wherein: The target detector includes a trained improved YOLOV8 model, which is used to obtain the target detection result of each frame image in the original video to be tracked; The target tracker is based on the BYTE association algorithm and is used to process the target detection result of the target detector to obtain the target tracking result of the construction personnel.