A multi-category building material video counting method, system and counting device

Through the robot video counting method, the number of building materials is automatically calculated using the YOLOv4 model and the double-line algorithm, which solves the problems of cumbersome manual counting and large errors on the construction site, and achieves efficient and accurate statistics on the quantity of building materials.

CN115171011BActive Publication Date: 2025-07-22CHINA UNIV OF GEOSCIENCES (WUHAN) +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210756710.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-07-22
Estimated Expiration
2042-06-30

Smart Images

  • Figure CN115171011B_ABST
    Figure CN115171011B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-category building material video counting method, system and counting device. The counting method includes: extracting a to-be-detected image of a video captured by a robot; inputting the to-be-detected image into a YOLOv4 model to extract features of the to-be-detected image; after performing three convolutions on the last feature layer of the backbone feature extraction network, using multi-scale max pooling processing to separate the context features in the to-be-detected image; performing multi-scale prediction on the obtained features, and decoding to obtain the positions of the prediction boxes in the to-be-detected image; inputting all box information into an NMS module to obtain the filtered box information; inputting the box coordinate sequences of the front and rear frames in the frame sequence output by the target detector into a sort tracking module to output the inter-frame target ids. The present invention adopts a neural network method and uses a multi-category multi-object tracking to associate the inter-frame information of the video, overcome target occlusion, and finally calculates the quantity and types of building materials in the entire video through a double-crossing counting algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a multi-category building material video counting method, system, and counting device. Background Art

[0002] With the proposal of the concept of "digital construction site", robot intelligent monitoring technology has been widely applied in the construction industry, gradually realizing the inspection of building materials at the construction site, the detection of the quantity of building materials, and the real-time feedback of the demand for building materials, so as to reduce the occurrence of construction site accidents and improve the implementation efficiency of the construction industry.

[0003] Currently, after building materials enterprises transport building materials to the construction site by transport vehicles, generally, three parties of construction site personnel, namely the supplier, the labor team material clerk, and the project department material clerk, need to count the quantity of building materials to complete the goods acceptance. And the construction site generally adopts the manual counting method. For example, different colored pigments or electronic automatic counting pens are generally used to distinguish and mark the building materials to be counted.

[0004] Although the manual counting method is simple, the work intensity is large, the counting process is cumbersome and boring, and the staff will be in a highly tense state for a long time, which is likely to lead to counting errors. In addition, the whole process often needs to be repeatedly proofread. It generally takes about several hours for workers to count the building materials, and the counting efficiency is very low, which can no longer meet the needs of rapid production of modern construction enterprises. Summary of the Invention

[0005] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides a multi-category building material video counting method, system, and counting device to solve the technical problems of large work intensity, cumbersome counting process, easy counting errors, and low work efficiency caused by the current manual counting method at the construction site.

[0006] To solve the above problems, the first object of the present invention is to provide a multi-category building material video counting method, which is applied to the estimation of the quantity of building materials at the construction site. The video counting method includes:

[0007] S 100 : Extract the image to be measured in the video captured by the robot;

[0008] S 200 : Input the image to be measured in the captured video into the YOLOv4 model, and extract the features of the image to be measured through the backbone feature extraction network CSPDarknet53;

[0009] S 300 : After performing three convolutions on the last feature layer of the backbone feature extraction network CSPdarknet53, process it respectively by using the maximum pooling method of multiple different scales to separate the most significant context features in the image to be measured;

[0010] S 400 : After extracting the features, use YOLOv3Head to perform multi-scale prediction on the obtained features to obtain the prediction results of 3 effective feature layers. The positions of the prediction boxes in the image to be measured are obtained by decoding the 3 effective feature layers.

[0011] S 500 : Input all the box information output by the prediction head into the NMS module to obtain the filtered box information.

[0012] S 600 : Input the box coordinate sequences of the front and rear frames in the frame sequence output by the target detector into the sort tracking module, and the sort tracking module outputs the target IDs between frames.

[0013] S 700 : Calculate the number of building material targets in the video through the double-crossing algorithm and print it in the output video.

[0014] Optionally, in step S 200 : The specific operation of extracting the features of the image to be measured is as follows:

[0015] Extract 3 effective feature layers (76, 76, 256), (38, 38, 512), and (19, 19, 1024) in the image to be measured. The 3 effective feature layers are located at different positions in the backbone feature extraction network CSPDarknet53 for detecting small, medium, and large targets to be measured respectively.

[0016] Optionally, in step S 300 : After performing three DarknetConv2D_BN_Leaky convolutions on the last output feature layer in the backbone feature extraction network CSPDarknet53, process it using max pooling kernels of four different scales (13, 13), (9, 9), (5, 5), and (1, 1) respectively to improve the receptive field size and separate the most significant context features.

[0017] Optionally, in step S 400 : The specific operation of using YOLOv3Head to perform multi-scale prediction on the obtained features includes:

[0018] Use YOLOv3Head to perform multi-scale prediction on the obtained features to obtain the prediction results of 3 effective feature layers, thereby outputting three encoded tensor values of (19, 19, 33), (38, 38, 33), and (76, 76, 33), and the positions of the three prediction boxes can be determined.

[0019] Obtain the coordinates of (19 * 19 + 38 * 38 + 76 * 76) * 3 boxes, and the coordinate structure is [x, y, w, h, confidence, class1, class2, …, class N]

[0020] Where: x and y represent the upper left coordinates of each prior box, w and h represent the width and height of the prior box respectively, confidence represents the confidence that the network determines the prior box belongs to class N, and class N represents N classes.

[0021] Optionally, in step S 500 The operation of inputting all the box information output by the prediction head into the NMS module to obtain the filtered box information specifically includes:

[0022] After obtaining several boxes from the yolov4 network, input the array containing the box information into the NMS module for non-maximum suppression, and output the final detection result.

[0023] Optionally, in step S 600 The operation that the sort tracking module outputs the inter-frame target id by inputting the box coordinate sequences of the front and rear frames in the frame sequence output by the target detector is as follows:

[0024] Input the box matrix filtered by the NMS module into the sort tracking module, and the sort tracking module assigns an id to all the targets in the current frame to determine whether the targets in two frames are the same target.

[0025] Optionally, in step S 700 The operation of calculating the number of building material targets in the video by the double-crossing line algorithm specifically includes:

[0026] S 701 : Lock whether the front and rear frames are the same target by the assigned id;

[0027] S 702 : Connect the box center coordinates of the current frame of each target with the center coordinates of the previous frame to form a vector;

[0028] S 703 : Judge the vector direction of each frame to determine which counting line of the double-crossing line it is. If the vector intersects the counting line, the target number is incremented by one.

[0029] Optionally, the loss function of the YOLOv3Head network includes coordinate loss coordError, confidence loss IOUError, and class prediction loss classError. The expression of the loss function of the YOLOv3Head network is as follows:

[0030]

[0031] Wherein: Indicates that the i-th cell contains a target. Indicates that the j-th bounding box of the i-th cell contains a target. Indicates that the j-th bounding box of the i-th cell does not contain a target, λ coord Indicates the weight value of the box regression loss, λ noobj Indicates the weight value occupied by the class without a target. Indicates the confidence that the predicted target is the i-th class, C i Represents the true confidence of the i-th class. Represents the probability of being predicted as the i-th class, p i (c) represents the true probability of the i-th class, and x, y, w, h respectively represent the center x, y coordinates of the predicted box and the width and height of the box.

[0032] The second object of the present invention is to provide a multi-category building material video counting device, including: a processor, a display, a memory, and computer program instructions stored on the memory and executable on the processor. When the processor executes the computer program instructions, it is used for the above-mentioned multi-category building material video counting method.

[0033] The third object of the present invention is to provide a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the multi-category building material video counting method as described above.

[0034] Compared with the prior art, the present invention has remarkable advantages and beneficial effects, which are specifically reflected in the following aspects:

[0035] The present invention proposes a method for counting building materials in a construction site video captured by a robot through an algorithm. This method uses the neural network method in deep learning, automatically detects the types and positions of building materials in each frame of the video by using a computer, and uses a multi-category multi-object tracking to associate the inter-frame information of the video and overcome target occlusion; finally, calculates the quantity and types of building materials in the entire video through a double-crossing line counting algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic flow chart of the multi-category building material video counting method in an embodiment of the present invention;

[0037] Figure 2 It is a schematic structural diagram of the multi-category building material video counting device in an embodiment of the present invention;

[0038] Figure 3Schematic diagram of the BLSTM in the embodiment of the present invention;

[0039] Figure 4 Schematic diagram of the confidence module in the embodiment of the present invention;

[0040] Figure 5 Schematic diagram of the PAN network in the embodiment of the present invention;

[0041] Figure 6 Effect diagram of the algorithm part of the multi-category building material video counting method in the first embodiment of the present invention;

[0042] Figure 7 Effect diagram of the algorithm part of the multi-category building material video counting method in the second embodiment of the present invention;

[0043] Figure 8 Effect diagram of the algorithm part of the multi-category building material video counting method in the third embodiment of the present invention;

[0044] Figure 9 Effect diagram of the algorithm part of the multi-category building material video counting method in the fourth embodiment of the present invention;

[0045] Figure 10 Effect diagram of the fourth part of the algorithm of the multi-category building material video counting method in the fifth embodiment of the present invention;

[0046] Figure 11 Effect diagram of the fourth part of the algorithm of the multi-category building material video counting method in the sixth embodiment of the present invention. Detailed implementation manners

[0047] To make the above objects, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.

[0048] Please refer to Figures 1-5 As shown, in the embodiment of the present invention, a multi-category building material video counting method is provided, which is applied to estimating the quantity of building materials at a construction site. The video counting method includes:

[0049] S 100 : Extract the image to be measured in the video captured by the robot;

[0050] Specifically, in this embodiment, the specific operation of extracting the image to be measured is: Compress the video captured by the robot from 1920*1080 per frame to 416*416 per frame, aiming to match the input dimension of the network.

[0051] S 200: Input the image to be measured in the video frame into the YOLOv4 network model, and extract the features of the image to be measured through the backbone feature extraction network CSPDarknet53;

[0052] Specifically, in this embodiment, the image to be measured is input into the backbone part of the YOLOv4 network model, and features of three different scales are extracted. The dimensions of the features of the three different scales are (19*19*1024), (38*38*512), and (76*76*256) respectively.

[0053] S 300 : After performing three convolutions on the last feature layer of the backbone feature extraction network CSPdarknet53, use multiple different-scale max pooling methods for processing to separate the most significant context features in the image to be measured;

[0054] It should be noted specifically that the structure of the SPP module in the YOLOv4 network model is as Figure 2 shown. The outputs of the backbone network go through 4 different-scale max pooling (Max Pooling) operations respectively. The pooling kernel sizes of the max pooling operations are 1*1 (no processing), 5*5, 9*9, and 13*13. Then, the feature maps of different scales are concatenated (Concat). The SPP module can generate an image with a fixed size from images of different sizes, greatly increasing the receptive field, separating the most significant context features, and playing a role in feature enhancement.

[0055] S 400 : After extracting the features, use YOLOv3Head to perform multi-scale prediction on the obtained features to obtain the prediction results of 3 effective feature layers. The positions of the prediction boxes in the image to be measured are obtained through decoding of the 3 effective feature layers;

[0056] During the multi-scale prediction process, the repeated extraction and fusion of features by the PANet module is an important method for multi-scale feature extraction.

[0057] Please refer to Figure 5 shown. The PANet module mainly includes two sub-modules, FPN and PAN. Based on the semantic features extracted by the neural network, the FPN sub-module performs a series of upsampling (Up Samping) to transfer the rich semantic information of the deep network to the shallow network; then, feature fusion is achieved through lateral connections (Lateral Connection) at the corresponding feature scales; the PAN sub-module transfers the localization information of the shallow network to the deep network through a series of downsampling (DownSamping); and then feature fusion is performed again.

[0058] Thus, through two feature pyramid operations, the PANet module fuses the strong semantic information conveyed by the FPN sub-module and the strong localization features conveyed by the PAN sub-module at the corresponding detection layers, enabling the simultaneous acquisition of accurate localization information and rich semantic information in both the shallow and deep networks. This leads to a dual improvement in localization accuracy and semantic information, enhancing the model's detection ability for different targets.

[0059] S 500 : Input all the box information output by the prediction head into the NMS module to obtain the filtered box information;

[0060] S 600 : Input the box coordinate sequences of the front and back frames in the frame sequence output by the object detector into the sort tracking module, and the sort tracking module outputs the object IDs between frames;

[0061] Thus, by adding the sort tracking module, the sort tracking module solves the single-frame nature of video counting and adds the function of global counting on the basis of real-time counting. It can not only predict the number of targets in the current frame but also predict the total number of all targets from the start of the video to the current frame, providing great convenience for counting building materials at the construction site.

[0062] Here, the sort tracking module is specifically described as follows:

[0063] Input the result sequence of a series of boxes obtained from the detector into a prediction model. Here, we use the Kalman filter as the prediction model, which is independent of other objects and the movement of the camera that captures the objects. The state of each object is modeled as:

[0064]

[0065] where: u and v represent the x and y coordinates of the center of the object, and s and r represent the size (area) and aspect ratio of the bounding box. The aspect ratio here is fixed, so the aspect ratios of the front and back frames are the same.

[0066] represents the coordinates of the predicted center of the next frame and the area of the detection box. The bounding box is used to update the object state, and the velocity component is solved using the Kalman filter. If there is no detection box associated with the object, a linear prediction model is used without correction.

[0067] When assigning detection boxes to existing objects, the shape of the bounding box for each object is estimated by predicting its new position in the current frame. Then, the assignment cost matrix is calculated, which is used as the intersection over union (IOU) between the object and the detection box. If the IOU is less than a certain threshold, the assignment of the detection box is rejected.

[0068] The target with the assigned detection frame is considered to be tracked successfully, and an id is assigned to it. If the ids of the targets in the previous and next frames are the same, they are considered to be the same target.

[0069] S 700 :The target number of building materials in the video is calculated through the double-crossing algorithm and printed in the output video.

[0070] In this embodiment, the movement of all target frames of two adjacent frames is counted. If the number of frames in one direction is greater than the number of frames in another direction, the frame is determined to be a frame moving in this direction; then it is determined whether all the current left-moving frames are greater than the right-moving frames. If greater, the counting result is counted according to the right line, otherwise, it is counted according to the left line.

[0071] Therefore, by adding a double-line counting strategy, this strategy can solve the counting error of the single-line counting strategy caused by the uncertain lens movement direction. The double-line counting strategy can adaptively determine the counting strategy according to the movement direction of the camera, greatly improving the counting accuracy.

[0072] In addition, the training data set used is photos taken by the camera carried by the construction site inspection robot.

[0073] It needs to be further explained here that the backbone network of YOLOv4 is CSPDarknet53, which adds a cross-stage primary network (CSPNet) based on the backbone network Darknet53 of YOLOv3.

[0074] Darknet53 is a fully convolutional network that uses a large number of residual connections (Resunit) and uses convolution operations with stride=2 instead of pooling layers for downsampling, which speeds up the operation while ensuring network performance.

[0075] See also Figure 2 As shown in the figure, the cross-stage primary network CSPNet mainly solves the problem of excessive computation caused by deep networks. The cross-stage primary network CSPNet first divides the feature map of the base layer into two parts. One part is residually connected to alleviate the gradient explosion and overfitting problems, and the other part is skipped to reduce computation. Then they are merged through skip connections to speed up training.

[0076] Specifically, in the embodiment of the present invention, in step S 200 In the above, the specific operation of extracting the features of the image to be tested is:

[0077] Extract three effective feature layers (76, 76, 256), (38, 38, 512), and (19, 19, 1024) from the image to be measured. The three effective feature layers are located at different positions of the backbone feature extraction network CSPDarknet53 respectively, and are used to detect small, medium, and large objects to be measured.

[0078] Specifically, in the embodiment of the present invention, step S 300 In this step, after performing three DarknetConv2D_BN_Leaky convolutions on the last output feature layer in the backbone feature extraction network CSPDarknet53, it is processed respectively using max-pooling kernels of four different scales (13, 13), (9, 9), (5, 5), and (1, 1) to improve the size of the receptive field and separate the most significant context features.

[0079] Therefore, by processing with max-pooling kernels of four different scales, the purpose is to significantly improve the size of the receptive field and separate the most important context features.

[0080] Specifically, in the embodiment of the present invention, step S 400 In this step, the specific operations of performing multi-scale prediction on the obtained features using YOLOv3Head include:

[0081] Performing multi-scale prediction on the obtained features using YOLOv3Head to obtain the prediction results of the three effective feature layers, thereby outputting three encoded tensor values of (19, 19, 33), (38, 38, 33), and (76, 76, 33), and the positions of the three prediction boxes can be determined.

[0082] Obtain the coordinates of (19 * 19 + 38 * 38 + 76 * 76) * 3 boxes, and the coordinate structure is [x, y, w, h, confidence, class1, class2,..., class N]

[0083] Among them: x and y represent the upper left coordinates of each prior box, w and h represent the width and height of the prior box respectively, confidence represents the confidence that the network determines the prior box belongs to class N, and class N represents N categories.

[0084] The classification regression layer mainly completes the object detection tasks at different scales. The feature map is divided into grids in three different ways to detect objects at different scales.

[0085] Among them, the three different grid divisions are as follows:

[0086] Each grid area of the 13×13 grid division is the largest and is used to predict large objects;

[0087] Each grid in the 26×26 grid division has a medium size and is used to predict medium-sized objects;

[0088] Each grid in the 52×52 grid division has the smallest size and is used to predict small objects.

[0089] After obtaining the prior boxes at three scales, the model further obtains the category to which the target belongs through the regression loss function and the classification loss function, returns the bounding box of the target, and obtains the final detection result.

[0090] Specifically, in the embodiments of the present invention, in step S 500 The operation of inputting all the box information output by the prediction head into the NMS module to obtain the filtered box information specifically includes:

[0091] After obtaining several boxes from the YOLOv4 network model, the array containing the box information is input into the NMS module for non-maximum suppression, and the final detection result is output.

[0092] Specifically, in the embodiments of the present invention, step S 600 The operation of inputting the box coordinate sequences of the front and rear frames in the frame sequence output by the target detector into the sort tracking module, and the specific operation of the sort tracking module for outputting the target id between frames is:

[0093] The box matrix filtered by the NMS module is input into the sort tracking module, and the sort tracking module assigns an id to all the targets in the current frame to determine whether the targets in two frames are the same target.

[0094] Specifically, in the embodiments of the present invention, step S 700 The operation of calculating the number of building material targets in the video through the double-crossing line algorithm specifically includes:

[0095] S 701 : Lock whether the front and rear frames are the same target through the assigned id;

[0096] S 702 : Connect the center coordinates of the box of each target in the current frame to the center coordinates of the previous frame to form a vector;

[0097] S 703 : Judge the direction of the vector of each frame to determine which counting line of the double-crossing line it is. If the vector intersects the counting line, the target number is incremented by one.

[0098] Specifically, in the embodiments of the present invention, the loss function of the YOLOv3Head network includes coordinate loss coordError, confidence loss IOUError, and class prediction loss classError. The expression of the loss function of the YOLOv3Head network is as follows:

[0099]

[0100] Where: Indicates that the i-th cell contains a target; Indicates that the j-th bounding box in the i-th cell contains a target; Indicates that the j-th bounding box in the i-th cell does not contain a target. λ coord Represents the weight value of the box regression loss, λ noobj Represents the weight value of the classes without targets, Represents the confidence that the predicted target is the i-th class, C i Represents the true confidence of the i-th class. Represents the probability predicted to be the i-th class, p i (c) represents the true probability of the i-th class. x, y, w, h respectively represent the center x, y coordinates of the predicted box and the width and height of the box.

[0101] Please refer to Table 1 below. In the embodiments of the present invention, the counting metrics are as shown in the following table:

[0102]

[0103] Table 1

[0104] Note: * in the table indicates that the video contains steel bars and steel rings.

[0105] Figures 6-11 This is the effect diagram of the algorithm part in the embodiments of the present invention. By using the method of counting building materials in the site through the algorithm for the video of construction site building materials captured by the robot, and adopting the neural network method in deep learning, it solves the problem of automatically detecting the types and positions of building materials in each frame of the video by computer, and uses a multi-class multi-object tracking to associate the inter-frame information of the video, overcome target occlusion, and finally calculates the quantity and types of building materials in the whole video through the double-crossing counting algorithm.

[0106] Please refer to Figure 2 As shown, the embodiments of the present invention also provide a multi-class building material video counting device, including: a processor, a display, a memory, and computer program instructions stored on the memory and executable on the processor. When the processor executes the computer program instructions, it is used to implement the above-mentioned multi-class building material video counting method.

[0107] The video counting device provided by the embodiment of the present application can be used to execute the multi-category building material video counting method provided by any of the above method embodiments. The implementation principle and technical effects are similar and will not be elaborated here.

[0108] The embodiment of the present invention also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions run on a computer, the computer is enabled to execute the multi-category building material video counting method described above.

[0109] It should be noted that the above computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0110] Optionally, the readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application-specific integrated circuit. Of course, the processor and the readable storage medium can also exist as discrete components in the device.

[0111] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Without departing from the spirit and scope of the present disclosure, those skilled in the art can make various changes and modifications, and these changes and modifications will all fall within the protection scope of the present invention.

Claims

1. A multi-category building material video counting method, applied to the estimation of the quantity of building materials at a construction site, characterized in that, The video counting method includes: S 100 : Extract the image to be measured from the video captured by the robot; S 200 : Input the image to be measured in the captured video into the YOLOv4 model, and extract the features of the image to be measured through the backbone feature extraction network CSPDarknet53; S 300 : After performing three convolutions on the last feature layer of the backbone feature extraction network CSPdarknet53, it is processed using multiple max pooling methods with different scales respectively to isolate the most prominent context features in the image to be tested. S 400 : After extracting the features, YOLOv3Head is used to perform multi-scale prediction on the obtained features to obtain the prediction results of 3 effective feature layers. The positions of the prediction boxes in the image to be measured are obtained by decoding the 3 effective feature layers. S 500 : Input all the box information output by the prediction head into the NMS module to obtain the filtered box information; S 600 : Input the box coordinate sequences of the front and rear frames in the frame sequence output by the target detector into the sort tracking module, and the sort tracking module outputs the target IDs between frames; S 700 : Calculate the number of building material targets in the video through the double-crossing algorithm and print it in the output video; Among them, the specific process of calculating the number of building material targets in the video by the double crossing line algorithm includes: S 701 : Lock whether the front and back frames are the same target by the assigned ID; S 702 : Connect the box center coordinates of the current frame of each target to the center coordinates of the previous frame to form a vector; S 703 : Determine the vector direction of each frame to identify which counting line is the double-crossing line. If the vector intersects the counting line, increment the target count by one; Among them, the specific process of judging the vector direction of each frame to determine which counting line of the double crossing line is as follows: statistically analyze the movement of all target boxes in two adjacent frames. If the number of boxes in one direction is greater than that in the other direction, then determine that this frame is a frame moving in this direction; then judge whether the current number of left-moving frames is greater than that of right-moving frames. If it is greater, the counting result is statistically analyzed according to the right line, otherwise, it is statistically analyzed according to the left line.

2. The multi-category building material video counting method according to claim 1, characterized in that, In step S 200 the specific operation of extracting the features of the image to be measured is as follows: Extract three effective feature layers (76, 76, 256), (38, 38, 512), and (19, 19, 1024) from the image to be measured. The three effective feature layers are located at different positions of the backbone feature extraction network CSPDarknet53 respectively, and are used to detect small, medium, and large targets to be measured.

3. The multi-category building material video counting method according to claim 1, characterized in that, In step S 300 the last output feature layer in the backbone feature extraction network CSPDarknet53 is subjected to three DarknetConv2D_BN_Leaky convolutions and then processed using max pooling kernels of four different scales (13, 13), (9, 9), (5, 5), and (1, 1) respectively to improve the receptive field size and isolate the most prominent context features.

4. The multi-category building material video counting method according to claim 1, wherein In step S 400 The specific operation of performing multi-scale prediction on the obtained features using YOLOv3Head includes: Use YOLOv3Head to perform multi-scale prediction on the obtained features to obtain the prediction results of the three effective feature layers, so as to output three encoded tensor values of (19, 19, 33), (38, 38, 33), and (76, 76, 33), and the positions of the three prediction boxes can be determined; Obtain the coordinates of (19 * 19 + 38 * 38 + 76 * 76) * 3 boxes, and the coordinate structure is [x, y, w, h, confidence, class1, class2,..., class N]; Among them: x and y represent the upper left coordinates of each prior box, w and h represent the width and height of the prior box respectively, confidence represents the confidence that the network determines that the prior box belongs to class N, and class N represents N categories.

5. The multi-category building material video counting method according to claim 1, wherein In step S 500 wherein, inputting all the box information output by the prediction head into the NMS module to obtain the filtered box information specifically includes: After obtaining several boxes from the yolov4 network, input the array containing box information into the NMS module for non-maximum suppression to output the final detection result.

6. The multi-category building material video counting method according to claim 1, characterized in that In step S 600 the specific operation of inputting the box coordinate sequences of the front and rear frames in the target detector output frame sequence into the sort tracking module and the sort tracking module outputting the target id between frames is as follows: Input the box matrix filtered by the NMS module into the sort tracking module. The sort tracking module assigns an id to all targets in the current frame to determine whether the targets in two frames are the same target.

7. The multi-category building material video counting method according to claim 1, characterized in that, The loss function of the YOLOv3Head network includes coordinate loss coordError, confidence loss IOUError, and class prediction loss classError. The expression of the loss function of the YOLOv3Head network is as follows: Wherein: Indicates that the i-th cell contains a target, Indicates that the j-th bounding box of the i-th cell contains a target, Indicates that the j-th bounding box of the i-th cell does not contain a target, λ coord Indicates the weight value of the box regression loss, λ noobj Indicates the weight value occupied by the class without a target, Indicates the confidence that the predicted target is the i-th class, C i Represents the true confidence of the i-th class, Represents the probability predicted as the i-th class, p i (c) represents the true probability of the i-th class, and x, y, w, h respectively represent the center x, y coordinates of the predicted box and the width and height of the box.

8. A multi-category building material video counting device, comprising: A processor, a display, a memory, and computer program instructions stored on the memory and executable on the processor, characterized in that when the processor executes the computer program instructions, it is used to implement the multi-category building material video counting method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium. When the computer-executable instructions are executed by the processor, they are used to implement the multi-category building material video counting method according to any one of claims 1 to 7.