Method for identifying third-party construction based on video

Through the deep learning detection model based on video recognition, the traditional recognition method lacks the recognition rate and real-time monitoring capabilities of large-scale construction machinery, and realizes high-precision and automated construction monitoring, reducing safety hazards and misjudgment risks.

CN120164154AInactive Publication Date: 2025-06-17ANHUI HUICAI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510077900.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional identification method lacks the recognition rate and real-time monitoring capabilities of large construction machinery, making it difficult to effectively solve safety hazards and misjudgment problems.

Method used

Using a deep learning detection model based on video recognition, a detection model including Backbone, FPN and YoloHead is constructed by acquiring and labeling image data of large machinery, for training and evaluation, and deploying it to the TensorRT inference engine.

Benefits of technology

It improves the recognition rate and real-time monitoring capabilities of large-scale construction machinery, reduces misjudgment and disputes, and can operate stably in various environments to ensure continuous and automatic monitoring of the construction area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164154A_ABST
    Figure CN120164154A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of construction supervision, in particular to a method for identifying third-party construction based on videos. The method for identifying the third-party construction based on the video comprises the following steps that image data of large machines are collected, the large machines comprise a road roller, a paver, a loader, an excavator and a crane, and the image data cover different angles and illumination conditions; marking the image data to obtain the category of each machine and the position information of the machine in the image, and representing the position information in the form of a bounding box; increasing the diversity of the image data using a data enhancement technique; dividing the marked image data into a training set, a verification set and a test set according to a ratio of 8: 1: 1; according to the invention, large machinery constructed by a third party is effectively supervised by using a deep learning video recognition technology, the construction safety and efficiency are ensured, and a safer and more efficient large machinery supervision solution is provided for the field of road construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of construction supervision, and more specifically, to a method for identifying third-party construction based on video recognition. Background Art

[0002] With the acceleration of urbanization and the increasing density of road networks, third-party construction activities in the vicinity are carried out frequently. During this process, the widespread application of large construction machinery such as rollers, pavers, loaders, excavators, and cranes has become a key factor in road construction. These machines undertake important tasks in different construction stages, such as excavating earthwork, laying road surface materials, and lifting heavy objects. However, traditional identification and monitoring methods face many challenges in accurately identifying these large machines, especially in terms of potential safety hazards such as the protection of underground pipelines and the impact on road bearing capacity. Improper construction may lead to serious consequences.

[0003] Existing research has used video recognition methods based on Faster R-CNN and combined with the introduction of image spatial features to improve the recognition rate of large construction machinery in third-party construction. In addition, an analysis and processing flow based on the Apriori association analysis algorithm to construct potential strong association relationships between the construction environment, equipment dynamics monitoring, and hidden danger data has further improved the discovery and disposal efficiency of hidden dangers at the construction site.

[0004] Currently, the development of technologies such as computer vision and artificial intelligence provides comprehensive technical support for third-party construction identification based on video. Deep learning technology has made remarkable progress in the field of image recognition, and the feasibility of automatically identifying third-party construction using video images is gradually increasing. Accurate monitoring and identification can be achieved through processes such as data preprocessing, feature extraction, and classification recognition. After entering the 5G era, the data transmission between cameras and servers is more efficient, laying a good foundation for third-party construction identification based on video and meeting the requirements of construction identification tasks with high real-time requirements.

[0005] There are mainly two existing identification methods: one is manual identification, that is, relying on the observation and judgment of on-site staff. However, this method requires a large amount of human input and is easily affected by factors such as fatigue and distraction, resulting in a decrease in the accuracy and timeliness of identification. In addition, in the face of a complex construction environment, it may be difficult for manual identification to distinguish similar machines or handle occluded situations. The other traditional method is simple sensor technology, such as proximity sensors at fixed positions or weight-based sensors, etc. These sensors are relatively single, with insufficient versatility and adaptability, and cannot meet the identification requirements for different large construction machines in complex road construction environments.

[0006] With the rapid development of information technology, deep learning has demonstrated unprecedented advantages in the fields of image and video processing. By constructing multi-layer neural network models, deep learning can automatically learn and extract features. During the video recognition of large-scale machinery in third-party construction, deep learning can collect rich video data through various means such as on-site construction surveillance cameras and drone photography. These videos cover different lighting conditions, weather changes, and complex construction environments. Deep learning algorithms can analyze this massive amount of data, thereby learning the unique visual features and appearance changes of each large-scale machinery, establishing a high-precision and highly reliable recognition model, providing solid technical support for the intelligent management of road construction, construction progress monitoring, and safety hazard investigation, and significantly improving the efficiency and quality of road construction management.

[0007] Based on the above background, the present invention proposes a method for third-party construction based on video recognition, aiming to solve the deficiencies of traditional recognition methods and improve the recognition rate and real-time monitoring ability of large-scale construction machinery. Summary of the Invention

[0008] The object of the present invention is to provide a method for third-party construction based on video recognition to solve the problems of insufficient recognition rate and real-time monitoring ability of large-scale construction machinery in the traditional recognition method proposed in the above background technology.

[0009] To achieve the above object, the present invention aims to provide a method for third-party construction based on video recognition, including the following steps:

[0010] S1. Collect image data of large-scale machinery, where the large-scale machinery includes rollers, pavers, loaders, excavators, and cranes, and the image data covers different angles and lighting conditions;

[0011] S2. Label the image data to obtain the category of each machinery and its position information in the image, and the position information is represented in the form of a bounding box;

[0012] S3. Use data augmentation techniques to increase the diversity of image data;

[0013] S4. Divide the labeled image data into a training set, a validation set, and a test set at a ratio of 8:1:1;

[0014] S5. Construct a deep learning detection model including Backbone, FPN, and YoloHead, extract image features through training, fuse feature information at different scales, and perform predictions on category and position information;

[0015] S6. Evaluate the performance of the trained model using metrics such as mean average precision (mAP), recall rate, precision rate, and intersection over union (IoU);

[0016] S7. Deploy the trained model to the TensorRT inference engine.

[0017] As a further improvement of this technical solution, the image data of large machinery is acquired through the monitoring cameras and drones at the construction site.

[0018] As a further improvement of this technical solution, the use of data augmentation techniques to increase the diversity of image data in step S3 specifically includes steps of rotation, scaling, flipping, adding noise, and randomly cropping and color adjusting the images.

[0019] As a further improvement of this technical solution, the convolutional layers used in the Backbone adopt depthwise separable convolutions.

[0020] As a further improvement of this technical solution, the FPN fuses feature maps from different scales through a top - down and lateral connection manner.

[0021] As a further improvement of this technical solution, the YoloHead performs class prediction through convolutional operations and uses the binary cross - entropy loss function to obtain class probabilities.

[0022] As a further improvement of this technical solution, when evaluating the performance of the trained model in step S6, the accuracy of object detection is analyzed according to different IoU thresholds, and a PR curve is plotted to show the model performance.

[0023] As a further improvement of this technical solution, during the process of model deployment in step S7, model optimization is carried out through TensorRT, including model quantization and the creation of an inference engine.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0025] 1. The method for identifying third - party construction based on video recognition proposed in the present invention solves the problems of low efficiency and easy omission of manual inspection of third - party construction. Through video recognition, continuous and automatic monitoring of the construction area can be realized, greatly improving the monitoring efficiency and accuracy.

[0026] 2. The method for identifying third - party construction based on video recognition proposed in the present invention solves the problem that it is difficult to detect potential safety hazards caused by third - party construction in a timely manner. It can quickly identify construction behaviors, give early warnings before dangerous situations occur, and gain time for taking measures to ensure the safety of the construction environment.

[0027] 3. The method for identifying third - party construction based on video recognition proposed in the present invention solves the problem of vague definition of third - party construction behaviors. Through precise feature extraction and model training, it can clearly judge whether it is third - party construction, avoiding misjudgment and disputes.

[0028] 4. The method for identifying third-party construction based on video recognition proposed in the present invention solves the problem of high difficulty in recognition under different environments. Whether it is light change, bad weather or complex background scenes, this method can effectively extract key features for accurate recognition, ensuring stable operation under various working conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is the overall flowchart of the method for identifying third-party construction based on video recognition of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] In a specific embodiment, as Figure 1 shown, the present invention provides a method for identifying third-party construction based on video recognition, including the following steps:

[0032] I. Data preparation.

[0033] Collect images of large machinery (rollers, pavers, loaders, excavators, cranes). These images have diversity, including different angles, lighting conditions, etc. Then perform annotation. For each target object in the image, accurately annotate its category and location. The location information is usually given in the form of the coordinates of the bounding box. In order to improve the generalization ability of the model, data augmentation techniques are adopted, and the richness of the data is increased by means of rotation, scaling, flipping, adding noise, etc. At the same time, the training set, validation set and test set are reasonably divided, and divided according to a certain ratio of 8:1:1 to ensure sufficient training data and the validation and test sets can effectively evaluate the model performance.

[0034] II. Model construction.

[0035] Backbone: It is mainly responsible for extracting rich and representative features from the input image. It consists of a series of convolutional layers and pooling layers. The convolutional kernels used in these convolutional layers are 3×3 in size and have a stride of 2. Through continuous convolutional operations, it gradually extracts low-level features such as the texture and edges of the image and more abstract high-level semantic features. In the design, it focuses on reducing the computational amount and the number of parameters, and uses the depthwise separable convolution technique. It also has downsampling operations to reduce the size of the feature map and increase the receptive field, enabling the network to capture image information in a larger range. And through the network structure design, it ensures the effective transmission and fusion of features at different levels, provides high-quality features for the subsequent detection heads, and improves the detection ability of the entire model for targets at different scales.

[0036] FPN: FPN (Feature Pyramid Network) plays a crucial role in detection. It is mainly used to fuse feature information at different levels to address the multi-scale object detection problem. FPN obtains feature maps at different scales from the Backbone. These feature maps have different semantic information and resolutions. The high-level feature maps have strong semantic information but low resolution, while the low-level feature maps have high resolution but weak semantic information. Through the top-down and lateral connection methods, it fuses the strong semantic features of the high level and the detailed features of the low level, so that the features at each scale contain rich semantic and location information. In this way, when performing object prediction, the feature maps at different scales can better detect objects of different sizes. Small objects can be more accurately detected in the low-level feature maps with high resolution, and large objects can be accurately located and classified with the assistance of the relevant information in the fused high-level feature maps.

[0037] YoloHead: YoloHead is the key module for object detection prediction. It receives the feature maps fused by FPN and makes predictions on object categories, positions, and confidences based on these feature maps. YoloHead usually contains multiple parallel branches, and each branch is responsible for outputting detection-related information for objects at a specific scale. For category prediction, it maps the features to the corresponding category dimension space through convolutional operations and uses the binary cross-entropy loss function to obtain the category probabilities. In terms of position prediction, it calculates the coordinates of the object's bounding box, comprehensively considering the position information and relative scale information in the feature map to accurately determine the position of the object in the image. Confidence prediction determines whether there is an object in the predicted bounding box. Combining the category and position information, it finally outputs high-quality detection results to achieve accurate detection of objects at different scales and different categories.

[0038] III. Model Training.

[0039] Determine hyperparameters, such as learning rate, batch size, number of iterations, etc. After the training starts, the model extracts features through the Backbone based on the input image data, fuses features with FPN, and makes predictions with YoloHead. Calculate the loss between the prediction results and the true labels. The loss function includes object confidence loss, class loss, and coordinate regression loss. During the training process, the optimization algorithm SGD is used to continuously adjust the model parameters to reduce the loss. At the same time, the validation set is used to monitor the model performance to prevent overfitting. When overfitting is found, techniques such as regularization are adopted.

[0040] IV. Model Evaluation.

[0041] Its performance is mainly measured by a variety of metrics. The mean average precision (mAP) is used. mAP comprehensively considers the precision of different category targets at different recall rates and can accurately reflect the accuracy of the model in detecting various targets. When calculating mAP, analyze the matching degree between the target categories and positions predicted by the model and the true annotations. At the same time, the recall rate will also be concerned, that is, the ratio of the number of targets correctly detected by the model to the actual number of targets, which reflects the ability of the model to find all targets, and the precision rate, that is, the proportion of the number of correctly predicted targets in all predicted targets, which reflects the accuracy of the prediction. In addition, the selection of the intersection over union (IoU) threshold is also crucial, and the performance of the model is different under different thresholds. The PR curve is plotted to visually display the performance of the model at different recall rates and precision rates, so as to comprehensively evaluate the detection effect of the model.

[0042] V. Model Deployment.

[0043] Use TensorRT to deploy the YOLOv7 model. First, convert the trained YOLOv7 model into a format supported by TensorRT. This process involves parsing the model structure and parameters and converting them into the computation graph of TensorRT. When converting, utilize the optimization function of TensorRT to quantize the model, convert floating-point operations into low-precision integer operations, reduce the computational amount and storage requirements, and at the same time do not significantly reduce the precision. Then create an inference engine, set the relevant configurations such as input and output parameters and memory allocation. Next, input the image data into the TensorRT inference engine, and perform fast forward inference calculation through its optimized computation core, and output the target detection results, including information such as the category, position, and confidence of the target, so as to achieve the efficient deployment and fast inference of the model on TensorRT.

[0044] The method for third-party construction based on video recognition proposed by the present invention has the following advantages:

[0045] From the perspective of efficiency, the fast detection ability of YOLOv7 can process video streams in real time, can quickly judge construction behaviors, is much more efficient than manual inspections, and can avoid long-term monitoring blanks;

[0046] In terms of accuracy, YOLOv7 is trained with a large amount of data, accurately extracts the features of construction workers, tools and scenes, can accurately identify third-party construction, and reduce misjudgment. It has strong generalization ability and can be effectively identified regardless of different lighting, weather or complex backgrounds, adapting to various construction environments. Moreover, this automatic identification method can work continuously for 24 hours, continuously monitor the construction area, and once abnormal construction is found, it can give an early warning in time, enabling relevant parties to respond quickly, reducing the safety risks and damage to the project caused by third-party construction, and ensuring construction safety and progress.

[0047] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A method for identifying third-party construction based on video, characterized in that: The following steps are involved: S1. Collect image data of large machinery, including rollers, pavers, loaders, excavators and cranes, and the image data covers different angles and lighting conditions; S2. Annotating the image data to obtain the category of each machine and its location information in the image, where the location information is represented in the form of a bounding box; S3. Use data augmentation techniques to increase the diversity of image data; S4, dividing the annotated image data into training set, validation set and test set in a ratio of 8:1:1; S5. Build a deep learning detection model including Backbone, FPN and YoloHead, extract image features through training, fuse feature information of different scales, and predict category and location information; S6. Evaluate the performance of the trained model using mean average precision (mAP), recall, precision, and intersection over union (IoU) metrics; S7. Deploy the trained model to the TensorRT inference engine.

2. The method for video-based identification of third-party construction according to claim 1, characterized in that: The image data of large machinery is acquired through monitoring cameras and drones at the construction site.

3. The method for video-based identification of third-party construction according to claim 1, characterized in that: The step S3 uses data enhancement technology to increase the diversity of image data, specifically including the steps of rotating, scaling, flipping, adding noise, and performing random cropping and color adjustment on the image.

4. The method for video-based identification of third-party construction according to claim 1, characterized in that: The convolutional layer used in Backbone adopts depthwise separable convolution.

5. The method for video-based identification of third-party construction according to claim 1, characterized in that: The FPN fuses feature maps from different scales through top-down and lateral connections.

6. The method for video-based identification of third-party construction according to claim 1, characterized in that: The YoloHead performs category prediction through convolution operations and uses a binary cross entropy loss function to obtain category probabilities.

7. The method for video-based identification of third-party construction according to claim 1, characterized in that: When evaluating the performance of the trained model in step S6, the accuracy of target detection is analyzed according to different IoU thresholds, and a PR curve is drawn to display the model performance.

8. The method for video-based identification of third-party construction according to claim 1, characterized in that: During the model deployment process in step S7, model optimization is performed through TensorRT, including model quantization and creation of an inference engine.

Citation Information

Cited By

  • Integrated intelligent power distribution control system of electric loader

    CN120697778A