An embedded system for spot-checking semi-finished steel bars based on point counting

By using an embedded system on the construction site, equipped with a camera and image processing algorithm, combined with the target detection network P2Pnet, the rapid identification and counting of semi-finished steel bars is achieved, solving the problems of traditional manual counting inefficient and high error rates, and improving the construction efficiency and the accuracy of material management.

CN119832317BActive Publication Date: 2025-06-10ANHUI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411901147.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-06-10
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The counting and point inspection of traditional semi-finished steel bars rely on manual operations, which are inefficient and have high error rates, making it difficult to adapt to the fast-paced construction environment, affecting the control of project costs and quality assurance.

Method used

A point-test embedded system for semi-finished steel bars based on point counting is developed. The embedded development board is equipped with a camera and image processing algorithm to quickly identify and count semi-finished steel bars through the target detection network P2Pnet.

Benefits of technology

It significantly improves the point inspection speed of semi-finished steel bars, shortens working time, improves counting efficiency, ensures the accuracy and safety of material management, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832317B_ABST
    Figure CN119832317B_ABST
Patent Text Reader

Abstract

The present invention discloses an embedded system for spot-checking semi-finished steel bars based on point counting, which includes the following steps: using an embedded development board Orange Pi device equipped with an external visible light camera to capture the diagonal inclined plane image of the semi-finished steel bars at the construction site; after preprocessing the collected image, sending it to the NPU processor of the embedded development board Orange Pi for processing; using the detection and counting function in the system to detect and count the semi-finished steel bar image, and the model adopted by the detection and counting function is based on the improved object detection network P2Pnet; counting the detected quantity and inflection point position, and printing the final result in the display window. The invention uses the object detection method based on deep learning to intelligentize the counting of semi-finished steel bars, solves the disadvantages of manual counting of semi-finished steel bars, and the system significantly improves the spot-checking speed of semi-finished steel bars through the automated point counting technology, and greatly shortens the working time compared with the traditional manual counting method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of construction site building image processing, and specifically to an embedded system for spot-checking semi-finished steel bars based on point counting. Background Art

[0002] In the construction industry, semi-finished steel bars are important materials indispensable in the construction process. Their quality and quantity are directly related to the safety, stability and construction efficiency of construction projects. However, in the traditional construction process, the counting and spot-checking of semi-finished steel bars have long relied on manual operations. Manual counting is not only inefficient and difficult to adapt to the fast-paced construction environment, but also human errors are inevitable, resulting in chaos and inaccuracy in material management, further affecting the control of project costs and the guarantee of quality.

[0003] With the development of smart construction sites and the progress of technology, manual counting can no longer meet the requirements of efficient, accurate and intelligent construction. To address this issue, an embedded system for spot-checking semi-finished steel bars based on point counting has been developed. Through embedded technology and point counting algorithms, it fundamentally solves the problems of low efficiency and high error rate caused by manual counting, thereby improving construction efficiency, reducing labor costs, and at the same time ensuring more accurate and safe material management for construction projects, providing solid technical support for the quality and safety of construction projects. Summary of the Invention

[0004] (I) Technical Problems to be Solved

[0005] In view of the deficiencies of the prior art, the present invention provides an embedded system for spot-checking semi-finished steel bars based on point counting. Through the camera and image processing algorithm carried by the embedded system, it can achieve rapid identification and counting of semi-finished steel bars, improving the counting efficiency.

[0006] (II) Technical Solutions

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: An embedded system for spot-checking semi-finished steel bars based on point counting, including the following steps:

[0008] S1: Use a device with a domestic embedded development board Orange Pi equipped with an external visible light camera to capture the diagonal inclined plane image of the semi-finished steel bars at the construction site;

[0009] S2: After preprocessing the collected image, send it to the NPU processor of the domestic embedded development board Orange Pi for processing;

[0010] S3: Import the trained weights into the development board, and use the detection and counting function in the system to detect and count the semi-finished steel bar image. The network used is based on the improved object detection network P2Pnet;

[0011] S4: Count the number of detections and the inflection point positions, and print the final results in the display window.

[0012] Among them, the training of the object detection network used in the counting function in step S3 includes the following steps:

[0013] The first step is to collect training images: Use a camera on the construction site to take pictures of the diagonal slopes of the semi-finished steel bars on site to obtain a dataset, and perform manual annotation on the obtained images to obtain the training labels corresponding to the images. The annotation object is the inflection point of the semi-finished steel bars, and the annotation type is point annotation;

[0014] The second step is to train the optimal network: Divide the obtained at least 500 semi-finished steel bar datasets into a training set, a validation set, and a test set; perform data augmentation on the training set before training, and fill the shapes of all images into squares, and then scale their sizes to 640×640; Use the augmented and preprocessed training set to train the object detection network, and use the divided validation set and test set to verify and test the network to evaluate the performance of the network.

[0015] Among them, the specific steps of the network structure module training in the second step are as follows:

[0016] (1). Construct a backbone network for extracting the features of semi-finished steel bars: The backbone network uses the VGG19 network. The structure of the VGG19 network is a classic convolutional neural network, consisting of multiple convolutional layers, pooling layers, and fully connected layers. VGG19 focuses on achieving efficient feature extraction through smaller convolutional kernels (3×3), and at the same time obtaining richer feature expressions by deepening the number of layers. The encoder part consists of 19 learnable convolutional layers and several pooling layers, and gradually performs downsampling and feature extraction on the feature maps. In the first two stages, each stage contains two convolutional layers and one pooling layer. In the third, fourth, and fifth stages, each stage contains four convolutional layers and one pooling layer to extract deeper features, and retain the feature information of the three stages.

[0017] (2). Construct a Feature Pyramid Transformer module for feature information fusion: It includes RenderingTransformer (RT), Grounding Transformer (GT), and Self-Transformer (ST), and enhances the image representation ability through the interaction of multiple layers of features: RT realizes the adaptive enhancement from low-level details to high-level semantics, and its input is high-level features and low-level features Combined with x through the attention mechanism high_mask = BN(Conv3x3(x high)) and x low_gp =ReLU(BN(Conv1x1(AvgPool(x low ))) and generates a fused output; GT reversely guides low-level detail features from high-level abstract semantics, and the input is low-level x low and high-rise x high , calculate the attention f(x) through various methods such as dot product or Gaussian low , x high )=Softmax(θ(x low )·φ(x high ) T ), and get the enhanced output z = BN (Conv1x1 (y)) + x low ; ST uses multi-head self-attention to model the global interaction within the same feature layer, and the input is After projection q, k, v = Conv(x), 1 calculates attention The final output is combined with the residual connection to obtain out = Norm (Conv (output)) + x. Overall, FPT constructs rich spatial and semantic feature representations through cross-layer and intra-layer feature interactions.

[0018] (3) Construct a regression-classification branch attention (R-CBA) module: The two branches of R-CBA optimize the positioning and classification tasks respectively, so that the model can simultaneously achieve high-precision point annotation positioning and accurate target classification, especially in dense target scenes. This module contains a regression branch and a classification branch, and the two share a feature extraction network to ensure feature consistency and full utilization. In the regression branch, the input is the shared feature x, and the feature extraction and enhancement are performed through multi-layer convolution and embedded CBAM module (combining channel attention mechanism and spatial attention mechanism), and the final output is Where B is the batch size, N is the number of predicted points, and the output of each point contains its two-dimensional coordinates in the image, which is used to accurately locate the point annotation position of the target. In the classification branch, the shared feature x is also processed, and the key features are enhanced by CBAM and then output Where C is the number of categories, and the classification score marked on each point represents the probability distribution of the category to which it belongs.

[0019] As an optimization, in step S1, ensure that the interface between the Orange Pi development board and the external visible light camera matches, such as USB, MIPI, etc., and prepare the corresponding connecting cables. Fix the camera when shooting the diagonal inclined surface image of the semi-finished steel bar at the construction site to ensure that its shooting angle and position can accurately capture the image of the diagonal inclined surface of the semi-finished steel bar.

[0020] As an optimization, during the preprocessing in step S2, all image shapes are filled into squares and then scaled proportionally to 640×640 to achieve the best detection effect.

[0021] As an optimization, in step S3, in the second step, the data is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1. The methods for data augmentation of the training set include horizontal mirroring, adjustment of image brightness, and salt-and-pepper noise with random intensity. The brightness adjustment factor ranges from 0.5 to 1.5, and the noise adjustment is set such that there is a 5% probability for each pixel to be replaced with salt-and-pepper noise.

[0022] As an optimization, in step S4, the detection results output by the NPU processor of the embedded development board are statistically analyzed, including summing up the number of predicted points detected to obtain the quantity of semi-finished steel bars, and plotting the inflection point position information on the result graph and displaying it on the touch screen of the Orange Pi development board.

[0023] (III) Beneficial Effects

[0024] The present invention provides an embedded system for spot-checking semi-finished steel bars based on point counting, having the following

[0025] Beneficial Effects:

[0026] The present invention uses an object detection method based on deep learning to intelligentize the counting of semi-finished steel bars, solving the drawbacks of manual counting of semi-finished steel bars. This system significantly improves the spot-checking speed of semi-finished steel bars through an automated point counting technology. Compared with the traditional manual counting method, the working time is greatly shortened. The embedded system for spot-checking semi-finished steel bars based on point counting brings improvements and enhancements to the management of semi-finished steel bars at the construction site through automated and intelligent technical means. Description of the Drawings

[0027] Figure 1 It is a flow chart of the semi-finished steel bar detection method of the present invention;

[0028] Figure 2 It is a framework diagram of network optimization and semi-finished steel bar detection of the present invention;

[0029] Figure 3 It is a framework diagram of the feature pyramid transformer module of the present invention;

[0030] Figure 4 It is a framework diagram of the regression-classification branch attention detection module of the present invention;

[0031] Figure 5 It is an effect diagram of detecting semi-finished steel bars of the present invention. Detailed Embodiment

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] Please refer to Figures 1-5 , the present invention provides a technical solution:

[0034] As Figure 2 shown, the model training and testing are completed in the following steps:

[0035] The first step is to collect training images: Use a camera on the construction site to take pictures of the diagonal slopes of the semi-finished steel bars on site to obtain a data set, and perform manual annotation on the obtained images to obtain the training labels corresponding to the images. The annotation object is the inflection points of the semi-finished steel bars, and the annotation type is point annotation;

[0036] The second step is to train the optimal network: Divide the obtained at least 500 semi-finished steel bar data sets into a training set, a validation set, and a test set; perform data augmentation on the training set before training, and fill the shapes of all images into squares, and then scale their sizes to 640×640; use the augmented and pre-processed training set to train the target detection network, and use the divided validation set and test set to verify and test the network to evaluate the performance of the network. The specific steps for the network to be trained are as follows:

[0037] (1). Construct a backbone network for extracting the features of semi-finished steel bars: The backbone network uses the VGG19 network. The structure of the VGG19 network is a classic convolutional neural network, consisting of multiple convolutional layers, pooling layers, and fully connected layers. VGG19 focuses on achieving efficient feature extraction through smaller convolutional kernels (3×3), and at the same time obtaining richer feature expressions by deepening the number of layers. The encoder part consists of 19 learnable convolutional layers and several pooling layers, and gradually performs downsampling and feature extraction on the feature maps. In the first two stages, each stage contains two convolutional layers and one pooling layer. In the third, fourth, and fifth stages, each stage contains four convolutional layers and one pooling layer to extract deeper features, and retain the feature information of the three stages.

[0038] (2) Construct a Feature Pyramid Transformer module for feature information fusion: It includes RenderingTransformer (RT), Grounding Transformer (GT), and Self-Transformer (ST), which enhance the image representation ability through the interaction of multi-layer features. RT realizes the adaptive enhancement from low-level details to high-level semantics, and its input is high-level features and low-level features Combined with x through the attention mechanism high_mask = BN(Conv3x3(x high )) and x low_gp = ReLU(BN(Conv1xl(AvgPool(x low ))), and generate the fused output; GT guides the low-level detail features from high-level abstract semantics in the reverse direction, and the input is low-level x low and high-level x high , calculates the attention f(x low , x high ) = Softmax(θ(x low )·φ(x high ) T ) through various methods such as dot product or Gaussian, and obtains the enhanced output z = BN(Conv1x1(y)) + x low ; ST models the global interaction within the same feature layer through multi-head self-attention, and the input is After projection q, k, v = Conv(x), calculate the attention Finally, the output is combined with the residual connection to get out = Norm(Conv(output)) + x. Overall, FPT constructs rich spatial and semantic feature representations through cross-layer and intra-layer feature interactions.

[0039] (3) Construct a Regression-Classification Branch Attention (R-CBA) module: The two branches of R-CBA optimize the localization and classification tasks respectively, enabling the model to simultaneously achieve high-precision point annotation localization and accurate target classification, especially showing superiority in dense target scenarios. This module includes a regression branch and a classification branch, and the two share a feature extraction network to ensure feature consistency and full utilization. In the regression branch, the input is the shared feature x, and through multi-layer convolution and the embedded CBAM module (combining channel attention mechanism and spatial attention mechanism), feature extraction and enhancement are performed, and finally the output where B is the batch size and N is the number of predicted points. The output of each point contains its two-dimensional coordinates in the image, which is used to accurately locate the point annotation position of the target. In the classification branch, the shared feature x is also processed, and after enhancing the key features through CBAM, the output is Where C is the number of categories, and the classification scores marked for each point represent the probability distribution of the category to which it belongs.

[0040] The third step, the semi-finished steel bar detection function: Input the inclined plane image of the semi-finished steel bar taken on-site into the trained point annotation detection network, use the network to process the inclined plane image, extract the feature information therein, and obtain the predicted inflection point positions and quantities of the semi-finished steel bars using the positioning and classification branches, and output the results on the touch screen of the Orange Pi development board.

[0041] Based on the above steps, as Figure 1 shown, an embedded system for spot-checking semi-finished steel bars based on point counting has the following specific operation steps:

[0042] S1: Use a device with a domestic embedded development board Orange Pi equipped with an external visible light camera to take the diagonal inclined plane image of the semi-finished steel bars at the construction site;

[0043] S2: After preprocessing the collected images, send them to the NPU processor of the domestic embedded development board Orange Pi for processing;

[0044] S3: Import the trained weights into the development board, and use the detection and counting function in the system to detect and count the semi-finished steel bar images. The network used is the improved object detection network P2Pnet;

[0045] S4: Statistically count the detection quantity and inflection point positions, and print the final results in the display window.

[0046] In step S1, ensure that the interface between the Orange Pi development board and the external visible light camera matches, such as USB, MIPI, etc., and prepare the corresponding connecting wires. Fix the camera when taking the diagonal inclined plane image of the semi-finished steel bars at the construction site to ensure that its shooting angle and position can accurately capture the diagonal inclined plane image of the semi-finished steel bars.

[0047] During the preprocessing in step S2, fill the shapes of all images into squares, and scale them proportionally to 640×640 to achieve the best detection effect.

[0048] In step S3, in the second step, divide the data into a training set, a validation set, and a test set according to the ratio of 7:2:1. The methods for data augmentation of the training set include horizontal mirroring, adjustment of image brightness, and salt-and-pepper noise with random intensity. Among them, the brightness adjustment magnification is from 0.5 to 1.5, and the noise adjustment is set so that each pixel has a 5% probability of being replaced by salt-and-pepper noise.

[0049] In step S4, the detection results output by the NPU processor of the embedded development board are statistically analyzed, including summing up the number of detected prediction points to obtain the quantity of semi-finished steel bars, and plotting the inflection point position information on the result graph and displaying it on the touch screen of the OrangePi development board.

[0050] In the optimization process of the entire network of the present invention, the network gradually learns the multi-scale features and global context information of different semi-finished steel bars. When the model finds the optimal network model through continuous optimization, the entire network can be used as a tool for spot-checking semi-finished steel bars. At this time, the model has the ability to perform spot-check counting of semi-finished steel bars in the construction site field.

[0051] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An embedded system for checking semi-finished steel bars based on point counting, characterized in that: The following steps are involved: S1: Use Orange Pi, an embedded development board equipped with an external visible light camera, to capture images of the diagonal slope of semi-finished steel bars at the construction site; S2: After pre-processing, the collected image is sent to the NPU processor of the embedded development board Orange Pi for processing; S3: Use the detection and counting function in the system to detect and count the semi-finished steel bar images. The model used by the detection and counting function is based on the improved target detection network P2Pnet; S4: Count the number of tests and the position of the inflection points, and print out the final results in the display window; The training of the target detection network P2Pnet used for the detection counting function model in step S3 includes the following steps: The first step is to collect training images: use a camera on the construction site to shoot the diagonal slopes of the semi-finished steel bars to obtain a data set, manually annotate the acquired images, and obtain the training labels corresponding to the images. The annotated objects are the inflection points of the semi-finished steel bars, and the annotation type is point annotation. The second step is optimal network training: the obtained dataset of at least 500 semi-finished steel bars is divided into training set, validation set and test set; before training, the training set is augmented and all image shapes are filled with squares, and then the size is scaled to 640×640; the augmented and preprocessed training set is used to train the object detection network, and the divided validation set and test set are used to verify and test the network to evaluate the network performance; Among them, the specific steps of target detection network training in the second step are as follows: (1) Construct a backbone network for extracting features of semi-finished steel bars: The backbone network uses the VGG19 network. The encoder part consists of 19 learnable convolutional layers and several pooling layers, which gradually downsample and extract features from feature maps. In the first two stages, each stage contains two convolutional layers and one pooling layer. In the third, fourth, and fifth stages, each stage contains four convolutional layers and one pooling layer, which extract deeper features and retain the feature information of the three stages. (2) Construct a feature pyramid Transformer module that fuses feature information: It includes RT, GT, and ST. It enhances the image representation capability through the interaction of multiple layers of features: RT realizes adaptive enhancement from low-level details to high-level semantics, and its input is high-level features. and low-level features Combine x through the attention mechanism high_mask =BN(Conv3x3(x high )) and x low_gp =ReLU(BN(Conv1x1(AvgPool(x low ))) and generates a fused output; GT reversely guides low-level detail features from high-level abstract semantics, and the input is low-level x low and high-rise x high , calculate the attention f(x) by dot product or Gaussian method low , x high )=Softmax(θ(x low )·φ(x high ) T ), and get the enhanced output z = BN (Conv1x1 (y)) + x low ; ST uses multi-head self-attention to model the global interaction within the same feature layer, and the input is After projection q, k, v = Conv(x), calculate the attention The final output is combined with the residual connection to obtain out = Norm (Conv (output)) + x. Overall, FPT constructs rich spatial and semantic feature representations through cross-layer and intra-layer feature interactions; (3) Construct a regression-classification branch attention module: The module contains a regression branch and a classification branch. The two branches share a feature extraction network to ensure the consistency and full utilization of features. In the regression branch, the input is the shared feature x. Through multi-layer convolution and embedded CBAM modules, the channel attention mechanism and the spatial attention mechanism are combined to extract and enhance features, and the final output is Where B is the batch size, N is the number of predicted points, and the output of each point contains its two-dimensional coordinates in the image, which is used to accurately locate the point annotation position of the target; in the classification branch, the shared feature x is also processed, and the key features are enhanced by CBAM and then output Where C is the number of categories, and the classification score marked on each point represents the probability distribution of the category to which it belongs.

2. According to claim 1, a semi-finished steel bar inspection embedded system based on point counting is characterized in that: In step S1, ensure that the interface between the Orange Pi development board and the external visible light camera matches, and prepare the corresponding connecting wires. Fix the camera when shooting the diagonal inclined surface of the semi-finished steel bars on the construction site to ensure that its shooting angle and position can accurately capture the image of the diagonal inclined surface of the semi-finished steel bars.

3. According to claim 1, a semi-finished steel bar inspection embedded system based on point counting is characterized in that: During the preprocessing in step S2, all image shapes are filled into squares and then proportionally scaled to 640×640 to achieve the best detection effect.

4. According to claim 1, a semi-finished steel bar inspection embedded system based on point counting is characterized in that: In the second step, step S3 divides the data into a training set, a validation set, and a test set in a ratio of 7:2:

1. The method of data augmentation for the training set includes horizontal mirroring, brightness adjustment of the image, and salt and pepper noise of random intensity, where the brightness adjustment ratio is 0.5 to 1.5, and the noise adjustment is set to a 5% probability that each pixel is replaced with salt and pepper noise.

5. The embedded system for checking semi-finished steel bars based on point counting according to claim 1 is characterized in that: Step S4 counts the detection results output by the NPU processor of the embedded development board, including summing the number of detected prediction points to obtain the number of semi-finished steel bars, plotting the inflection point position information on the result graph and displaying it on the touch screen of the Orange Pi development board.

Citation Information

Patent Citations

  • Small sample steel defect detection method based on attention feature pyramid mechanism

    WO2025010883A1