Deep learning-based concrete compression test block automatic positioning method

Through the DeepLabV3+ network combined with local weighted convolution kernel and morphological processing, the problem of insufficient positioning accuracy of concrete test blocks is solved, and efficient and automated test block positioning is achieved, which is suitable for large-scale inspection.

CN120339568APending Publication Date: 2025-07-18ZHONG JIAO YI HANG JU DI SAN GONG CHENG YOU XIAN GONG SI +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510427586.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing concrete compressive test block positioning method relies on manual operation, is inefficient and susceptible to human factors. The existing deep learning methods lack positioning accuracy under complex backgrounds and low-quality images, making it difficult to adapt to large-scale batch detection.

Method used

The DeepLabV3+ network is used to combine local weighted convolution kernel, spatial attention mechanism and morphological processing to achieve automated positioning of test blocks through image preprocessing, edge feature extraction, rectangular enclosure box optimization and Hough transformation.

Benefits of technology

It significantly improves the accuracy and efficiency of the positioning of concrete test blocks, reduces human error, is suitable for large-scale inspection, and improves the degree of automation and production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339568A_ABST
    Figure CN120339568A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning-based concrete compression test block automatic positioning method. The method comprises the following steps of S1, acquiring test block image data through image acquisition equipment; s2, preprocessing the image data; s3, inputting the preprocessed image into a DeepLabV3 + network encoder, extracting edge features by using a local weighted convolution kernel and a space attention mechanism, and generating a preliminary region position; s4, reconstructing edge information through a DeepLabV3 + decoder and a convolution operation with a fixed step length, and outputting a rectangular bounding box and a region identifier; s5, performing morphological processing on the rectangular bounding box, optimizing the position and the size of the bounding box, and calculating a positioning error; and S6, in combination with the positioning error and the region identifier, carrying out secondary optimization on the positioning result by adopting Hough transformation. According to the method, the DeepLabV3 + network and the morphological processing technology are combined, the edge features of the concrete compression resistance test block are accurately extracted, the positioning process is optimized, manual intervention and errors are reduced, and the positioning precision and efficiency of the test block are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of concrete and building material testing, and particularly to an automatic positioning method for concrete compressive test blocks based on deep learning. Background Art

[0002] In the concrete compressive test, the accurate positioning of test blocks is the key to ensuring the reliability and accuracy of test results. Traditional methods for positioning concrete compressive test blocks mainly rely on manual operations. Usually, professional personnel determine the position of test blocks through visual observation and manual operations. Although this method can meet the requirements in small-scale tests, there are many problems when facing large-scale and batch concrete quality inspections. First of all, manual positioning is not only inefficient but also easily affected by human factors, resulting in fluctuations in positioning accuracy. In addition, manual operations are easily interfered by factors such as the experience and fatigue of operators, further increasing the probability of errors, and ultimately affecting the test results and quality.

[0003] With the development of intelligent technologies, more and more research has begun to attempt to apply image processing technologies to the positioning of concrete test blocks. Image processing technologies identify and position test blocks through computer vision technologies, thereby improving the efficiency and accuracy of positioning. Existing positioning methods based on image processing usually use technologies such as edge detection and image segmentation to identify the edges of test blocks. Although these methods have improved the positioning efficiency to a certain extent, there are still some problems. First of all, traditional edge detection methods (such as Canny edge detection) are greatly affected by noise, resulting in inaccurate edge extraction. Especially in complex lighting or poor image quality conditions, it is easy to cause edge information loss or misidentification, thereby affecting the accurate positioning of test blocks. Secondly, traditional image processing methods for estimating the position of test blocks usually rely on threshold segmentation or other rule-based algorithms. These methods lack adaptability and are difficult to cope with changes in different environments. Therefore, how to improve the robustness and accuracy of image processing technologies has become an important issue for improving the positioning accuracy of concrete compressive test blocks.

[0004] In recent years, the rapid development of deep learning technology has brought revolutionary breakthroughs to the field of image recognition and processing. Deep learning algorithms, especially convolutional neural networks (CNNs), have become important tools in the field of image processing. Through convolutional neural networks, the system can automatically learn the important features in images and gradually extract more complex image information through deep network structures. Image processing methods based on deep learning have achieved remarkable results in many fields, such as object detection, image segmentation, etc. However, in the automatic positioning of concrete compressive test blocks, existing deep learning methods still face some challenges. First, although deep learning methods can effectively extract features in images, in practical applications, due to the diversity of factors such as the shape, size, and position of the test blocks, existing deep learning methods are still insufficient in terms of positioning accuracy and robustness, especially in complex backgrounds and low-quality images, and the performance of the model is not as expected. Second, most traditional deep learning models rely on large-scale datasets for training. However, in the field of concrete compressive test block detection, relevant labeled data is relatively limited. How to maintain a high positioning accuracy in the case of scarce data is an urgent problem to be solved.

[0005] Currently, some studies have attempted to use deep learning-based image segmentation methods for the positioning of test blocks, such as using network models like U-Net for image segmentation and feature extraction. These methods estimate the position of the test blocks by segmenting the image and extracting the regional information of the test blocks. Compared with traditional image processing methods, this method has indeed improved the positioning accuracy and automation to a certain extent, but there are still some obvious deficiencies. First, the existing image segmentation methods have limited ability to refine the contour of the test blocks. In some complex scenarios, such as uneven lighting and strong background interference, the segmentation effect is not ideal, resulting in inaccurate positioning of the test blocks. Second, although deep learning methods can extract edge features in images, existing network models often cannot effectively handle the interference between multiple test blocks or similar backgrounds in the image, resulting in misidentification and mispositioning of the test block edges.

[0006] In order to further improve the accuracy and efficiency of concrete test block positioning, in recent years, researchers have begun to combine deep learning with traditional image processing techniques. For example, combining edge detection algorithms and deep learning models, using edge information for test block positioning, and attempting to improve the accuracy and robustness of edge recognition through deep learning networks. However, existing technologies still face problems such as inaccurate edge feature extraction, increased complexity of test block shapes, and large fluctuations in image quality. In existing technologies, although deep learning networks have made some progress in edge extraction and region segmentation, due to excessive reliance on the feature extraction ability of deep convolutional networks, the network fails to capture sufficient detailed information in the image. Especially when dealing with complex or low-quality images, the positioning accuracy and efficiency are still not ideal.

[0007] Therefore, how to provide an automatic positioning method for concrete compressive test blocks that can make full use of the advantages of deep learning, combine edge feature extraction and image processing technologies, and have high robustness and accuracy is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0008] An object of the present invention is to propose an automatic positioning method for concrete compressive test blocks based on deep learning. The present invention fully combines a deep learning network and image processing technology, adopts the encoder and decoder structures of the DeepLabV3+ network, and through technologies such as local weighted convolution kernels, spatial attention mechanisms, and morphological processing, can accurately identify the edge features of the test blocks and effectively optimize the positioning results in complex environments. This method has a high degree of automation, can significantly improve the accuracy and efficiency of positioning, reduce human errors, and is especially suitable for large-scale concrete quality inspection.

[0009] An automatic positioning method for concrete compressive test blocks based on deep learning according to an embodiment of the present invention includes the following steps:

[0010] S1. Obtain image data of the concrete compressive test block through an image acquisition device;

[0011] S2. Preprocess the obtained image data;

[0012] S3. Input the preprocessed image data into the encoder of the DeepLabV3+ network, extract the edge features of the concrete compressive test block through local weighted convolution kernels and spatial attention mechanisms, and generate the preliminary regional position of the test block;

[0013] S4. According to the preliminary regional position of the test block, reconstruct the edge information of the test block through the convolution operation with a fixed step size in the decoder of the DeepLabV3+ network, output the rectangular bounding box of the test block and generate the corresponding regional identifier;

[0014] S5. Perform morphological processing on the output rectangular bounding box, remove the noise area through opening operation, optimize the position and size of the rectangular bounding box using the maximum area method, locate the edge of the test block and calculate the positioning error;

[0015] S6. Combine the positioning error and the regional identifier, and perform secondary optimization on the positioning result using the line detection method based on the Hough transform, and output the position of the concrete compressive test block.

[0016] Optionally, the preprocessing includes image denoising, grayscale adjustment, and edge detection.

[0017] Optionally, S3 includes the following specific steps:

[0018] S31. Input the preprocessed image data into the encoder of the DeepLabV3+ network;

[0019] S32. In the encoder, perform a convolution operation on the image using a locally weighted convolution kernel, and the weight calculation formula of the locally weighted convolution kernel is:

[0020]

[0021] where, W i,j represents the weighted value between the i-th pixel and the j-th pixel, x i ,k and x j ,k are the gray values of the i-th and j-th pixels in the k-th channel of the image respectively, α is a hyperparameter for adjusting the weighted sensitivity, n is the number of image channels, and is used as the inter-channel distance metric during calculation;

[0022] S33. Adopt a spatial attention mechanism to generate a spatial attention map:

[0023]

[0024] where, A i,j is the attention weight at the image position (i,j), f(x i,j ,m) is the feature value at the image position (i,j) in the m-th channel, w m is the weight coefficient corresponding to the channel, γ is a hyperparameter for controlling the intensity of attention distribution, the ReLU function is used to enhance the feature response, M is the number of channels, and N is the total number of pixels in the image;

[0025] S34. Fuse the output of the locally weighted convolution kernel and the spatial attention map. The specific fusion method is: multiply the convolution operation result and the spatial attention map element by element according to the pixel position, and the calculation formula is:

[0026]

[0027] where, F fuse (i,j) represents the final feature value at the pixel position (i,j) in the fused feature map, F conv (i,j,c) represents the pixel value in the output feature map after the locally weighted convolution operation, the convolution result of the c-th channel, A i,j,c is the attention weight corresponding to the channel in the spatial attention map, and C is the number of channels of the convolution image;

[0028] S35. According to the fused feature map, combined with the weight of the locally weighted convolution kernel, adjust the accuracy of the region boundary through the weighted response value:

[0029]

[0030] Among them, represents the response value after the weight adjustment of the local weighted convolution kernel based on the fused feature map, F fuse (i, j, c) is the pixel value in the fused feature map, the value of the c-th channel, W i,j is the weight of the local weighted convolution kernel, and C is the number of channels of the convolution image;

[0031] S36. Output the preliminary regional position of the concrete compression test block according to the adjusted response value.

[0032] Optionally, the S4 includes the following specific steps:

[0033] S41. Extract the feature map from the decoder of the DeepLabV3+ network according to the preliminary regional position of the test block. The feature map contains multi-channel outputs obtained by layer-by-layer upsampling through the decoder;

[0034] S42. Perform a convolution operation with a fixed stride on the feature map output by the decoder. The convolution operation is gradually weighted layer by layer:

[0035]

[0036] Among them, F l (i, j) is the pixel value of the output feature map after the e-th layer convolution operation, I l-1 (i, j) is the input feature map of the previous layer, W l (m, n) is the convolution kernel weight of the l-th layer, b l is the bias term of the l-th layer, and the convolution kernel size is k×k;

[0037] S43. After the convolution operation, perform a pixel-level weighted and spatial attention mechanism to calculate the weight coefficient α l :

[0038]

[0039] Among them, α l (i, j) is the weighted coefficient at the position (i, j), is the normalization factor of the weighted coefficient at the position (i, j), is the normalization factor of all position weighted coefficients;

[0040] S44. Input the weighted feature map into the activation function for processing. The activation function is a non-linear transformation with LeakyReLU:

[0041]

[0042] Among them, α is a constant less than 1, which is used to handle negative values to avoid the problem of gradient disappearance;

[0043] S45. Extract the edge features of the test block by applying the max pooling operation to the activated feature map, and calculate P l (i,j):

[0044]

[0045] Among them, P l (i,j) is the maximum value at the position (i, j) after pooling, {i-k,…,i+k} and {j-k,…,j+k} are the sizes of the pooling window, and F l (m,n) is the pixel value within the window in the feature map of the l-th layer;

[0046] S46. Input the pooled edge feature map into the fully connected layer for final prediction of the rectangular bounding box parameters. The coordinate calculation formula of the rectangular bounding box is:

[0047] B l =(σ(x center ),σ(y center ),σ(w box ),σ(h box ));

[0048] Among them, B l is the rectangular bounding box, x center and y center are the coordinates of the center point of the bounding box respectively, w box and h box are the width and height of the bounding box respectively, and σ(x) is the Sigmoid activation function to ensure that the output value is within the interval (0, 1);

[0049] S47. Calculate the position of the test block through the rectangular bounding box and generate the corresponding region identifier, and output the bounding box and region identifier of the test block.

[0050] Optionally, the S5 includes the following specific steps:

[0051] S51. Perform morphological processing on the output rectangular bounding box, and remove the noise area through opening operation to retain the edge information:

[0052]

[0053] Among them, O is the image after the opening operation, A is the input rectangular bounding box image, B is the structuring element, represents the dilation operation, represents the erosion operation;

[0054] S52. Optimize the rectangular bounding box after morphological processing using the maximum area method, and select the region with the largest area as the final positioning box:

[0055]

[0056] Among them, A max is the area of the maximum region, and are the pixel values of the input rectangular box A and the structural element B at the position (i, j) respectively, and N and M are the number of rows and columns of the rectangular bounding box;

[0057] S53. Fine-tune the position and size of the optimized rectangular bounding box, and correct the error through the introduction of a weighted loss function during the fine-tuning process:

[0058]

[0059] Among them, is the loss function after fine-tuning, α and β are the weight coefficients of the position error, and are the x and y coordinates of the center of the predicted rectangular bounding box respectively, and are the x and y coordinates of the center of the true rectangular bounding box, γ is the regularization coefficient, is the shape regularization term of the rectangular bounding box, which is used to constrain the shape and size of the box to be consistent;

[0060] S54. Optimize the edges and angles of the rectangular bounding box through the gradient descent method, and the gradient calculation formula during the optimization process is:

[0061]

[0062] Among them, θ opt is the parameter of the optimized rectangular box, θ init is the parameter of the initial rectangular box, η is the learning rate, is the gradient of the loss function with respect to the rectangular box parameter θ;

[0063] S55. Calculate the positioning error by comparing the position and size of the optimized rectangular bounding box with the true box:

[0064]

[0065] Among them, Error x and Error y represent the positioning errors of the rectangular bounding box in the horizontal and vertical directions respectively, with the unit of pixel, x opt and y opt are the center coordinates of the optimized rectangular box, x gtand y gt are the center coordinates of the true rectangular box, w gt and h gt are the width and height of the true rectangular box.

[0066] Optionally, S6 includes the following specific steps:

[0067] S61. Input the edge information of the rectangular bounding box into the Hough transform algorithm for line detection according to the positioning error and region identifier:

[0068] ρ = xcosθ + ysinθ;

[0069]

[0070] where ρ is the polar coordinate distance of the line, θ is the polar coordinate angle of the line, P(ρ,θ) is the cumulative voting value on the Hough plane, is the edge pixel value at the coordinate (x i , y i ) in the image, δ(·) is the Dirac δ function, and N is the number of edge pixels in the image;

[0071] S62. Calculate the maximum cumulative vote V max of the line and select the line with the maximum vote:

[0072]

[0073] where V max is the maximum vote value, w(x i , y i , θ) is the weight of the coordinate point (x i , y i ) at the angle θ, is the edge pixel value at the coordinate (x i , y i ) in the image, N is the number of edge pixels, and θ is the angle calculated in the Hough plane;

[0074] S63. Based on the line parameters ρ max and θ max of the maximum vote, calculate the intersection points of the line and the rectangular bounding box:

[0075]

[0076] where (x intersect , y intersect ) respectively represent the intersection coordinates of the line and the rectangular bounding box;

[0077] S64. Calculate the center position of the test block based on multiple intersection points, and perform secondary optimization on the positioning result. The final position of the optimized test block is:

[0078]

[0079] (x final , y final ) = (x center -Δx, y center -Δy);

[0080] where, (x center , y center ) is the center coordinate of the optimized test block, M is the number of selected straight lines, x intersecti and y intersecti are the coordinates of the intersection point of the i-th straight line and the rectangular frame, (x final , y final ) is the center coordinate of the finally located test block, and Δx and Δy are the corrected horizontal and vertical deviations;

[0081] S65. Based on the final result of the secondary optimization, output the position of the concrete compressive test block.

[0082] The beneficial effects of the present invention are as follows:

[0083] First of all, by introducing an image processing method based on deep learning, the present invention effectively improves the accuracy of automatic positioning of concrete compressive test blocks. Using the encoder and decoder structures of the DeepLabV3+ network, the present invention can accurately extract the edge features of the test block, and enhance the edge information through local weighted convolution kernels and spatial attention mechanisms, thereby realizing accurate prediction of the preliminary area position of the test block. This process has a high degree of automation and overcomes the errors and instabilities that may occur in traditional manual positioning.

[0084] Secondly, through further morphological processing and optimization algorithms, the present invention significantly improves the accuracy and stability of the positioning result. Specifically, after obtaining the preliminary rectangular bounding box, opening operation is used to remove the noise area, and the maximum area method is used to optimize the bounding box. This optimization method can effectively remove the invalid area, ensure more accurate positioning of the test block, and calculate the positioning error, providing a reliable basis for subsequent optimization.

[0085] Finally, the present invention combines the straight line detection method of the Hough transform to further optimize the positioning result and output the accurate position of the concrete test block. Through this method, the position of the test block can be secondarily optimized according to the positioning error and area identification, and finally ensure that the positioning result has high accuracy. In addition, the present invention is applicable to large-scale and batch concrete quality detection, has a high degree of automation, can effectively reduce manual intervention, improve work efficiency, and reduce labor costs. Description of the Drawings

[0086] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0087] Figure 1 is a flowchart of an automatic positioning method for concrete compression test blocks based on deep learning proposed by the present invention. Detailed Embodiment

[0088] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.

[0089] Reference Figure 1 , an automatic positioning method for concrete compression test blocks based on deep learning, includes the following steps:

[0090] S1. Obtain image data of the concrete compression test block through an image acquisition device;

[0091] S2. Preprocess the obtained image data;

[0092] S3. Input the preprocessed image data into the encoder of the DeepLabV3+ network, and extract the edge features of the concrete compression test block through a locally weighted convolution kernel and a spatial attention mechanism to generate a preliminary regional position of the test block;

[0093] S4. According to the preliminary regional position of the test block, reconstruct the edge information of the test block through the convolution operation with a fixed step size in the decoder of the DeepLabV3+ network, output a rectangular bounding box of the test block and generate a corresponding regional identifier;

[0094] S5. Perform morphological processing on the output rectangular bounding box, remove the noise area through opening operation, optimize the position and size of the rectangular bounding box using the maximum area method, locate the edge of the test block and calculate the positioning error;

[0095] S6. Combine the positioning error and the regional identifier, and use a line detection method based on the Hough transform to perform secondary optimization on the positioning result, and output the position of the concrete compression test block.

[0096] In this embodiment, the preprocessing includes image denoising, grayscale adjustment, and edge detection.

[0097] In this embodiment, S3 includes the following specific steps:

[0098] S31. Input the preprocessed image data into the encoder of the DeepLabV3+ network;

[0099] S32. In the encoder, a local weighted convolutional kernel is used to perform a convolution operation on the image. The weight calculation formula of the local weighted convolutional kernel is as follows:

[0100]

[0101] where Wi,j represents the weighted value between the i-th pixel and the j-th pixel, xi,k and xj,k are the gray values of the i-th and j-th pixels in the k-th channel of the image respectively, α is a hyperparameter for adjusting the weighted sensitivity, n is the number of image channels, and is used as the inter-channel distance metric during calculation;

[0102] S33. Adopt a spatial attention mechanism to generate a spatial attention map:

[0103]

[0104] where A i,j is the attention weight at the image position (i,j), f(x i,j ,m) is the feature value at the image position (i,j) in the m-th channel, w m is the weight coefficient corresponding to the channel, γ is a hyperparameter for controlling the intensity of attention distribution, the ReLU function is used to enhance the feature response, M is the number of channels, and N is the total number of pixels in the image;

[0105] S34. Fuse the output of the local weighted convolutional kernel and the spatial attention map. The specific fusion method is as follows: Multiply the convolution operation result and the spatial attention map element by element according to the pixel position. The calculation formula is:

[0106]

[0107] where F fuse (i,j) represents the final feature value at the pixel position (i,j) in the fused feature map, F conv (i,j,c) represents the pixel value in the output feature map after the local weighted convolution operation, the convolution result of the c-th channel, and A i,j,c is the attention weight of the corresponding channel in the spatial attention map, and C is the number of channels of the convolution image;

[0108] S35. According to the fused feature map, combined with the weight of the local weighted convolutional kernel, adjust the accuracy of the region boundary through the weighted response value:

[0109]

[0110] where represents the response value after being adjusted by the weight of the local weighted convolutional kernel based on the fused feature map, and Ffuse (i, j, c) is the pixel value in the fused feature map, the value of the c-th channel, and W i,j is the weight of the local weighted convolution kernel, and C is the number of channels of the convolutional image;

[0111] S36. Output the preliminary regional position of the concrete compressive test block according to the adjusted response value.

[0112] In this embodiment, S4 includes the following specific steps:

[0113] S41. Extract the feature map from the decoder of the DeepLabV3+ network according to the preliminary regional position of the test block. The feature map contains the multi-channel output obtained by layer-by-layer upsampling through the decoder;

[0114] S42. Perform a convolution operation with a fixed stride on the feature map output by the decoder. The convolution operation is gradually weighted in each layer:

[0115]

[0116] where F l (i, j) is the pixel value of the output feature map after the l-th layer of convolution operation, and I l-1 (i, j) is the input feature map of the previous layer, and W l (m, n) is the weight of the l-th layer convolution kernel, and b l is the bias term of the l-th layer, and the convolution kernel size is k×k;

[0117] S43. After the convolution operation, perform a pixel-level weighted and spatial attention mechanism to calculate the weight coefficient α l :

[0118]

[0119] where α l (i, j) is the weighted coefficient at the position (i, j), is the normalization factor of the weighted coefficient at the position (i, j), is the normalization factor of all position weighted coefficients;

[0120] S44. Input the weighted feature map into the activation function for processing. The activation function is a non-linear transformation with LeakyReLU:

[0121]

[0122] where α is a constant less than 1, used to process negative values to avoid the problem of gradient disappearance;

[0123] S45. Extract the edge features of the test block by applying the max pooling operation to the activated feature map, and calculate P l (i,j):

[0124]

[0125] where P l (i,j) is the maximum value at position (i, j) after pooling, {i - k, …, i + k} and {j - k, …, j + k} are the sizes of the pooling window, and F l (m,n) is the pixel value within the window in the feature map of the l-th layer;

[0126] S46. Input the pooled edge feature map into the fully connected layer for final prediction of the rectangular bounding box parameters. The coordinate calculation formula for the rectangular bounding box is:

[0127] B l =(σ(x center ), σ(y center ), σ(w box ), σ(h box ));

[0128] where B l is the rectangular bounding box, x center and y center are the coordinates of the center point of the bounding box respectively, w box and h box are the width and height of the bounding box respectively, and σ(x) is the Sigmoid activation function to ensure that the output value is within the range (0, 1);

[0129] S47. Calculate the position of the test block through the rectangular bounding box and generate the corresponding region identifier, and output the bounding box and region identifier of the test block.

[0130] In this embodiment, the S5 includes the following specific steps:

[0131] S51. Perform morphological processing on the output rectangular bounding box, and remove the noise area through opening operation to retain the edge information:

[0132]

[0133] where O is the image after opening operation, A is the input rectangular bounding box image, B is the structuring element, represents the dilation operation, represents the erosion operation;

[0134] S52. Optimize the rectangular bounding box after morphological processing using the largest area method, and select the region with the largest area as the final positioning box:

[0135]

[0136] Among them, A max is the area of the largest region, and are the pixel values of the input rectangle A and the structural element B at the position (i, j) respectively, and N and M are the number of rows and columns of the rectangular bounding box;

[0137] S53. Fine-tune the position and size of the optimized rectangular bounding box. The fine-tuning process corrects the error by introducing a weighted loss function:

[0138]

[0139] Among them, is the fine-tuned loss function, α and β are the weight coefficients of the position error, and are the x and y coordinates of the center of the predicted rectangular bounding box respectively, and are the x and y coordinates of the center of the true rectangular bounding box, γ is the regularization coefficient, is the shape regularization term of the rectangular bounding box, which is used to constrain the shape and size of the box to be consistent;

[0140] S54. Optimize the edges and angles of the rectangular bounding box by the gradient descent method. The gradient calculation formula in the optimization process is:

[0141]

[0142] Among them, θ opt is the parameter of the optimized rectangular box, θ init is the parameter of the initial rectangular box, η is the learning rate, is the gradient of the loss function with respect to the rectangular box parameter θ;

[0143] S55. Calculate the positioning error by comparing the position and size of the optimized rectangular bounding box with the true box:

[0144]

[0145] Among them, Error x and Error y represent the positioning errors of the rectangular bounding box in the horizontal and vertical directions respectively, with the unit of pixel, x opt and y opt are the center coordinates of the optimized rectangular box, x gt and y gt are the center coordinates of the true rectangular box, w gt and h gt are the width and height of the true rectangular box.

[0146] In this embodiment, S6 includes the following specific steps:

[0147] S61. Input the edge information of the rectangular bounding box into the Hough transform algorithm for line detection according to the positioning error and region identifier:

[0148] ρ = xcosθ + ysinθ;

[0149]

[0150] where ρ is the polar coordinate distance of the line, θ is the polar coordinate angle of the line, P(ρ,θ) is the cumulative voting value on the Hough plane, is the edge pixel value at the coordinate (x i , y i ) in the image, δ(·) is the Dirac δ function, and N is the number of edge pixels in the image;

[0151] S62. Calculate the maximum cumulative vote V of the line max and select the line with the maximum vote:

[0152]

[0153] where V max is the maximum vote value, w(x i , y i , θ) is the weight of the coordinate point (x i , y i ) at the angle θ, is the edge pixel value at the coordinate (x i , y i ) in the image, N is the number of edge pixels, and θ is the angle calculated in the Hough plane;

[0154] S63. Based on the line parameters ρ max and θ max of the maximum vote, calculate the intersection points of the line and the rectangular bounding box:

[0155]

[0156] where (x intersect , y intersect ) respectively represent the intersection coordinates of the line and the rectangular bounding box;

[0157] S64. Calculate the center position of the test block based on multiple intersection points, perform secondary optimization on the positioning result, and the final position of the optimized test block is:

[0158]

[0159] (x final , y final ) = (x center - Δx, y center - Δy);

[0160] Where (x center , y center ) is the center coordinate of the optimized test block, M is the number of selected straight lines, and are the coordinates of the intersection points of the i-th straight line and the rectangular frame, (x final , y final ) is the center coordinate of the finally located test block, and Δx and Δy are the corrected horizontal and vertical deviations;

[0161] S65. Output the position of the concrete compressive test block based on the final result of the secondary optimization.

[0162] Example 1:

[0163] To verify the feasibility of the present invention in implementation, the present invention is applied to

[0164] To verify the effectiveness of the present invention, we applied the present invention to the production site of a certain precast concrete factory for an automatic positioning experiment of concrete compressive test blocks. The factory is located in the building materials production park of a certain city, mainly produces various concrete products, and provides relevant quality inspection services for construction projects. In the past, the positioning of test blocks mainly relied on manual operation. The operators measured and positioned according to the surface characteristics of the test blocks. This method was not only inefficient but also easily affected by human factors, resulting in large positioning errors. Especially in the process of mass production, a large number of test blocks could not be processed timely and accurately, resulting in a long detection cycle and affecting production efficiency and quality control.

[0165] In practical applications, the present invention can achieve automatic positioning of concrete compressive test blocks through the combination of a deep learning network and an image processing algorithm, greatly improving the positioning accuracy and efficiency. In the experiment, concrete test block samples generated during the production process of the factory were used, and a total of 500 concrete compressive test blocks in different batches, including test blocks with different shapes, sizes, and surface textures, were tested. During the implementation process, the test block images were collected by an installed high-resolution camera, and the collected image data was input into the method of the present invention for processing.

[0166] The specific application process is as follows: First, the acquisition device takes pictures of concrete test blocks, and the image data is input into the encoder of the DeepLabV3+ network after preprocessing. Through local weighted convolution kernels and spatial attention mechanisms, this network can automatically identify the edge features of the test blocks and predict the preliminary regional positions of the test blocks. Then, the decoder of the network optimizes the preliminary region, reconstructs the edge information of the test blocks using convolutional operations with a fixed stride, outputs the rectangular bounding boxes of the test blocks, and generates region identifiers. Subsequently, through morphological processing and optimization using the largest region method, invalid regions are further removed, the edges of the test blocks are accurately located, and the positioning error is calculated. Finally, the accurate positions of the concrete test blocks are output.

[0167] To verify the feasibility and superiority of the method of the present invention in an actual production environment, we set up a comparative experiment. The automatic positioning method of the present invention was compared with the traditional manual positioning method, and data such as the time required for each method to process a single test block and the positioning error were recorded. The following is the experimental data table of the embodiment.

[0168] Table 1 Experimental data of automatic positioning of concrete test blocks

[0169]

[0170]

[0171] The experimental results show that the time for automatic positioning using the method of the present invention is much lower than that of manual positioning, and the positioning error is significantly reduced. Table 1 shows that the positioning time for each test block using the method of the present invention is 12 seconds, while manual positioning requires an average of 45 seconds. In terms of positioning accuracy, the positioning error of the method of the present invention is within 5 mm, while the error of manual positioning is generally above 10 mm.

[0172] As can be seen from Table 1, the method of the present invention shows significant advantages in both positioning time and accuracy. Especially in the positioning of a large number of test blocks, the advantages of the automated method are even more obvious, which can effectively improve production efficiency and greatly reduce the errors caused by manual operations. The experiment also shows that even in the case of complex test block shapes, the method of the present invention can still maintain a high positioning accuracy, while manual positioning faces higher errors and longer processing times.

[0173] In addition, the method of the present invention also has strong adaptability and scalability. On concrete test blocks of different sizes and shapes, the system can automatically adjust the positioning strategy, ensuring stability under various environmental conditions. During mass production, the present invention can quickly process a large number of test block images and automatically complete the positioning process, which not only improves efficiency but also reduces manual intervention and enhances the automation level of the production line.

[0174] In summary, through the application experiment in a certain precast concrete factory, the present invention has proven its feasibility and advantages in the automatic positioning of concrete compressive test blocks. By the method of the present invention, the positioning accuracy of concrete test blocks has been effectively improved, the positioning time has been greatly shortened, and the positioning error has been significantly reduced, especially suitable for large-scale production and batch detection scenarios. In practical applications, the present invention can effectively improve production efficiency, reduce labor costs, and enhance the accuracy of quality control, providing a reliable and automated solution for concrete quality detection.

[0175] As described above, only the preferred specific embodiments of the present invention are given, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, with equivalent replacement or change, should be covered within the protection scope of the present invention.

Claims

1. An automatic positioning method for concrete compressive test blocks based on deep learning, characterized in that, It includes the following steps: S1. Obtain the image data of the concrete compressive test block through an image acquisition device; S2. Preprocess the obtained image data; S3. Input the preprocessed image data into the encoder of the DeepLabV3+ network, extract the edge features of the concrete compressive test block through a locally weighted convolution kernel and a spatial attention mechanism, and generate the preliminary region position of the test block; S4. According to the preliminary region position of the test block, reconstruct the edge information of the test block through the convolution operation with a fixed stride in the decoder of the DeepLabV3+ network, output the rectangular bounding box of the test block and generate the corresponding region identifier; S5. Perform morphological processing on the output rectangular bounding box, remove the noise region through opening operation, optimize the position and size of the rectangular bounding box using the maximum area method, locate the edge of the test block and calculate the positioning error; S6. Combine the positioning error and the region identifier, and use the line detection method based on the Hough transform to perform secondary optimization on the positioning result, and output the position of the concrete compressive test block.

2. The automatic positioning method for concrete compressive test blocks based on deep learning according to claim 1, wherein, The preprocessing includes image denoising, grayscale adjustment and edge detection.

3. The automatic positioning method for concrete compressive test blocks based on deep learning according to claim 1, characterized in that, The S3 includes the following specific steps: S31. Input the preprocessed image data into the encoder of the DeepLabV3+ network; S32. In the encoder, perform a convolution operation on the image using a locally weighted convolution kernel, and the weight calculation formula of the locally weighted convolution kernel is: Among them, W i,j represents the weighted value between the i-th pixel and the j-th pixel, x i ,k and x j ,k are the gray values of the i-th and j-th pixels in the k-th channel of the image respectively, α is a hyperparameter for adjusting the weighted sensitivity, n is the number of image channels, and is used as the inter-channel distance metric; S33. Adopt a spatial attention mechanism to generate a spatial attention map: Among them, A i,j is the attention weight at the image position (i, j), f(x i,j , m) is the feature value at the image position (i, j) in the m-th channel, w m is the weight coefficient of the corresponding channel, γ is a hyperparameter that controls the intensity of the attention distribution, the ReLU function is used to enhance the feature response, M is the number of channels, and N is the total number of pixels in the image; S34. Fuse the output of the locally weighted convolution kernel and the spatial attention map. The specific fusion method is: multiply the convolution operation result and the spatial attention map element by element according to the pixel position, and the calculation formula is: Among them, F fuse (i,j) represents the final feature value of the pixel position (i,j) in the fused feature map, F conv (i,j,c) represents the pixel value in the output feature map after local weighted convolution operation, the convolution result of the c-th channel, A i,j,c is the attention weight of the corresponding channel in the spatial attention map, C is the number of channels of the convolutional image; S35. According to the fused feature map, combine the weight of the locally weighted convolution kernel, and adjust the accuracy of the region boundary through the weighted response value: Among them, represents the response value after the weight adjustment of the local weighted convolution kernel based on the fused feature map, F fuse (i, j, c) is the pixel value in the fused feature map, the value of the c-th channel, W i,j is the weight of the local weighted convolution kernel, and C is the number of channels of the convolution image; S36. According to the adjusted response value, output the preliminary region position of the concrete compressive test block.

4. A method for automatically locating concrete compressive test blocks based on deep learning according to claim 1, characterized in that, The S4 includes the following specific steps: S41. According to the preliminary region position of the test block, extract the feature map from the decoder of the DeepLabV3+ network. The feature map contains multi-channel outputs obtained by layer-by-layer upsampling through the decoder; S42. Perform a convolution operation with a fixed stride on the feature map output by the decoder, and the convolution operation is gradually weighted in each layer: Among them, F l (i, j) is the pixel value of the output feature map after the l-th layer of convolution operation, I l-1 (i, j) is the input feature map of the previous layer, W l (m, n) is the convolutional kernel weight of the l-th layer, b l is the bias term of the l-th layer, and the convolutional kernel size is k×k; S43. After the convolution operation, perform pixel-level weighting and spatial attention mechanism to calculate the weight coefficient α l : where α l (i, j) is the weighting coefficient at position (i, j), is the normalization factor of the weighting coefficient at position (i, j), is the normalization factor of all position weighting coefficients; S44. Input the weighted feature map into the activation function for processing. The activation function is a non-linear transformation with LeakyReLU: Among them, α is a constant less than 1, which is used to process negative values to avoid the problem of gradient disappearance; S45. Extract the edge features of the test block by applying the max pooling operation to the activated feature map, and calculate P l (i,j): where P l (i,j) is the maximum value at the pooled position (i, j), {i-k, …, i+k} and {j-k, …, j+k} are the sizes of the pooling window, and F l (m,n) is the pixel value within the window in the l-th layer feature map; S46. Input the pooled edge feature map into the fully connected layer to perform the final rectangular bounding box parameter prediction. The coordinate calculation formula of the rectangular bounding box is: B l = (σ(x center ), σ(y center ), σ(w box ), σ(h box )); Among them, B e is a rectangular bounding box, and x center and y center are the coordinates of the center point of the bounding box respectively, w box and h box are the width and height of the bounding box respectively, and σ(x) is the Sigmoid activation function to ensure that the output value is within the interval (0, 1); S47. Calculate the position of the test block through the rectangular bounding box and generate the corresponding region identifier, and output the bounding box and region identifier of the test block.

5. The automatic positioning method for concrete compressive test blocks based on deep learning according to claim 1, characterized in that, The S5 includes the following specific steps: S51. Perform morphological processing on the output rectangular bounding box, remove the noise region through opening operation, and retain the edge information: Among them, O is the image after opening operation, A is the input rectangular bounding box image, B is the structuring element, ⊕ represents the dilation operation, represents the erosion operation; S52. Optimize the rectangular bounding box after morphological processing using the maximum area method, and select the region with the largest area as the final positioning box: Among them, A max is the area of the largest region, and are the pixel values of the input rectangle A and the structural element B at the position (i, j), respectively; N and M are the number of rows and columns of the rectangular bounding box; S53. Fine-tune the position and size of the optimized rectangular bounding box. The error correction in the fine-tuning process is performed by introducing a weighted loss function: Among them, is the loss function after fine-tuning, where α and β are the weight coefficients of the position error, and are the x and y coordinates of the center of the predicted rectangular bounding box respectively, and are the x and y coordinates of the center of the ground-truth rectangular bounding box, and γ is the regularization coefficient, is the shape regularization term of the rectangular bounding box, which is used to constrain the shape and size of the box to be consistent; S54. Optimize the edges and angles of the rectangular bounding box by the gradient descent method. The gradient calculation formula in the optimization process is: Among them, θ opt is the optimized rectangular box parameter, θ init is the initial rectangular box parameter, η is the learning rate, is the gradient of the loss function with respect to the rectangular box parameter θ; S55. Calculate the positioning error by comparing the position and size of the optimized rectangular bounding box with the ground truth box: Among them, Error x and Error y respectively represent the positioning errors of the rectangular bounding box in the horizontal and vertical directions, with the unit of pixel, x opt and y opt are the center coordinates of the optimized rectangular box, x gt and y gt are the center coordinates of the true rectangular box, w gt and h gt are the width and height of the true rectangular box.

6. The automatic positioning method for concrete compressive test blocks based on deep learning according to claim 1, characterized in that, The said S6 includes the following specific steps: S61. Input the edge information of the rectangular bounding box into the Hough transform algorithm for line detection according to the positioning error and region identification: ρ = xcosθ + ysinθ; where ρ is the polar coordinate distance of the straight line, θ is the polar coordinate angle of the straight line, and P(ρ, θ) is the cumulative voting value on the Hough plane. is the edge pixel value at the coordinate (x i , y i ) in the image, δ(·) is the Dirac δ function, and N is the number of edge pixels in the image. S62. Calculate the maximum cumulative vote V of the straight line max And select the straight line with the maximum vote: Among them, V max is the maximum voting value, w(x i , y i , θ) is the weight of the coordinate point (x i , y i ) at the angle θ, is the edge pixel value at the coordinate (x i , y i ) in the image, N is the number of edge pixels, and θ is the angle calculated in the Hough plane; S63. Straight line parameter ρ based on maximum voting max and θ max , calculate the intersection point of the straight line and the rectangular bounding box: Among them, (x intersect , y intersect ) respectively represent the intersection coordinates of the straight line and the rectangular bounding box; S64. Calculate the center position of the test block based on multiple intersection points, and perform secondary optimization on the positioning result. The final position of the optimized test block is: (x final , y final ) = (x center - Δx, y center - Δy); Among them, (x center , y center ) is the center coordinate of the optimized test block, M is the number of selected straight lines, and are the coordinates of the intersection points of the i-th straight line and the rectangular frame, (x final , y final ) is the center coordinate of the finally located test block, and Δx and Δy are the corrected horizontal and vertical deviations; S65. Output the position of the concrete compressive test block based on the final result of the secondary optimization.

Citation Information

Cited By

  • High-precision edge extraction method based on automatic image processing

    CN121095587A