An image recognition and positioning method for component embedded parts based on target detection

Through the improved YOLOV8 deep learning network model and QR code pixel size calibration technology, the problem of low detection speed and accuracy in quality inspection of prefabricated components in prefabricated buildings is solved, and more efficient and accurate embedded parts detection is achieved.

CN119672717BActive Publication Date: 2025-05-13ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510195361.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-13
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The prior art has problems such as low detection speed and accuracy, missed inspection and missed inspection in the quality inspection of prefabricated components in prefabricated buildings, especially in engineering projects with large construction area and large production quantity.

Method used

The image recognition and positioning method of component embedded parts based on target detection is adopted, and the plane image of the prefabricated components is detected through the improved YOLOV8 deep learning network model, the pixel coordinates of the category and center position of the embedded parts are identified, and the pixel coordinates are converted into actual coordinates through QR code pixel size calibration technology.

Benefits of technology

It significantly improves the accuracy and efficiency of inspection, reduces the occurrence of missed inspections and missed inspections, and can more accurately locate and classify embedded parts, meeting the accuracy requirements of factory acceptance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672717B_ABST
    Figure CN119672717B_ABST
Patent Text Reader

Abstract

The present invention provides an image recognition and positioning method for component embedded parts based on target detection, improves the YOLOV8 deep learning network model, adds a CBAM attention mechanism module in the Backbone network, and replaces the SPPF module with the Improve_SPPF module; replaces the Conv module with the GSConv module in the Neck network; optimizes the bounding box regression loss function in the Head network; uses the improved YOLOV8 deep learning network model to identify and locate embedded parts in the plane image data set of prefabricated components, and uses the two-dimensional code pixel size calibration technology to calibrate the actual coordinates of the center position of the embedded parts. The present invention uses the above-mentioned improvements to effectively improve the recognition efficiency and accuracy of embedded parts such as wire boxes and embedded pipes in prefabricated components, and realizes accurate positioning of embedded parts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of prefabricated building detection technology, and in particular to an image recognition and positioning method for component embedded parts based on target detection. Background Art

[0002] Prefabricated buildings are an important future direction for the construction industry. They are made by pre-processing various types of building components in the factory and then transporting them to the site for reliable connection and assembly. Compared with traditional cast-in-place structural buildings, prefabricated buildings have the advantages of large-scale production, rapid construction and low cost.

[0003] In the prefabricated component production factory of assembled buildings, prefabricated components are the main building components produced. It is extremely important to control the production quality of prefabricated components, which mainly involves the inspection of the prefabricated component upper wire boxes and embedded pipes; if the number of wire boxes and embedded pipes is lost, it will cause rework and repair, affecting the production progress, and more seriously affecting the later construction use of the project, which may cause immeasurable losses to the entire construction project.

[0004] At present, the quality inspection of prefabricated components mostly adopts traditional methods, which are carried out manually on-site. However, for engineering projects with large construction areas and large production quantities, the manual input is large, the detection speed and accuracy are low, and there are phenomena of missed inspections and false inspections. Summary of the invention

[0005] The purpose of the present invention is to provide an image recognition and positioning method for component embedded parts based on target detection to solve the deficiencies of the above-mentioned prior art.

[0006] To achieve the above object, the present invention adopts the following technical solution:

[0007] An image recognition and positioning method for component embedded parts based on target detection comprises the following steps:

[0008] S1. Collect the plane images of prefabricated components and perform data enhancement processing and embedded parts annotation to construct a plane image dataset;

[0009] S2. Build an improved YOLOV8 deep learning network model: Based on the YOLOV8 deep learning network model, add the CBAM attention mechanism module to the Backbone network and replace the SPPF module with the Improve_SPPF module; replace the Conv module with the GSConv module in the Neck network; optimize the bounding box regression loss function in the Head network;

[0010] S3. Use the improved YOLOV8 deep learning network model to detect the plane image data set and identify the category of embedded parts in prefabricated components and the pixel coordinates of the center position;

[0011] S4. Use the QR code pixel size calibration technology to convert the pixel coordinates of the center position of the embedded part into actual coordinates.

[0012] Furthermore, the planar image dataset is constructed by the following steps:

[0013] S11, collecting a plane image of the prefabricated component by taking a photo;

[0014] S12, performing data enhancement processing including but not limited to image flipping, image cropping, and color conversion on the collected plane image;

[0015] S13. Use the LabelImg labeling tool to label the embedded parts in the plane image after data enhancement processing, and construct a plane image dataset.

[0016] Furthermore, the use of the improved YOLOV8 deep learning network model to detect the plane image data set and identify the category of the embedded parts in the prefabricated components and the pixel coordinates of the center position is specifically achieved by the following steps:

[0017] S31, taking the plane image dataset as the input feature map, extracting image features from the input feature map through the Conv layer, C2f module, CBAM attention mechanism module and Improve_SPPF module in the Backbone network;

[0018] S32, the feature map extracted by the Backbone network is used as the input of the Neck network, and the feature fusion is performed through the Upsample layer, the Concat layer, the C2f module, and the GSConv module to enhance the feature representation capability of the feature map;

[0019] S33. The feature map obtained by fusing the Neck network is used as the input of the Head network, and the category of the embedded parts and the pixel coordinates of the center position are identified through the Detect layer.

[0020] Furthermore, performing image feature extraction on a planar image dataset through the Backbone network specifically includes the following steps:

[0021] S311, first use two Conv layers to perform convolution downsampling on the feature map in sequence and then input it into the C2f module;

[0022] S312, in the C2f module, the input feature map is evenly split into two parts in the channel dimension, one part is concatenated with the other part after being processed by convolution, normalization and activation operations through several Bottleneck modules, and then convolution is performed once in the Conv layer, and the concatenated feature map is restored to the original dimension and output;

[0023] The processing process of the C2f module is expressed by the following formula (1):

[0024] (1);

[0025] In formula (1): Represents the input of the C2f module; Represents the output of the C2f module; Represents the Bottleneck module; Indicates the number of Bottleneck modules; represents the convolutional layer; Indicates that the splicing operation is performed on the channel dimension;

[0026] S313, repeating the combination structure of a Conv layer and a C2f module three times in sequence, and inputting the output feature map into the CBAM attention mechanism module for processing;

[0027] The CBAM attention mechanism module includes a spatial attention mechanism and a channel attention mechanism. The specific processing process of the CBAM attention mechanism module is expressed by the following formulas (2)-(3):

[0028] (2);

[0029] (3);

[0030] in:

[0031]

[0032] (4);

[0033]

[0034] (5);

[0035] In formulas (2)-(5): is the input feature map; represents the channel attention weight matrix; Represents the cross product operation; represents the output of the channel attention mechanism, represents the spatial attention weight matrix; represents the output of the spatial attention mechanism; represents the average pooling operation; Represents the maximum pooling operation; for abbreviation of; for abbreviation of; represents a multi-layer perceptron; express Activation function; and express The weight of Represents a splicing operation; The filter is Convolution operation; for abbreviation of; for abbreviation of;

[0036] S314, input the feature map processed by the CBAM attention mechanism module into the Improve_SPPF module for processing;

[0037] The Improve_SPPF module combines the average pooling method and the maximum pooling method to process the input feature map. The specific processing process of the Improve_SPPF module is expressed by the following formula (6):

[0038] (6);

[0039] in:

[0040] (7);

[0041] (8);

[0042] (9);

[0043] (10);

[0044] In formulas (6)-(10): Indicates output; Represents the convolution operation; Represents a splicing operation; Represents input; express Output after convolution operation; represents the average pooling operation; Represents the maximum pooling operation; Express The output after an average pooling operation and a maximum pooling operation; Express Output after two average pooling operations and a maximum pooling operation; Express Output after three average pooling operations and a max pooling operation.

[0045] Furthermore, performing image feature fusion on the feature map output by the Backbone network through the Neck network specifically includes the following steps:

[0046] S321, the feature map output by the Backbone network is used as the input of the Neck network, and is sequentially processed by the Upsample upsampling, Concat splicing, and C2f module loop twice, and then input into the GSConv module. At this time, the number of channels of the feature map is recorded as C1;

[0047] S322, in the GSConv module, the feature map is first subjected to the Conv convolution process and the DWConv depth-separable convolution process in sequence, at which time the number of channels is recorded as C2 / 2, and then the feature maps obtained by the two convolution processes are subjected to the Concat splicing operation to obtain a feature map with a channel number of C2;

[0048] S323, after the feature map with the number of channels C2 is subjected to Concat concatenation, C2f module processing and shuffling operation, the final feature map is output for detection.

[0049] Furthermore, the optimized bounding box regression loss function is expressed by the following formula (11):

[0050] (11);

[0051] in:

[0052] (12);

[0053] (13);

[0054] (14);

[0055] (15);

[0056] (16);

[0057] (17);

[0058] (18);

[0059] (19);

[0060] (20);

[0061] (twenty one);

[0062] In formula (11)-(21), represents intersection and union ratio; Represents the prediction box; represents the real frame; Represents the area of ​​the intersection between the predicted box and the true box, Represents the area of ​​the union of the predicted box and the true box; , , , Represent the prediction box The coordinates of the four vertices; , , , Represents the real box The coordinates of the four vertices; Represents the prediction box With real frame The distance between the upper left corner vertices; Represents the prediction box With real frame The distance between the upper right corner vertices; Represents the prediction box With real frame The distance between the lower left corner vertices; Represents the prediction box With real frame The distance between the lower right corner vertices; Represents the height of the input feature map; Represents the width of the input feature map; Represents the diagonal distance of the input feature map; Represents the prediction box The center point coordinates, ; Represents the real box The center point coordinates, ; Represents the prediction box With real frame The direct distance from the center point; It means that the prediction box can be With real frame The diagonal distance of the smallest enclosing frame; represents the weight parameter; Used to measure the consistency of aspect ratio; , Represents the prediction box width and height; , Represents the real box width and height; Represents the inverse tangent function.

[0063] Furthermore, the method of converting the pixel coordinates of the center position of the embedded part into actual coordinates by using the two-dimensional code pixel size calibration technology specifically includes the following steps:

[0064] S41. Use the pyzbar library in the python open source library to recognize the QR code on the plane image;

[0065] The two-dimensional code is a mark attached to each prefabricated component during the production process for recording production information;

[0066] S42, obtaining the coordinates of the four vertices of the two-dimensional code, and calculating the pixel sizes of the four sides of the two-dimensional code according to the distance formula;

[0067] Specifically, the four vertex coordinates of the QR code are marked as , , , , the pixel sizes of the four sides are recorded as , , , , according to the distance formula, the pixel sizes of the four edges are calculated by the following formulas (22)-(25):

[0068] (twenty two);

[0069] (twenty three);

[0070] (twenty four);

[0071] (25);

[0072] S43, taking the average of the pixel sizes of the four sides of the two-dimensional code as the pixel size of the side length of the two-dimensional code, and calculating the conversion ratio between the pixel size and the actual physical size in combination with the actual physical size of the two-dimensional code, thereby completing the pixel size calibration;

[0073] In this step, the specific calculation process is expressed by the following formulas (26)-(27):

[0074] (26);

[0075] (27);

[0076] In formulas (26)-(27): The pixel size of the side length, in pixels; It is the actual physical size in mm; Indicates the conversion ratio between pixel size and actual physical size;

[0077] S44, multiplying the pixel coordinates of the center position of the embedded component detected in the prefabricated component image by the calculated conversion ratio to obtain the actual coordinates of the center position of the embedded component in the prefabricated component.

[0078] It can be seen from the above technical solutions that the present invention has the following advantages compared with the prior art:

[0079] (1) The present invention introduces the CBAM attention mechanism module into the Backbone network to enhance the feature extraction of key information by the Backbone network, which can significantly improve the detection accuracy without significantly increasing the computational complexity.

[0080] (2) The present invention improves the SPPF module in the Backbone network and records the improved SPPF as Improve_SPPF. Improve_SPPF introduces the AvgPool average pooling method, combines the average pooling and maximum pooling methods to improve the extraction of global feature information, and enhances the information fusion between features at different levels.

[0081] (3) In the present invention, the Conv convolution layer is replaced by the GSConv grouped spatial convolution layer in the Neck network, which can retain multi-channel information, enhance the feature expression ability of the image, and reduce the amount of calculation.

[0082] (4) The present invention optimizes the bounding box regression loss function of the network model, which can more accurately reflect the difference between the predicted box and the real box, improve the positioning accuracy and classification performance of the embedded parts, and accelerate the speed of bounding box regression.

[0083] (5) The present invention uses the QR code pixel size calibration technology to convert the pixel size of the target center position detected in the image into the actual physical size in the real world, thereby completing the measurement of the center position of the prefabricated component upper wire box and the embedded pipe. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 It is a schematic flow chart of the method of the present invention;

[0085] Figure 2It is a schematic diagram of the network structure of the improved YOLOV8 deep learning network model of the present invention;

[0086] Figure 3 It is a schematic diagram of the network structure of the CBAM module of the present invention;

[0087] Figure 4 It is a schematic diagram of the network structure of the Improve_SPPF module of the present invention;

[0088] Figure 5 Schematic diagram of the network structure of the GSConv module of the present invention;

[0089] Figure 6 This is a visual inspection and recognition effect diagram in an embodiment of the present invention;

[0090] Figure 7 This is a visual inspection and recognition effect diagram in an embodiment of the present invention;

[0091] Figure 8 It is a visual inspection and recognition effect diagram in an embodiment of the present invention. DETAILED DESCRIPTION

[0092] A preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings.

[0093] like Figure 1 The image recognition and positioning method of component embedded parts based on target detection shown includes the following steps:

[0094] S1. Collect the plane images of prefabricated components and perform data enhancement processing and embedded parts annotation to construct a plane image dataset;

[0095] This is achieved through the following steps:

[0096] S11, collecting a plane image of the prefabricated component by taking a photo;

[0097] S12. Perform data enhancement processing on the collected plane images, including but not limited to image flipping, image cropping, and color conversion. Image flipping allows the model to receive samples from more angles, helping the model learn more different perspectives. Image cropping can remove unnecessary information in the image, making the detection target more prominent and improving the image quality and recognition accuracy. Color conversion can change the color, brightness, contrast and other attributes of the image, create new image samples, and enhance the model's robustness to lighting conditions and color changes.

[0098] S13. Use the LabelImg labeling tool to label the embedded parts in the plane image after data enhancement processing, and construct a plane image data set; the embedded parts include but are not limited to wire boxes and embedded pipes.

[0099] S2. Construct an improved YOLOV8 deep learning network model; the YOLOV8 deep learning network model includes a Backbone network, a Neck network and a Head network; the improvement of the YOLOV8 deep learning network model specifically includes: adding a CBAM attention mechanism module to the Backbone network, and replacing the SPPF module with an Improve_SPPF module; replacing the Conv module with a GSConv module in the Neck network; optimizing the bounding box regression loss function in the Head network; the network structure of the improved YOLOV8 deep learning network model is as follows Figure 2 shown.

[0100] S3, using the improved YOLOV8 deep learning network model to detect the plane image data set, identifying the category of embedded parts in the prefabricated components and the pixel coordinates of the center position; specifically including the following steps:

[0101] S31. Take the plane image dataset as the input feature map, and extract image features from the input feature map through the Conv layer, C2f module, CBAM attention mechanism module and Improve_SPPF module in the Backbone network:

[0102] S311, first use two Conv layers to perform convolution downsampling on the feature map in sequence and then input it into the C2f module;

[0103] S312, in the C2f module, the input feature map is evenly split into two parts in the channel dimension, one part is concatenated with the other part after being processed by convolution, normalization and activation operations through several Bottleneck modules, and then convolution is performed once in the Conv layer, and the concatenated feature map is restored to the original dimension and output;

[0104] The processing process of the C2f module is expressed by the following formula (1):

[0105] (1);

[0106] In formula (1): Represents the input of the C2f module; Represents the output of the C2f module; Represents the Bottleneck module; Indicates the number of Bottleneck modules; represents the convolutional layer; Indicates that the splicing operation is performed on the channel dimension;

[0107] S313, repeating the combination structure of a Conv layer and a C2f module three times in sequence, and inputting the output feature map into the CBAM attention mechanism module for processing;

[0108] like Figure 3 As shown in FIG. 1 , the CBAM attention mechanism module includes a channel attention mechanism (Channel Attention Module) and a spatial attention mechanism (Spatial Attention Module). The channel attention mechanism can adaptively learn the importance of each channel feature of the image to enhance the expression ability of the feature image, while the spatial attention mechanism enhances the discrimination ability of the feature image by learning the importance of each spatial feature of the image. By introducing the CBAM attention mechanism, the feature extraction of key information by the Backbone network can be effectively enhanced, which can significantly improve the detection accuracy without significantly increasing the computational complexity. The specific processing process of the CBAM attention mechanism module is expressed by the following formulas (2)-(3):

[0109] (2);

[0110] (3);

[0111] in:

[0112]

[0113] (4);

[0114]

[0115] (5);

[0116] In formulas (2)-(5): is the input feature map; represents the channel attention weight matrix; Represents the cross product operation; represents the output of the channel attention mechanism, represents the spatial attention weight matrix; represents the output of the spatial attention mechanism; represents the average pooling operation; Represents the maximum pooling operation; for abbreviation of; for abbreviation of; represents a multi-layer perceptron; express Activation function; and express The weight of Represents a splicing operation; The filter is Convolution operation; for abbreviation of; for abbreviation of;

[0117] S314, input the feature map processed by the CBAM attention mechanism module into the Improve_SPPF module for processing;

[0118] like Figure 4 As shown, the Improve_SPPF module combines the average pooling method and the maximum pooling method to process the input feature map to improve the extraction of global feature information and enhance the information fusion between features at different levels. The specific processing process of the Improve_SPPF module is expressed by the following formula (6):

[0119] (6);

[0120] in:

[0121] (7);

[0122] (8);

[0123] (9);

[0124] (10);

[0125] In formulas (6)-(10): Indicates output; Represents the convolution operation; Represents a splicing operation; Represents input; express Output after convolution operation; represents the average pooling operation; Represents the maximum pooling operation; Express The output after an average pooling operation and a maximum pooling operation; Express Output after two average pooling operations and a maximum pooling operation; Express Output after three average pooling operations and maximum pooling operations;

[0126] S32. The feature map extracted by the Backbone network is used as the input of the Neck network. The feature fusion is performed through the Upsample layer, Concat layer, C2f module, and GSConv module to enhance the feature representation capability of the feature map:

[0127] S321, the feature map output by the Backbone network is used as the input of the Neck network, and is processed by the Upsample module, the Concat module, and the C2f module twice in sequence, and then input into the GSConv module. At this time, the number of channels of the feature map is recorded as C1;

[0128] S322, such as Figure 5 As shown, in the GSConv module, the feature map is first subjected to Conv convolution processing and DWConv depth separable convolution processing in sequence, at which time the number of channels is recorded as C2 / 2, and then the feature maps obtained by the two convolution processes are subjected to Concat splicing operation to obtain a feature map with a channel number of C2; the GSConv module can retain multi-channel information, enhance the feature expression ability of the image, and reduce the amount of calculation;

[0129] S323, after the feature map with the number of channels C2 is subjected to Concat concatenation, C2f module processing and shuffle operation (Shuffle), the final feature map is output for detection; through the shuffle operation, the channels of the feature map can be shuffled and rearranged, so that the model can capture more complex features and enhance the recognition ability of the network model for the target features;

[0130] S33. The feature map obtained by fusing the Neck network is used as the input of the Head network, and the category of the embedded parts and the pixel coordinates of the center position are identified through the Detect layer.

[0131] S4. Use the QR code pixel size calibration technology to convert the pixel coordinates of the center position of the embedded part into actual coordinates.

[0132] The optimized bounding box regression loss function of this preferred embodiment can more accurately reflect the difference between the predicted box and the real box, improve the positioning accuracy and classification performance of the embedded parts, and speed up the bounding box regression. The optimized bounding box regression loss function is expressed by the following formula (11):

[0133] (11);

[0134] in:

[0135] (12);

[0136] (13);

[0137] (14);

[0138] (15);

[0139] (16);

[0140] (17);

[0141] (18);

[0142] (19);

[0143] (20);

[0144] (twenty one);

[0145] In formula (11)-(21), represents intersection and union ratio; Represents the prediction box; represents the real frame; Represents the area of ​​the intersection between the predicted box and the true box, Represents the area of ​​the union of the predicted box and the true box; , , , Represent the prediction box The coordinates of the four vertices; , , , Represents the real box The coordinates of the four vertices; Represents the prediction box With real box The distance between the upper left corner vertices; Represents the prediction box With real box The distance between the upper right corner vertices; Represents the prediction box With real box The distance between the lower left corner vertices; Represents the prediction box With real box The distance between the lower right corner vertices; Represents the height of the input feature map; Represents the width of the input feature map; Represents the diagonal distance of the input feature map; Represents the prediction box The center point coordinates, ; Represents the real box The center point coordinates, ; Represents the prediction box With real frame The direct distance from the center point; It means that the prediction box can be With real frame The diagonal distance of the smallest enclosing frame; represents the weight parameter; Used to measure the consistency of aspect ratio; , Represents the prediction box width and height; , Represents the real box width and height; Represents the inverse tangent function.

[0146] The method of calibrating the center position of the embedded part using the two-dimensional code pixel size calibration technology in this preferred embodiment specifically includes the following steps:

[0147] S41. Use the pyzbar library in the python open source library to recognize the QR code on the plane image;

[0148] The two-dimensional code is a mark attached to each prefabricated component during the production process for recording production information;

[0149] S42, obtaining the coordinates of the four vertices of the two-dimensional code, and calculating the pixel sizes of the four sides of the two-dimensional code according to the distance formula;

[0150] Specifically, the four vertex coordinates of the QR code are marked as , , , , the pixel sizes of the four sides are recorded as , , , , according to the distance formula, the pixel sizes of the four edges are calculated by the following formulas (22)-(25):

[0151] (twenty two);

[0152] (twenty three);

[0153] (twenty four);

[0154] (25);

[0155] S43, taking the average of the pixel sizes of the four sides of the two-dimensional code as the pixel size of the side length of the two-dimensional code, and calculating the conversion ratio between the pixel size and the actual physical size in combination with the actual physical size of the two-dimensional code, thereby completing the pixel size calibration;

[0156] In this step, the specific calculation process is expressed by the following formulas (26)-(27):

[0157] (26);

[0158] (27);

[0159] In formulas (26)-(27): The pixel size of the side length, in pixels; It is the actual physical size in mm; Indicates the conversion ratio between pixel size and actual physical size;

[0160] S44, multiplying the pixel coordinates of the center position of the embedded component detected in the prefabricated component image by the calculated conversion ratio to obtain the actual coordinates of the center position of the embedded component in the prefabricated component.

[0161] Training and testing:

[0162] In specific use, in order to ensure the recognition effect of the improved YOLOV8 deep learning network model, the improved YOLOV8 deep learning network model can be trained and tested, specifically including: referring to the above steps S11-S13, the plane image of the prefabricated component is collected by taking pictures, and the collected plane image is subjected to data enhancement processing including but not limited to image flipping, image cropping, and color conversion; then the LabelImg annotation tool is used to separately annotate the embedded parts including but not limited to the wire box and the embedded pipe in the plane image after data enhancement processing, and a plane image data set for model training and testing is constructed, and divided into a training set and a test set according to a ratio of 4:1; and the divided training set and test set are used to train and test the improved YOLOV8 deep learning network model.

[0163] Effect evaluation:

[0164] This preferred embodiment uses three evaluation indicators, namely, average precision (P), average recall (R), and average mAP50, to evaluate the original YOLOv8 deep learning network model and the improved YOLOv8 deep learning network model proposed in the present invention. The evaluation results are shown in Table 1 below. Specifically, in Table 1, YOLOv8 is the original YOLOv8 deep learning network model, and the improved YOLOv8 is the improved YOLOV8 deep learning network model described in the present invention.

[0165] In the detection results of the original YOLOV8 deep learning network model, the average precision (P) of the wire box and the pre-buried pipe reached 90.7%, the average recall (R) reached 91.5%, and the average mAP50 reached 92.3%; in the detection results of the improved YOLOV8 deep learning network model, the average precision (P) of the wire box and the pre-buried pipe reached 95.2%, the average recall (R) reached 95.5%, and the average mAP50 reached 96.1%. The visual inspection and recognition effect can be seen in Figure 6 , Figure 7 , Figure 8 ; In the figure, embedded_pipe means embedded pipe, and wire_box means wire box.

[0166] Table 1

[0167]

[0168] Table 2 shows the measured values ​​of the junction box and embedded pipe calculated by the QR code pixel size calibration technology, and compared with the theoretical values. By calculating the absolute error, it can be found that the errors are all within 5 mm, which meets the actual factory acceptance standards and verifies the effectiveness of the QR code pixel size calibration technology.

[0169] Table 2

[0170]

[0171] The above-described embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. The image recognition and positioning method of component embedded parts based on target detection is characterized in that: The following steps are involved: S1. Collect the plane images of prefabricated components and perform data enhancement processing and embedded parts annotation to construct a plane image dataset; S2. Build an improved YOLOV8 deep learning network model: Based on the YOLOV8 deep learning network model, add the CBAM attention mechanism module to the Backbone network and replace the SPPF module with the Improve_SPPF module; replace the Conv module with the GSConv module in the Neck network; optimize the bounding box regression loss function in the Head network; S3. Use the improved YOLOV8 deep learning network model to detect the plane image data set and identify the category of embedded parts in the prefabricated components and the pixel coordinates of the center position. This is achieved by the following steps: S31, taking the plane image dataset as the input feature map, extracting image features from the input feature map through the Conv layer, C2f module, CBAM attention mechanism module and Improve_SPPF module in the Backbone network; S311, first use two Conv layers to perform convolution downsampling on the feature map in sequence and then input it into the C2f module; S312, in the C2f module, the input feature map is evenly split into two parts in the channel dimension, one part is concatenated with the other part after being processed by convolution, normalization and activation operations through several Bottleneck modules, and then convolution is performed once in the Conv layer, and the concatenated feature map is restored to the original dimension and output; The processing process of the C2f module is expressed by the following formula (1): (1); In formula (1): Represents the input of the C2f module; Represents the output of the C2f module; Represents the Bottleneck module; Indicates the number of Bottleneck modules; represents the convolutional layer; Indicates that the splicing operation is performed on the channel dimension; S313, repeating the combination structure of a Conv layer and a C2f module three times in sequence, and inputting the output feature map into the CBAM attention mechanism module for processing; The CBAM attention mechanism module includes a spatial attention mechanism and a channel attention mechanism. The specific processing process of the CBAM attention mechanism module is expressed by the following formulas (2)-(3): (2); (3); in: (4); (5); In formulas (2)-(5): is the input feature map; represents the channel attention weight matrix; Represents the cross product operation; Represents the output of the channel attention mechanism; represents the spatial attention weight matrix; represents the output of the spatial attention mechanism; represents the average pooling operation; Represents the maximum pooling operation; for abbreviation of; for abbreviation of; represents a multi-layer perceptron; express Activation function; and express The weight of Represents a splicing operation; The filter is Convolution operation; for abbreviation of; for abbreviation of; S314, input the feature map processed by the CBAM attention mechanism module into the Improve_SPPF module for processing; The Improve_SPPF module combines the average pooling method and the maximum pooling method to process the input feature map. The specific processing process of the Improve_SPPF module is expressed by the following formula (6): (6); in: (7); (8); (9); (10); In formulas (6)-(10): Indicates output; Represents the convolution operation; Represents a splicing operation; Represents input; express Output after convolution operation; represents the average pooling operation; Represents the maximum pooling operation; Express The output after an average pooling operation and a maximum pooling operation; Express Output after two average pooling operations and a maximum pooling operation; Express Output after three average pooling operations and maximum pooling operations; S32, the feature map extracted by the Backbone network is used as the input of the Neck network, and the feature fusion is performed through the Upsample layer, the Concat layer, the C2f module, and the GSConv module to enhance the feature representation capability of the feature map; S33, using the feature map obtained by fusing the Neck network as the input of the Head network, and identifying the category of the embedded parts and the pixel coordinates of the center position through the Detect layer; S4. Use the QR code pixel size calibration technology to convert the pixel coordinates of the center position of the embedded part into actual coordinates.

2. The image recognition and positioning method for component embedded parts based on target detection according to claim 1 is characterized in that: The planar image dataset is constructed by the following steps: S11, collecting a plane image of the prefabricated component by taking a photo; S12, performing data enhancement processing including but not limited to image flipping, image cropping, and color conversion on the collected plane image; S13. Use the LabelImg labeling tool to label the embedded parts in the plane image after data enhancement processing, and construct a plane image dataset.

3. The image recognition and positioning method of component embedded parts based on target detection according to claim 1 is characterized in that: The image feature fusion of the feature map output by the Backbone network through the Neck network specifically includes the following steps: S321, the feature map output by the Backbone network is used as the input of the Neck network, and is sequentially processed by the Upsample upsampling, Concat splicing, and C2f module loop twice, and then input into the GSConv module. At this time, the number of channels of the feature map is recorded as C1; S322, in the GSConv module, the feature map is first subjected to the Conv convolution process and the DWConv depth-separable convolution process in sequence, at which time the number of channels is recorded as C2 / 2, and then the feature maps obtained by the two convolution processes are subjected to the Concat splicing operation to obtain a feature map with a channel number of C2; S323, after the feature map with the number of channels C2 is subjected to Concat concatenation, C2f module processing and shuffling operation, the final feature map is output for detection.

4. The image recognition and positioning method of component embedded parts based on target detection according to claim 1 is characterized in that: The optimized bounding box regression loss function is expressed by the following formula (11): (11); in: (12); (13); (14); (15); (16); (17); (18); (19); (20); (21); In formula (11)-(21), represents intersection and union ratio; Represents the prediction box; represents the real frame; Represents the area of ​​the intersection between the predicted box and the true box, Represents the area of ​​the union of the predicted box and the true box; , , , Represent the prediction box The coordinates of the four vertices; , , , Represents the real box The coordinates of the four vertices; Represents the prediction box With real frame The distance between the upper left corner vertices; Represents the prediction box With real frame The distance between the upper right corner vertices; Represents the prediction box With real frame The distance between the lower left corner vertices; Represents the prediction box With real frame The distance between the lower right corner vertices; Represents the height of the input feature map; Represents the width of the input feature map; Represents the diagonal distance of the input feature map; Represents the prediction box The center point coordinates, ; Represents the real box The center point coordinates, ; Represents the prediction box With real frame The direct distance from the center point; It means that the prediction box can be With real frame The diagonal distance of the smallest enclosing frame; represents the weight parameter; Used to measure the consistency of aspect ratio; , Represents the prediction box width and height; , Represents the real box width and height; Represents the inverse tangent function.

5. The image recognition and positioning method of component embedded parts based on target detection according to claim 1 is characterized in that: The method of converting the pixel coordinates of the center position of the embedded part into actual coordinates by using the two-dimensional code pixel size calibration technology specifically includes the following steps: S41. Use the pyzbar library in the python open source library to recognize the QR code on the plane image; The two-dimensional code is a mark attached to each prefabricated component during the production process for recording production information; S42, obtaining the coordinates of the four vertices of the two-dimensional code, and calculating the pixel sizes of the four sides of the two-dimensional code according to the distance formula; Specifically, the four vertex coordinates of the QR code are marked as , , , , the pixel sizes of the four sides are recorded as , , , , according to the distance formula, the pixel sizes of the four edges are calculated by the following formulas (22)-(25): (22); (23); (24); (25); S43, taking the average of the pixel sizes of the four sides of the two-dimensional code as the pixel size of the side length of the two-dimensional code, and calculating the conversion ratio between the pixel size and the actual physical size in combination with the actual physical size of the two-dimensional code, thereby completing the pixel size calibration; In this step, the specific calculation process is expressed by the following formulas (26)-(27): (26); (27); In formulas (26)-(27): The pixel size of the side length, in pixels; It is the actual physical size in mm; Indicates the conversion ratio between pixel size and actual physical size; S44, multiplying the pixel coordinates of the center position of the embedded component detected in the prefabricated component image by the calculated conversion ratio to obtain the actual coordinates of the center position of the embedded component in the prefabricated component.

Citation Information

Patent Citations

  • Embedded part detection method, device and system and storage medium

    CN118447023A

  • Deep learning-based pointer type instrument reading identification method

    CN119049026A