Precast concrete slab non-contact online size quality detection method and system based on multi-modal fusion

Through a non-contact detection method based on multimodal fusion, combined with image acquisition and deep learning, real-time, efficient, and non-destructive dimensional quality detection of precast concrete panels is achieved, solving the problems of time-consuming and labor-intensive traditional detection methods and high costs of automated equipment, and improving detection efficiency and accuracy.

CN120689336APending Publication Date: 2025-09-23CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510853215.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve efficient, accurate and seamless dimensional quality inspection of precast concrete panels. Traditional methods are time-consuming, labor-intensive and error-prone, while automated inspection equipment is costly, complex and difficult to promote.

Method used

A non-contact detection method based on multimodal fusion is adopted, combining image acquisition, deep learning and BIM model. The YOLOv8m model is used for target recognition and classification, HQ-SAM is used for pixel-level segmentation, and laser triggering is used to achieve automated detection.

Benefits of technology

It realizes real-time, non-destructive and efficient detection of precast concrete panels, and can simultaneously measure the characteristics of multiple types of components such as laminates, holes and electrical boxes, meeting the needs of industrial mass production and improving detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689336A_ABST
    Figure CN120689336A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of quality detection, in particular to a precast concrete slab non-contact online size quality detection method and system based on multi-modal fusion. The method comprises the following steps: carrying out image acquisition on a precast concrete slab, and preprocessing the image; selecting a YOLOv8m model as a reference model, constructing a pyramid network for enhancing context feature fusion, introducing shape-IOU, identifying and classifying targets in the image, and generating an ROI region; carrying out pixel-level segmentation by adopting HQ-SAM as a segmentation model to obtain pixel-level segmentation masks of the laminated board and the holes, and adjusting boundary frame coordinates of the ROI; respectively carrying out size measurement on the prediction frames of the laminated board, the hole and the electrical box; and extracting size design data from the BIM model, and comparing a visual detection result as production data of the component with design information. According to the technical scheme, the size quality of the prefabricated part in the production process can be detected in real time, and detection efficiency and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of quality inspection, and in particular to a non-contact online dimensional quality inspection method and system for precast concrete slabs based on multimodal fusion. Background Art

[0002] As the most prefabricated horizontal component in precast building systems, the dimensional quality of precast concrete panels (PCS) is crucial to the functional performance of the building system. Accurate dimensional quality not only ensures component compatibility during assembly, avoiding costly repairs and delays caused by dimensional mismatches, but also significantly improves construction efficiency and reduces costs. According to research by the Construction Industry Research Institute (CII), the cost of reworking components due to construction defects averages 5% of total construction costs. Therefore, systematic dimensional quality inspection (DQI) of precast concrete components is crucial to ensuring project success and on-time completion.

[0003] Traditional methods for inspecting the dimensional quality of precast concrete panels rely primarily on manual sampling, with workers using tools like tape measures and steel rulers to check component dimensions against design specifications. This method is not only time-consuming and labor-intensive, but also prone to errors, omissions, and misdetection, making it difficult to meet the demands of large-scale industrial production.

[0004] In recent years, although researchers have begun to explore automated DQI methods, these methods are difficult to promote and apply in actual projects due to the limitations of the inspection methods and equipment themselves.

[0005] Photogrammetry: Early researchers attempted to use photogrammetry to measure the quality of target dimensions in engineering applications. However, this method requires a certain amount of control point data, which typically requires additional effort and time to obtain, increasing the complexity and cost of measurement. Furthermore, poor lighting conditions can significantly affect photogrammetry results.

[0006] 3D laser scanning technology: Due to its high precision and high-speed data acquisition capabilities, 3D laser scanning has become a mainstream technology for prefabricated component dimensional inspection. However, this technology is subject to high equipment costs, complex data processing, and the need for specialized technicians and specialized testing sites, which limits its widespread adoption in practical applications.

[0007] Image Processing Technology (IPT): While IPT methods can process captured images to detect and analyze part dimensional variations, their robustness can be significantly degraded when noise (such as distortion and lighting) significantly impacts image content. Furthermore, IPT methods are difficult to apply to large quantities of continuously produced parts.

[0008] With the rise of computer technology and big data, deep learning (DL) is increasingly being used in civil engineering, providing a key technology for addressing challenges such as concrete appearance quality inspection, steel strip quality inspection, and sewage pipe quality inspection. For prefabricated component dimensional quality inspection, deep learning can build accurate recognition and segmentation models by learning component features, enabling automated and efficient inspection. However, current research on deep learning models primarily serves as auxiliary tools for post-processing point cloud models using 3D laser scanning technology, limiting their scope of application. Summary of the Invention

[0009] The purpose of the present invention is to propose a non-contact online dimensional quality detection method and system for precast concrete slabs based on multimodal fusion, which can detect the dimensional quality of precast components in real time during the production process and improve the detection efficiency and accuracy.

[0010] To achieve the above objectives, the present invention provides a non-contact online dimensional quality detection method for precast concrete panels based on multimodal fusion, comprising: Capture images of precast concrete panels and preprocess the captured original images; Build a model, select the YOLOv8m model as the baseline model, build a pyramid network that enhances context feature fusion, and introduce shape-IOU to identify and classify objects in the image and generate ROI areas; HQ-SAM is used as the segmentation model to perform pixel-level segmentation, obtain the pixel-level segmentation mask of the laminate and the hole, and adjust the coordinates of the ROI region bounding box according to the segmentation mask; Measure the dimensions of the laminate, hole, and electrical box prediction frame respectively to obtain their respective dimension data; Automatically extract dimensional design data from the BIM model, use the visual inspection results as component production data, and compare them with the design information to determine whether the component's dimensional quality is qualified.

[0011] Beneficial effects of the basic solution: This solution eliminates the need for manual contact with precast concrete slabs, enabling real-time inspection through direct image acquisition. This avoids the risk of component damage associated with traditional contact measurement. It also seamlessly integrates with production lines, automating the entire process from production to quality inspection, significantly improving inspection efficiency and meeting the quality inspection requirements of industrialized mass production. By combining image acquisition with BIM model data, production data is captured through visual inspection and automatically compared with design information, eliminating the tedious steps of manual data entry and comparison.

[0012] Traditional measurement methods require different measurement tools and methods for different component features such as laminates, holes, and electrical boxes, which is cumbersome and time-consuming. However, this technical solution can simultaneously measure the dimensions of multiple component features such as laminates, holes, and electrical boxes, covering the key geometric parameters of precast concrete panels. There is no need to frequently change measurement tools or adjust measurement processes, which greatly saves measurement time and improves overall work efficiency. This technical solution improves upon the YOLOv8m baseline model, constructing a pyramid network that enhances contextual feature fusion and introduces shape-IOU. The result is a specialized model specifically designed for multi-target recognition and classification of prefabricated laminates. This model effectively captures detailed features such as edges and holes in prefabricated panels. Combined with the pixel-level segmentation capabilities of HQ-SAM, this achieves dual security from target localization to refined boundary segmentation. Adjusting the ROI bounding box coordinates using a segmentation mask avoids measurement errors caused by blurred object edges. Design data is automatically extracted from the BIM model, enabling direct comparison and verification of inspection results against design intent, completing quality inspections.

[0013] As an implementable optimal solution, an industrial linear array camera is selected as the image acquisition device. The image acquisition is based on the workbench. The image acquisition device is set at the exit of the maintenance kiln and laser triggering is used to achieve automatic acquisition. When the workbench has not passed, the laser emitted by the laser transmitter is continuously received by the receiver and the camera is in a non-working state; when the workbench passes, the front buffer pad blocks the laser and the receiver cannot receive the signal, triggering the camera to start continuous scanning; after the workbench has completely passed, the receiver receives the laser signal again and the camera stops working.

[0014] As an implementable preferred solution, image acquisition is performed on the precast concrete slab, and the acquired original image is preprocessed, including the following: The shooting equipment is set at the exit of the maintenance kiln and the laser trigger method is used to realize automatic data collection; Use the LabelImg tool to label the target area and generate a true bounding box; Perform data enhancement processing on the original image, including rotation, flipping, scaling, and brightness adjustment; The original images are randomly divided into training set, validation set and test set.

[0015] As an implementable preferred solution, the YOLOv8m model is selected as the baseline model, including the following: The head network of the YOLOv8m model adopts a decoupled head structure, which separates the detection head from the classification structure and replaces the anchor-based structure with an anchor-free structure; the loss function uses CIoU and distributed focal loss for loss calculation.

[0016] As an implementable and optimal solution, a pyramid network for enhancing context feature fusion is constructed, which includes the following contents: By constructing multi-directional channels, the features in the feature extraction network are fused with the features in each sub-path to achieve cross-scale connection and ensure that the features at each scale have detailed contextual information; the pyramid network also adds two-dimensional perception selection integration, and adjusts the fusion method and degree between features by considering the scale and dimensional information of the features.

[0017] As an implementable preferred solution, two-dimensional perception selection integration is used to transform high-dimensional features through convolution and interpolation. and low-dimensional features With the characteristics of the current layer Align them; then divide them into four equal parts in the channel dimension, and get 、 and ,in, 、 and Represent the first segmentation features of the low-level, high-level, and current-level features, respectively, and calculate the partition according to the following formula:

[0018]

[0019]

[0020]

[0021] in, The activation function is applied to The results obtained later, is the selective aggregation result of each partition; merged on the channel dimension To obtain ; Operations include 、 and ;when When , the model emphasizes fine-grained features, and when , giving priority to contextual features.

[0022] As an implementable and preferred solution, Shape-IoU is used to calculate the loss, taking into account the geometric constraints between the real box and the predicted box, and balancing the focus on the shape and scale of the bounding box. The Shape-IoU calculation method is as follows:

[0023]

[0024]

[0025]

[0026]

[0027]

[0028]

[0029] in, The scale factor is related to the scale of the objects in the dataset; is the vertical weighting coefficient, is the horizontal weighting coefficient, whose specific value depends on the shape of the true label box; the regression loss is calculated as follows:

[0030] The training process uses the SGD optimizer.

[0031] As an implementable and optimal solution, the bounding box output by the improved baseline model is input into the HQ-SAM model as the ROI region to obtain the segmentation mask of the laminate and the hole. The performance evaluation indicators of the segmentation model include pixel accuracy PA, DICE coefficient and mean intersection over union (mIoU), as shown in the following formula:

[0032]

[0033]

[0034] in, Indicates the number of pixel categories in the labeled image; The prediction is , and the actual Number of pixels; The prediction is , and the actual The number of pixels; The prediction is , but actually The number of pixels; is the DICE coefficient; The bounding box coordinates are adjusted according to the segmentation mask to eliminate the positioning deviation caused by occlusion.

[0035] As an implementable preferred solution, the dimensions of the laminate, hole, and predicted frame of the electrical box are measured respectively to obtain their respective dimensional data, including the following: Dimensional measurements of sandwich panels, corresponding to the longitudinal coordinate laminate width The formula is:

[0036] The formula for the average width of a laminate is:

[0037] in, is the number of vertical coordinate points, is the total sampling points, and Respectively indicate that in the nth row, the vertical coordinate value is When , the horizontal coordinate values ​​of the left and right edges of the laminate; is the pixel value of the average width of the laminate; Measure the center coordinates and diameter of holes, including: Use Graham scan method to calculate the convex hull on the Hole binary image and find the smallest convex polygon containing all the points. Use small edges to split the convex hull, approximate each small segment, and connect all the approximated segments; Traverse all edges of the approximate polygon and find the smallest outer rectangle; The calculation formula for the center coordinates of the hole is as follows:

[0038] The diameter is calculated from the equation:

[0039] in, represents the pixel coordinates of the hole center, represents the pixel diameter of the hole; For electrical box coordinate measurement, the local box information generated by the improved benchmark model is directly used as the coordinates of the predicted box of the electrical box. The calculation formula is as follows:

[0040] in, is the pixel coordinate of the center point of the electrical box prediction box, and are the minimum horizontal and vertical coordinate values ​​of the prediction box, respectively.

[0041] As an implementable and optimal solution, a Revit plug-in is developed to implement the IExternalCommand interface, and a C# class library is created through the compiler to define the initial project settings and classes to implement the function of extracting PLS size information. Data comparison and evaluation include: Laminate outline dimension error, the formula is as follows:

[0042] Hole diameter error, the formula is as follows:

[0043] Coordinate error, the formula is as follows:

[0044] Among them, and and are the design values ​​of length and width, and are the differences between the test values ​​and the design values ​​in the length and width directions respectively; is the difference between the detected hole diameter and the actual value; is the difference between the detected horizontal coordinate and the actual horizontal coordinate, is the difference between the detected vertical coordinate and the actual vertical coordinate. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A schematic flow chart of a non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion provided in one embodiment of the present invention.

[0046] Figure 2 Schematic diagram of the CEFFPN structure.

[0047] Figure 3 Schematic diagram of the dual-stream FPN structure.

[0048] Figure 4 Schematic diagram of DASI structure.

[0049] Figure 5 Schematic diagram of the improved network structure of the PLS recognition model.

[0050] Figure 6 Schematic diagram of the calculation relationship of Shape-IoU.

[0051] Figure 7 An example diagram of the recognition performance of the benchmark model on PLS.

[0052] Figure 8 Schematic diagram of PLS ​​width calculation.

[0053] Figure 9 Schematic diagram for hole center coordinates and hole diameter calculation.

[0054] Figure 10 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.

[0055] Reference numerals: electronic device 500 , processor 501 , communication interface 502 , memory 503 , bus 504 . DETAILED DESCRIPTION

[0056] In order to make the technical solution and advantages of the present application clearer, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It will be understood that the specific embodiments described herein are only partial embodiments of the present invention, which are only used to explain the present application, rather than to limit the present application. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered to be isolated, and they can be combined with each other to achieve better technical effects. The same reference numerals appearing in the drawings of the following embodiments represent the same features or components, which can be applied to different embodiments.

[0057] In addition, unless otherwise defined, technical or scientific terms used in the description of the present invention should have the common meanings understood by those skilled in the art in the art to which the present invention belongs.

[0058] The present invention will be further described in detail below with reference to the accompanying drawings: Figure 1 ,A non-contact online size quality detection method for precast concrete slabs based on multi-modal fusion ,includes the following steps.

[0059] Step S100, data collection and preprocessing, includes: In step S101, an industrial linear array camera is selected as the main image acquisition device. The linear array camera has a single row of photosensitive cells and can achieve wide-range measurement while maintaining high measurement accuracy, making it suitable for continuous scanning applications.

[0060] Image acquisition is based on the workbench, which in this example measures 12,000 mm x 3,500 mm. The workbench runs at a constant speed of 13 m / min, and the maximum field of view perpendicular to the workbench's direction of travel is 3,500 mm. To allow for a certain margin, the field of view perpendicular to the workbench's direction of travel is set to 4,000 mm. The PLS's lateral detection accuracy is set to 0.5 mm per pixel.

[0061] The lateral resolution of the selected camera is calculated as follows:

[0062] in, Indicates the width of the field of view, represents the required detection accuracy, Indicates the minimum pixel value required for shooting. In this embodiment, the minimum pixel of the camera should be no less than 8000 pixels.

[0063] The line frequency is the scanning rate of the linear array camera. By controlling the line frequency, it can be ensured that the obtained PLS image is neither stretched nor compressed. The formula for calculating the line frequency is as follows:

[0064] in is the line frequency, The speed of the workbench.

[0065] The larger the focal length of the lens, the larger the field of view and the corresponding shooting field of view. and field of view The calculation formula is as follows:

[0066]

[0067] in, Indicates the focus distance, is the size of the target surface, is the field of view width.

[0068] The image acquisition device is located at the exit of the maintenance kiln and uses laser triggering for automatic acquisition. When the workbench is not passing, the laser emitted by the laser transmitter is continuously received by the receiver, and the camera is inactive. When the workbench passes, the front buffer blocks the laser, preventing the receiver from receiving the signal, triggering the camera to begin continuous scanning. After the workbench has completely passed, the receiver re-receives the laser signal, and the camera stops operating.

[0069] In step S102, the captured raw image is preprocessed and the target regions are labeled using the LabelImg tool to generate ground truth bounding boxes. This is a necessary step for supervised learning. The labeled targets include laminates, holes, and electrical boxes. In this example, a dataset containing 1592 laminates, 725 holes, and 1036 electrical boxes was constructed based on the captured images.

[0070] In order to improve the generalization ability of the model, data augmentation processing is performed on the original images, including rotation, flipping, scaling, brightness adjustment and other operations to expand the scale and diversity of the dataset.

[0071] Step S103 : The original images are randomly distributed into a training set, a validation set, and a test set. In this embodiment, the proportions are 70%, 15%, and 15%, respectively.

[0072] Step S200, multi-target recognition and classification, generating a region of interest (ROI), includes: Step S201: Establish a model and select the YOLOv8m model as the baseline model: including the backbone network, neck network and head network.

[0073] The backbone network adopts the idea of ​​CSP and replaces the C3 module of YOLOv5 with the efficient layer aggregation network (ELAN) and C2f structure, so that YOLOv8m can obtain richer gradient flow information while maintaining its own lightweight.

[0074] The neck network adopts a path aggregation network (PAN) and feature pyramid network (FPN) similar to YOLOv5, removes two convolutional connection layers, and is designed into a lightweight Dual Stream FPN. Its feature fusion efficiency and speed are better than PAN-FPN.

[0075] The head network adopts a decoupled head structure, which separates the detection head from the classification structure and replaces the anchor-based structure with an anchor-free structure. The loss function uses CIoU and distributed focal loss (DFL) for loss calculation. The calculation method of CIoU is as follows:

[0076]

[0077]

[0078] Among them, and respectively express The width and height of the real frame, and represents the predicted frame Width and height The variables in the formula represent penalty terms, which reflect the Predicted frame The difference in aspect ratio from the real frame, This is the penalty The weight value of .

[0079] The calculation method of DFL is as follows:

[0080] in, is the sigmoid output of the network, and is the interval order, It's a label.

[0081] This embodiment uses frames per second (FPS), FLOPs, and mean average precision (mAP) as detection performance evaluation indicators. FLOPs is used to measure the computational complexity of the model, indicating the number of floating-point operations required for the model to make predictions, which helps to understand the computing resources required by the model. mAP is a commonly used object detection metric that measures the average accuracy of a model in different categories or classes. mAP0.5:0.95 represents the average mAP value when the intersection-over-union (IOU) is between 0.5 and 0.95 with a step size of 0.05. The formula is as follows:

[0082]

[0083]

[0084]

[0085]

[0086] Among them, R is the recall rate, P is the precision rate, TP, FP, and FN are the number of positive cases, negative cases, and positive cases judged by the model, respectively.

[0087] Step S202: construct a pyramid network for enhanced context feature fusion (CEFFPN), the structure of which is as follows: Figure 2 CEFFPN constructs multi-directional channels to fuse features from the feature extraction network with features from each subpath, achieving cross-scale connections and ensuring that features at each scale have detailed contextual information, thereby promoting higher-level feature fusion.

[0088] Compared with dual-stream FPN (refer to Figure 3 ), CEFFPN not only redesigns the feature transmission path, but also adds two Dimensional Aware Selection Integration (DASI).

[0089] The channel partitioning selection mechanism enables DASI to adaptively select appropriate features for fusion based on the object's size and characteristics. This is crucial for detecting the predicted bounding box of an electrical box. The predicted bounding box occupies a small pixel ratio in the image, carries limited information, and is often occluded during the concrete pouring process, making it difficult to identify.

[0090] By considering the scale and dimensionality of features, DASI can precisely adjust the fusion method and degree between features, optimize the utilization efficiency of features, and improve the model's ability to extract and represent each target feature in PLS, which is especially beneficial for predicting electrical boxes that carry less information.

[0091] Reference Figure 4 DASI transforms high-dimensional features into and low-dimensional features With the characteristics of the current layer Then, they are divided into four equal parts in the channel dimension and obtained 、 and ,in, 、 and Represents the first segmentation feature of the low-level, high-level and current-level features respectively. The partition is calculated according to the following formula,

[0092]

[0093]

[0094]

[0095] in, The activation function is applied to The results obtained later, Is the selective aggregation result of each partition. Merge on the channel dimension To obtain The operation includes 、 and ,when When , the model emphasizes fine-grained features; when , giving priority to contextual features.

[0096] Improved baseline model network structure reference Figure 5 .

[0097] Step S203, model training, includes: The training parameter configuration is shown in Table 1.

[0098] Table 1. Hyperparameter settings.

[0099]

[0100] YOLOv8m uses CIoU and DFL to calculate the bounding box regression loss, but CIoU does not consider the balance between hard and easy examples in the loss calculation. In addition, CIoU uses aspect ratio as a penalty term in the loss function, which does not accurately reflect the actual difference between the predicted box and the ground truth (GT) box when they have the same aspect ratio but different width and height.

[0101] Therefore, in this embodiment, Shape-IoU is used instead of CIoU to calculate the loss. By considering the geometric constraints between the real box and the predicted box and balancing the attention to the shape and scale of the bounding box, the accuracy of bounding box regression is improved. Figure 6 As shown, the calculation method is as follows:

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108] in, The scale factor is related to the scale of the objects in the dataset; is the vertical weighting coefficient, Is the horizontal weighting coefficient, whose specific value depends on the shape of the true label (GT) box. The regression loss is calculated as follows:

[0109] The training process uses the SGD optimizer.

[0110] Step S204, refer to Figure 7 The trained model is used as a baseline model to generate predicted bounding boxes. The PLS image is input, and the bounding boxes of the laminate, hole, and electrical box output by the model serve as the input ROI (region of interest) of the subsequent segmentation model. By directing the attention of the subsequent model to the ROI in the PLS image, it helps to achieve accurate segmentation of the laminate, hole, and electrical box.

[0111] Step S300, segmentation model optimization and bounding box refinement based on SAM, includes: In step S301, HQ-SAM is used as the segmentation model. SAM consists of three modules: image encoder, cue encoder, and mask decoder. It can accurately predict the segmentation mask. Its core improvements include: Learn high-quality output tokens: Improve segmentation accuracy without significantly increasing the number of parameters.

[0112] Multi-scale feature adaptation: Achieve high-precision segmentation of small objects (holes, electrical boxes) through lightweight design.

[0113] In step S302, the bounding box output by the improved baseline model is input as the region of interest (ROI) into the HQ-SAM model to obtain a more accurate segmentation mask for the laminate and hole. Because concrete often obscures the electrical box during the PLS pouring and vibration process, making it difficult to generate an accurate segmentation mask in subsequent stages, this step only considers the pixel area of ​​the laminate and hole for segmentation.

[0114] In the pixel-level segmentation process of concrete slabs and holes, since the segmentation mainly generates a mask image, the evaluation indicators usually rely on the comparative analysis between the mask image and the true pixel-level label. The performance evaluation indicators of the segmentation model include pixel accuracy (PA), (DICE coefficient) and mean intersection over union (mIoU), the formula is as follows:

[0115]

[0116]

[0117] in, Indicates the number of pixel categories in the labeled image; The prediction is , and the actual Number of pixels; The prediction is , and the actual The number of pixels; The prediction is , but actually The number of pixels. is defined as the predicted segmentation ( ) and the true segmentation ( ), divided by and size.

[0118] In step S303, the ROI image is input, and the segmentation model outputs a pixel-level segmentation mask for the laminate and the hole. Since the electrical box is often occluded, only the bounding box information generated by the baseline model is used. The mask is a binary image with foreground pixel values ​​of 1 and background pixels of 0.

[0119] The bounding box coordinates are adjusted according to the segmentation mask to eliminate the positioning deviation caused by occlusion.

[0120] Step S400 involves multi-target dimensional measurement. Because the laminate, hole, and electrical box each have distinct structural characteristics and functional requirements, their dimensional inspection items differ significantly. Specifically, the laminate focuses on measuring outline dimensions, the hole focuses on measuring center coordinates and diameter, and the electrical box's prediction box focuses on measuring coordinate points. Therefore, different methods are required for dimensional inspection of these structural components, including: Step S401, measure the dimensions of the sandwich panel, refer to Figure 8 , corresponding to the vertical coordinate laminate width The formula is:

[0121] The formula for the average width of a laminate is:

[0122] in, is the number of vertical coordinate points, is the total sampling points, and Respectively indicate that in the nth row, the vertical coordinate value is , the horizontal coordinate values ​​of the left and right edges of the laminate. is the average width of the laminate in pixels, and similarly, the average length of the laminate in pixels Can also be obtained.

[0123] Step 402, measure the center coordinates and diameter of the hole, refer to Figure 9 ,include: Use Graham scan method to calculate the convex hull of the Hole binary image and find the smallest convex polygon containing all the points.

[0124] Split the convex hull with small edges, approximate each small segment, and connect all the approximated segments to approximate the original convex hull.

[0125] Traverse all the edges of the approximate polygon and find the two farthest vertices. These two vertices define one side of the rectangle. Next, find the two vertices perpendicular to this side and farthest away. These vertices will determine the other side of the rectangle, and finally form the smallest outer rectangle.

[0126] The calculation formula for the center coordinates of the hole is as follows:

[0127] The diameter is calculated from the equation:

[0128] in, represents the pixel coordinates of the hole center, Indicates the diameter of the hole in pixels.

[0129] Step S403, measuring the electrical box coordinates, includes: During the pouring process, the concrete float often obscures the electrical box. Since its outline is not clear, it is difficult to generate an accurate segmentation mask. Therefore, the local box information generated by the improved baseline model is directly used as the coordinates of the predicted box of the electrical box. The calculation formula is as follows:

[0130] in, is the pixel coordinate of the center point of the electrical box prediction box, and are the minimum horizontal and vertical coordinate values ​​of the prediction box, respectively.

[0131] Step S500 automatically extracts dimensional design data from the BIM model, uses the visual inspection results as component production data, and compares them with the design information to determine whether the component's dimensional quality is qualified, including: Step S501: Develop a Revit plug-in and implement the IExternalCommand interface. The development process is as follows: Create a new C# class library in the Visual Studio compiler and run it through the .NET Framework.

[0132] Define the initial settings for a new project, including data storage location, project name, and category.

[0133] Define a class that implements the IExternalCommand interface to extract PLS size information. The core of the IExternalCommand interface is the Execute function, which contains three internal parameters: CommandData, Message, and Elements. CommandData is the function's input parameter, used to store PLS size information; Message is an output parameter that returns the program's running status as a string; and Elements is an output parameter used to highlight graphical elements in the event of a program failure. After overriding the Execute function, a collection of PLS ​​information classes is created, as shown in Table 3.

[0134] Table 2. Program parameters corresponding to PLS.

[0135]

[0136] The code iterates over the elements of the PLS model, filters out the laminates, holes, and electrical boxes, gets the ID and size information of these elements, and adds the information of each element to the PLS information class collection.

[0137] Step S502, data comparison and evaluation, includes: Laminate outline dimension error, the formula is as follows:

[0138] Hole diameter error, the formula is as follows:

[0139] Coordinate error, the formula is as follows:

[0140] Among them, and and are the design values ​​of length and width, and are the differences between the detected values ​​and the designed values ​​in the length and width directions respectively. is the difference between the detected hole diameter and the actual value. is the difference between the detected horizontal coordinate and the actual horizontal coordinate, is the difference between the detected vertical coordinate and the actual vertical coordinate.

[0141] This embodiment also provides a non-contact online dimensional quality detection system for precast concrete slabs based on multimodal fusion, which is capable of executing the above-mentioned non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion.

[0142] This embodiment of the present application also provides an electronic device 500 that utilizes the aforementioned non-contact online dimensional quality inspection method for precast concrete panels based on multimodal fusion. The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the aforementioned non-contact online dimensional quality inspection method for precast concrete panels based on multimodal fusion are implemented. In this embodiment of the present application, the processor serves as the control center of the computer method and can be a processor of a physical machine or a processor of a virtual machine.

[0143] Reference Figure 10The electronic device 500 includes at least one processor 501, at least one communication interface 502, at least one memory 503, and at least one bus 504. Bus 504 is used to facilitate communication between these components, communication interface 502 is used to communicate signaling or data with other node devices, and memory 503 stores machine-readable instructions executable by processor 501. When the electronic device 500 is in operation, processor 501 communicates with memory 503 via bus 504. When the machine-readable instructions are invoked by processor 501, the steps of the non-contact online dimensional quality inspection method for precast concrete panels based on multimodal fusion are executed.

[0144] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor of an electronic device, it can implement the steps of the non-contact online size quality detection method of precast concrete panels based on multimodal fusion as described above.

[0145] Those skilled in the art will understand that all or part of the process steps in the method for non-contact online dimensional quality detection of precast concrete panels based on multimodal fusion can be implemented by instructing related hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of various embodiments of the method for non-contact online dimensional quality detection of precast concrete panels based on multimodal fusion. Among them, any reference to memory, storage, database or other media used in the various embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0146] The above contents are merely embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme is not described in detail here. A person of ordinary skill in the art is aware of all common technical knowledge in the technical field to which the invention belongs before the filing date or priority date, is able to obtain all existing technologies in the field, and has the ability to apply conventional experimental means before that date. A person of ordinary skill in the art can, under the guidance of this application, improve and implement this scheme in combination with his or her own abilities. Some typical known structures or known methods should not become an obstacle for a person of ordinary skill in the art to implement this application. It should be pointed out that for a person of ordinary skill in the art, several variations and improvements can be made without departing from the structure of the present invention, which should also be regarded as the scope of protection of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection claimed in this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A non-contact online dimensional quality detection method for precast concrete panels based on multimodal fusion, characterized in that: include: Capture images of precast concrete panels and preprocess the captured original images; Build a model, select the YOLOv8m model as the baseline model, build a pyramid network that enhances context feature fusion, and introduce shape-IOU to identify and classify objects in the image and generate ROI areas; HQ-SAM is used as the segmentation model to perform pixel-level segmentation, obtain the pixel-level segmentation mask of the laminate and the hole, and adjust the coordinates of the ROI region bounding box according to the segmentation mask; Measure the dimensions of the laminate, hole, and electrical box prediction frame respectively to obtain their respective dimension data; Automatically extract dimensional design data from the BIM model, use the visual inspection results as component production data, and compare them with the design information to determine whether the component's dimensional quality is qualified.

2. The non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion according to claim 1 is characterized in that: An industrial linear array camera is used as the image acquisition device. The image acquisition is based on the workbench and is set at the exit of the maintenance kiln. Laser triggering is used to achieve automatic acquisition. When the workbench has not passed, the laser emitted by the laser transmitter is continuously received by the receiver and the camera is in a non-working state; when the workbench passes, the front buffer pad blocks the laser and the receiver cannot receive the signal, triggering the camera to start continuous scanning; after the workbench has completely passed, the receiver receives the laser signal again and the camera stops working.

3. The non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion according to claim 2 is characterized in that: The precast concrete slabs are imaged and preprocessed, including the following: The shooting equipment is set at the exit of the maintenance kiln and the laser trigger method is used to realize automatic data collection; Use the LabelImg tool to label the target area and generate a true bounding box; Perform data enhancement processing on the original image, including rotation, flipping, scaling, and brightness adjustment; The original images are randomly divided into training set, validation set and test set.

4. The non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion according to claim 1 is characterized in that: The YOLOv8m model is selected as the baseline model, including the following: The head network of the YOLOv8m model adopts a decoupled head structure, which separates the detection head from the classification structure and replaces the anchor-based structure with an anchor-free structure; the loss function uses CIoU and distributed focal loss for loss calculation.

5. The non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion according to claim 4 is characterized in that: Constructing a pyramid network for enhanced context feature fusion, including the following: By constructing multi-directional channels, the features in the feature extraction network are fused with the features in each sub-path to achieve cross-scale connection and ensure that the features at each scale have detailed contextual information; the pyramid network also adds two-dimensional perception selection integration, and adjusts the fusion method and degree between features by considering the scale and dimensional information of the features.

6. The non-contact online dimensional quality detection method for precast concrete panels based on multimodal fusion according to claim 5 is characterized in that: Two-dimensional perception selection integration through convolution and interpolation, high-dimensional features and low-dimensional features With the characteristics of the current layer Align them; then divide them into four equal parts in the channel dimension, and get 、 and ,in, 、 and Represent the first segmentation features of the low-level, high-level, and current-level features, respectively, and calculate the partition according to the following formula: in, The activation function is applied to The results obtained later, is the selective aggregation result of each partition; merged on the channel dimension To obtain ; Operations include 、 and ;when When , the model emphasizes fine-grained features, and when , giving priority to contextual features.

7. The non-contact online dimensional quality detection method for precast concrete panels based on multimodal fusion according to claim 6 is characterized in that: The loss is calculated by Shape-IoU, which considers the geometric constraints between the real box and the predicted box and balances the attention to the shape and scale of the bounding box. The Shape-IoU calculation method is as follows: in, The scale factor is related to the scale of the objects in the dataset; is the vertical weighting coefficient, is the horizontal weighting coefficient, and its specific value depends on the shape of the true label box; the regression loss is calculated as follows: The training process uses the SGD optimizer.

8. The non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion according to claim 2 is characterized in that: The bounding box output by the improved baseline model is input into the HQ-SAM model as the ROI region to obtain the segmentation mask of the laminate and the hole. The performance evaluation indicators of the segmentation model include pixel accuracy PA, DICE coefficient and mean intersection over union (mIoU), as shown in the following formula: in, Indicates the number of pixel categories in the labeled image; The prediction is , and the actual Number of pixels; The prediction is , and the actual The number of pixels; The prediction is , but actually The number of pixels; is the DICE coefficient; The bounding box coordinates are adjusted according to the segmentation mask to eliminate the positioning deviation caused by occlusion.

9. The non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion according to claim 1 is characterized in that: The laminate, hole, and electrical box prediction boxes were measured to obtain their respective dimensional data, including the following: Dimensional measurements of sandwich panels, corresponding to the longitudinal coordinate laminate width The formula is: The formula for the average width of a laminate is: in, is the number of vertical coordinate points, is the total sampling points, and Respectively indicate that in the nth row, the vertical coordinate value is When , the horizontal coordinate values ​​of the left and right edges of the laminate; is the pixel value of the average width of the laminate; Measure the center coordinates and diameter of holes, including: Use Graham scan method to calculate the convex hull on the Hole binary image and find the smallest convex polygon containing all the points. Use small edges to split the convex hull, approximate each small segment, and connect all the approximated segments; Traverse all edges of the approximate polygon and find the smallest outer rectangle; The calculation formula for the center coordinates of the hole is as follows: The diameter is calculated from the equation: in, represents the pixel coordinates of the hole center, represents the pixel diameter of the hole; For electrical box coordinate measurement, the local box information generated by the improved benchmark model is directly used as the coordinates of the predicted box of the electrical box. The calculation formula is as follows: in, is the pixel coordinate of the center point of the electrical box prediction box, and are the minimum horizontal and vertical coordinate values ​​of the prediction box, respectively.

10. The non-contact online dimensional quality detection method for precast concrete slabs based on multimodal fusion according to claim 1, characterized in that: Automatically extract dimensional design data from the BIM model, use visual inspection results as component production data, and compare them with the design information to determine whether the component's dimensional quality is qualified, including the following: Developed a Revit plug-in, implemented the IExternalCommand interface, created a C# class library through the compiler, defined the project initial settings and classes, and implemented the function of extracting PLS size information; Data comparison and evaluation include: Laminate outline dimension error, the formula is as follows: Hole diameter error, the formula is as follows: Coordinate error, the formula is as follows: in, and are the design values ​​of length and width, and are the differences between the test values ​​and the design values ​​in the length and width directions respectively; is the difference between the detected hole diameter and the actual value; is the difference between the detected horizontal coordinate and the actual horizontal coordinate, is the difference between the detected vertical coordinate and the actual vertical coordinate.