AI-based preform multi-modal visual quality intelligent detection system and method
By using multimodal visual data acquisition and deep learning models to detect defects in prefabricated components, the problems of low efficiency and misjudgment in existing technologies are solved, and high-precision defect detection and quality assessment are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-03-24
AI Technical Summary
Existing methods for detecting defects in precast components are inefficient and lack sufficient intelligence, making it difficult to meet the quality requirements of high-precision components and prone to misjudgment.
A multimodal visual data acquisition unit is used to acquire multispectral imaging and three-dimensional structural scanning data of prefabricated components. Defect candidate areas are extracted through a multimodal feature fusion network, and a deep learning detection and segmentation model is used to classify defect types and segment boundaries. Quantitative indicators are calculated by combining three-dimensional point cloud and infrared images, and the defect judgment threshold is dynamically adjusted.
It significantly improves the intelligence and accuracy of defect detection in precast components, reduces misjudgments, provides accurate quantitative information on defects and quality assessment, and adapts to different environments and process differences.
Smart Images

Figure CN120890989B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent factory and workpiece quality inspection, and particularly relates to a prefabricated part multi-modal visual quality intelligent detection system and method based on AI. BACKGROUND
[0002] With the increasing demand for product quality, it is often necessary to detect defects in prefabricated components. Especially when the prefabricated component is a high-precision component, the production process may cause unevenness, edge collapse, scratches, color spots, cracks and other defects on the surface of the component, affecting its appearance, and even affecting its performance when the defects are serious.
[0003] At present, the existing prefabricated component defect detection methods have the following problems. Some methods use manual identification by naked eyes. The detection personnel identifies the component at multiple angles by turning it over under strong light, which has low detection efficiency and detection rate, and high labor intensity, and is easy to cause damage to the eyes of the detection personnel. Some methods use automatic defect detection equipment. However, the existing automatic detection equipment has complex structure, and it is difficult to detect the workpiece from multiple directions. The existing automatic detection equipment lacks the division of workpiece defect grades, and is prone to misjudgment due to environmental and process differences. The robustness of the component defect detection needs to be further improved. At the same time, the existing workpiece defect detection technology needs to be further improved in terms of intelligence. Ultimately, the workpiece detection rate is low, and it is difficult to meet the yield requirement of products. SUMMARY
[0004] To solve the problems in the prior art, the present application aims to solve the above-mentioned defects, and further provides a prefabricated part multi-modal visual quality intelligent detection system and method based on AI.
[0005] The present application adopts the following technical solutions.
[0006] The present application discloses a prefabricated part multi-modal visual quality intelligent detection system based on AI, which comprises:
[0007] A multi-modal visual data acquisition unit is configured to acquire multi-modal visual data of a detection area of a prefabricated component, and calibrate the multi-modal visual data to construct a multi-modal data set. The multi-view acquisition component is composed of a multi-spectral imaging device and a three-dimensional structure scanner.
[0008] A feature fusion and candidate region extraction unit is configured to input the multi-modal data set into a multi-modal feature fusion network to extract multi-modal features, and fuse the multi-modal features in a feature space to generate a defect candidate region of the prefabricated component. The multi-modal features include surface texture features extracted from an RGB image, temperature difference distribution features extracted from an infrared image, crack penetration fluorescence features extracted from an ultraviolet image, and geometric deformation and size features extracted from point cloud data.
[0009] a defect type recognition and quantitative analysis unit configured to invoke a deep learning detection and segmentation model to classify defect types and segment defect boundaries of the defect candidate region, and calculate quantitative indexes of defect regions within the segmented boundaries using three-dimensional point clouds and infrared images;
[0010] a defect level discrimination and component quality evaluation unit configured to dynamically determine thresholds of the defect regions according to defect types and quantitative indexes of the defect regions, to determine defect severity levels of the defect regions, and output quality evaluation results of the prefabricated components.
[0011] The second aspect of the present application discloses an AI-based prefabricated component multi-modal visual quality intelligent detection method, which is implemented by the AI-based prefabricated component multi-modal visual quality intelligent detection system of the first aspect. The method comprises:
[0012] acquiring multi-modal visual data of the prefabricated component detection region through multi-view acquisition components, and calibrating the multi-modal visual data to construct a multi-modal data set of a unified three-dimensional coordinate system;
[0013] invoking a multi-modal feature fusion network to extract multi-modal features from the multi-modal data set, and fusing the multi-modal features in a feature space to generate defect candidate regions of the prefabricated component;
[0014] utilizing a deep learning detection and segmentation model to classify defect types and segment defect boundaries of the defect candidate region, and calculating quantitative indexes of defect regions within the segmented boundaries using three-dimensional point clouds and infrared images;
[0015] dynamically determining thresholds of the defect regions according to defect types and quantitative indexes of the defect regions, to determine defect severity levels of the defect regions, and output quality evaluation results of the prefabricated components.
[0016] The multi-modal features include surface texture features extracted from RGB images, temperature difference distribution features extracted from infrared images, crack penetration fluorescence features extracted from ultraviolet images, and geometric deformation and size features extracted from point cloud data.
[0017] Further, the multi-view acquisition components are installed at the end of a mechanical arm and are composed of a multi-spectral imaging device and a three-dimensional structure scanner. The multi-spectral imaging device is used to acquire RGB images, infrared images, and ultraviolet images of the prefabricated component detection region through an RGB camera, an infrared thermal imager, and an ultraviolet fluorescence camera, respectively. The three-dimensional structure scanner is used to acquire point cloud data of the prefabricated component detection region.
[0018] The multi-modal visual data of the prefabricated component detection area is acquired by the multi-view acquisition component, and the multi-modal visual data is calibrated to construct a multi-modal data set of a unified three-dimensional coordinate system, comprising:
[0019] Based on the RGB image, infrared image, ultraviolet image, point cloud data of the prefabricated component detection area and the data acquisition timestamp, combined with the shape size of the prefabricated component and the kinematics parameters of the mechanical arm, a multi-modal original data set is constructed;
[0020] The same clock source is used for time synchronization processing of the multi-modal original data set, and the RGB image, infrared image and ultraviolet image are converted to a unified three-dimensional coordinate system with the point cloud data by Zhang's calibration method to construct the multi-modal data set.
[0021] Further, the multi-modal feature fusion network is called to extract multi-modal features from the multi-modal data set, and the multi-modal features are fused in the feature space to generate a defect candidate area of the prefabricated component, comprising:
[0022] The surface texture features of the prefabricated component are obtained by extracting features from the RGB image through a convolutional neural network;
[0023] Based on the infrared image and the ultraviolet image, the infrared temperature difference and the ultraviolet fluorescence are calculated respectively to obtain the temperature difference distribution features and the crack penetration fluorescence features;
[0024] Based on the point cloud data, the geometric deviation of the prefabricated component is calculated to obtain the geometric deformation and size features;
[0025] The surface texture features, temperature difference distribution features, crack penetration fluorescence features and geometric deformation and size features are fused by calling a feature weighting fusion formula to obtain fused features;
[0026] A preset fusion feature threshold is obtained, and when the fusion feature is not less than the fusion feature threshold, the prefabricated component detection area corresponding to the fusion feature is marked as a defect candidate area, and a spatial clustering algorithm is called to generate the region boundary of the defect candidate area.
[0027] Further, the deep learning detection and segmentation model is used to classify the defect type and segment the defect boundary of the defect candidate area, and the quantitative index of the defect area in the segmentation boundary is calculated by using the three-dimensional point cloud and the infrared image, comprising:
[0028] The fusion features corresponding to the defect candidate area are input into a multi-task convolutional neural network for classification to output the defect type probability corresponding to the defect candidate area, and the largest defect type probability is selected as the defect type of the defect candidate area;
[0029] The binary mask matrix corresponding to the segmented defect region is generated by performing boundary segmentation on the defect candidate region of the determined defect type through a Mask R-CNN segmentation network.
[0030] Based on the binary mask matrix, the segmentation boundaries are calculated on the RGB image, infrared image and ultraviolet image respectively, and the multiple segmentation results obtained are weighted and fused to obtain a defect region set.
[0031] Further, the defect type classification and defect boundary segmentation of the defect candidate region are performed by using the deep learning detection and segmentation model, and the quantitative indicators of the defect region in the segmentation boundary are calculated by using the three-dimensional point cloud and infrared image, which further comprises:
[0032] The corresponding segmentation boundary points in the binary mask matrix are mapped to the three-dimensional point cloud coordinate system corresponding to the point cloud data, the maximum boundary length, the maximum boundary width, the average depth of the defect region and the defect volume are calculated, and the thermal diffusion indicator in the defect region is determined according to the temperature difference distribution characteristics of the defect region;
[0033] All segmentation boundary points in the defect region are mapped to a unified three-dimensional coordinate system, and the maximum boundary length, the maximum boundary width, the average depth of the defect region, the defect volume and the thermal diffusion indicator are integrated to obtain the quantitative indicators.
[0034] Further, the dynamic threshold judgment is performed on the defect region according to the defect type and quantitative indicators of the defect region to determine the defect severity level of the defect region, and the quality evaluation result of the prefabricated component is output, which comprises:
[0035] The production information of the prefabricated component is obtained, and the quantitative indicators and production information are spliced to obtain the feature vector of each defect region;
[0036] A feature vector matrix is constructed based on the feature vectors of all defect regions, a reinforcement learning model based on proximal policy optimization is called to obtain a dynamically adjusted defect judgment threshold combined with historical defect region evaluation results, the defect severity score of each defect region is calculated according to different defect judgment thresholds, and the defect severity level of each defect region is determined according to the defect severity score.
[0037] Further, the dynamic threshold judgment is performed on the defect region according to the defect type and quantitative indicators of the defect region to determine the defect severity level of the defect region, and the quality evaluation result of the prefabricated component is output, which further comprises:
[0038] Based on the defect area, defect severity score and defect position weight, the comprehensive quality score of the prefabricated component is calculated by a comprehensive quality score formula, and the prefabricated component is marked as a scrapped component when the comprehensive quality score does not reach a first threshold value;
[0039] According to the comprehensive quality score and historical defect area evaluation results, the defect development trend of the defect area is estimated, and the prefabricated component is marked by the defect severity level, the comprehensive quality score and the defect development trend to obtain a component quality level label.
[0040] The component quality level label is written into a metadata node of a component digital twin model, and the three-dimensional coordinates and dimensions of each defect area are mapped to the model surface to output the quality evaluation results of the prefabricated component.
[0041] The second aspect of the application discloses a terminal comprising a processor and a storage medium.
[0042] The storage medium is used to store instructions.
[0043] The processor is used to operate according to the instructions to perform the steps of the method of the second aspect.
[0044] The third aspect of the application discloses a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the steps of the method of the second aspect.
[0045] The beneficial effects of the application are that, compared with the prior art, the application has the following advantages:
[0046] (1) The application arranges a multi-spectral imaging device (RGB camera, infrared thermal imager, ultraviolet fluorescence camera) and a three-dimensional structure scanner in the detection area of the prefabricated component, controls the mechanical arm to collect full-coverage images and point cloud data of six surfaces of the component from multiple perspectives, records the collection time stamp, uses a calibration board for spatial registration, and maps the collection results of different sensors to a unified three-dimensional coordinate system, which not only ensures the consistency of the data in position and scale, but also provides a unified and unbiased data basis for subsequent multi-modal fusion and defect detection.
[0047] (2) The application inputs the multi-modal image set and the point cloud data into a multi-modal feature fusion network, extracts surface texture features from the RGB images, extracts temperature difference distribution features from the infrared images, extracts crack penetration fluorescence features from the ultraviolet images, extracts geometric deformation and size features from the point cloud data, and fuses them in the feature space, generates the spatial position and shape range of the defect candidate area using the fusion results, and integrates the surface and internal information, which can significantly improve the integrity and accuracy of the preliminary screening of the component defects.
[0048] (3) The application inputs the defect candidate area into a deep learning detection and segmentation model, classifies the defect type (crack, honeycomb pitting, hole, exposed reinforcement, etc.) of each area and accurately segments the boundary, calculates the length, width, depth, volume and other quantitative indexes of the defect using three-dimensional point cloud and infrared image, and labels them in a unified component three-dimensional coordinate system, which not only improves the intelligent degree of automatic defect detection of the component, but also provides accurate defect quantitative information for quality grading and life prediction. In addition, through the reinforcement learning method, the defect judgment standard is adaptively adjusted according to the component category and production stage, the defect severity level is calculated, the misjudgment caused by environmental and process differences is reduced, and the robustness of the component defect detection is improved. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 is a structural schematic diagram of an AI-based prefabricated part multi-modal visual quality intelligent detection system provided by the application;
[0050] Figure 2 is a flowchart of an AI-based prefabricated part multi-modal visual quality intelligent detection method provided by the application. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0052] As shown in Figure 1 in one embodiment, an AI-based prefabricated part multi-modal visual quality intelligent detection system comprises:
[0053] A multi-modal visual data acquisition unit is configured to acquire multi-modal visual data of a detection area of a prefabricated component and calibrate the multi-modal visual data to construct a multi-modal data set. The multi-view acquisition component is composed of a multi-spectral imaging device and a three-dimensional structure scanner.
[0054] A feature fusion and candidate area extraction unit is configured to input the multi-modal data set into a multi-modal feature fusion network to extract multi-modal features and fuse the multi-modal features in a feature space to generate a defect candidate area of the prefabricated component.
[0055] The multi-modal features include surface texture features extracted from an RGB image, temperature difference distribution features extracted from an infrared image, crack penetration fluorescence features extracted from an ultraviolet image, and geometric deformation and size features extracted from point cloud data.
[0056] The defect type identification and quantitative analysis unit is used to call a deep learning detection and segmentation model to classify defect types and segment defect boundaries in the defect candidate area, and to calculate quantitative indicators of the defect area within the segmentation boundary using 3D point cloud and infrared image.
[0057] The defect level discrimination and component quality assessment unit is used to dynamically determine the defect severity level of the defect area based on the defect type and quantitative indicators, and output the quality assessment results of the precast component.
[0058] like Figure 2 As shown, in one embodiment, an AI-based intelligent detection method for multimodal visual quality of prefabricated components, implemented through the aforementioned AI-based intelligent detection system for multimodal visual quality of prefabricated components, includes the following steps:
[0059] Step S110: Acquire multimodal visual data of the prefabricated component detection area by acquiring components from multiple perspectives, and calibrate the multimodal visual data to construct a multimodal dataset with a unified three-dimensional coordinate system.
[0060] In some embodiments, the AI-based intelligent detection method for multimodal visual quality of prefabricated components provided by the present invention involves a multi-view acquisition component mounted at the end of a robotic arm, comprising a multispectral imaging device and a three-dimensional structure scanner. The multispectral imaging device is used to acquire RGB images, infrared images, and ultraviolet images of the detection area of the prefabricated component using an RGB camera, an infrared thermal imager, and an ultraviolet fluorescence camera, respectively. The three-dimensional structure scanner is used to acquire point cloud data of the monitoring area of the prefabricated component. Step S110 specifically includes the following steps:
[0061] Step S111: Based on the RGB image, infrared image, ultraviolet image, point cloud data and data acquisition timestamp of the detection area of the prefabricated component, and combined with the external dimensions of the prefabricated component and the kinematic parameters of the robotic arm, a multimodal raw dataset is constructed.
[0062] Step S112: The original multimodal dataset is time-synchronized using the same clock source, and the RGB images, infrared images, ultraviolet images and point cloud data are converted to a unified three-dimensional coordinate system using Zhang's calibration method to construct the multimodal dataset.
[0063] In a specific embodiment, the AI-based intelligent detection method for multimodal visual quality of prefabricated components provided by the present invention includes steps 1 to 5:
[0064] Step 1: Multimodal visual data acquisition and synchronous calibration.
[0065] A multispectral imaging device (RGB camera, infrared thermal imager, ultraviolet fluorescence camera) and a three-dimensional structure scanner are arranged in the prefabricated component detection area. The mechanical arm controls the multi-view acquisition of the full coverage images and point cloud data of the six surfaces of the component, records the acquisition timestamp, and uses the calibration board for spatial registration. The acquisition results of different sensors are mapped to a unified three-dimensional coordinate system to ensure the consistency of the data in position and scale, providing a unified and unbiased data basis for subsequent multi-modal fusion and defect detection.
[0066] The method comprises the following sub-steps:
[0067] Sub-step 1.1, multispectral and three-dimensional structure acquisition device deployment.
[0068] Specifically, first, the outer dimensions (length, width, height) of the prefabricated component are determined, the layout parameters (detection table size, lighting conditions) of the detection site are detected, and a sensor group is installed at the end of the mechanical arm in the detection area of the prefabricated component, including an RGB camera (resolution not less than 3840x2160 pixels, field of view angle 70°~90°, used for acquiring surface texture), an infrared thermal imager (temperature measurement accuracy ±0.5℃, thermal sensitivity ≤0.05℃, used for detecting internal temperature difference abnormalities), an ultraviolet fluorescence camera (wavelength range 320~400nm, used for fluorescent penetration crack detection), and a three-dimensional structure scanner (line laser scanning mode, accuracy ±0.2mm, used for generating point cloud data). These sensor devices are installed at the same end of the mechanical arm through a customized bracket, connected to an industrial computer through a gigabit Ethernet interface, and unified data acquisition control is realized.
[0069] Sub-step 1.2, mechanical arm multi-view scanning path planning and acquisition.
[0070] Specifically, according to the outer dimensions (length L, width W, height H) of the prefabricated component, combined with the kinematic parameters (joint angle range, maximum acceleration, and end load) of the mechanical arm, the path set P of the mechanical arm in six surface scanning is calculated, and the expression is:
[0071]
[0072] In the formula, is the pose parameter (combination of position and orientation) of the ith scanning position, is calculated by the forward kinematics equation of the mechanical arm, is the joint angle set of the mechanical arm; is the scanning distance, ranging from 0.5m to 1.5m, optimized according to the focal length of the sensor.
[0073] Then, the RGB camera, infrared thermal imager, ultraviolet fluorescence camera, and 3D scanner are triggered to collect at the same time at each pose point, ensuring timestamp synchronization.
[0074] Sub-step 1.3, time synchronization and space calibration.
[0075] Specifically, the same clock source (industrial computer system clock) is used to time each sensor, and the time stamp of each frame of data collected is recorded. The time synchronization deviation is calculated, which is the difference between the latest time stamp and the earliest time stamp in the same collection cycle. If the time synchronization deviation is not more than 5 ms, it is considered that the time synchronization is qualified, otherwise the collection is retriggered. The Zhang calibration method is used for space calibration. The camera internal and external parameter matrix and distortion coefficient are calculated by shooting the calibration board, and the multi-modal images (RGB image, infrared image, ultraviolet image) and point cloud data are converted to a unified three-dimensional coordinate system, and finally the multi-modal data set after time synchronization and preliminary space registration is obtained.
[0076] Sub-step 1.4, unified coordinate system mapping and accuracy verification.
[0077] Specifically, based on the multi-modal data set after time synchronization and preliminary space registration, the multi-modal data is projected to a unified three-dimensional coordinate system, and the mapping error is calculated by feature point matching , the expression is:
[0078]
[0079] In the formula, N is the number of feature points used for verification; is the real space coordinates corresponding to the i-th feature point in the point cloud data; is the space coordinates corresponding to the i-th feature point in the mapped image.
[0080] When the mapping error ≤0.3mm, it is considered that the registration accuracy is up to standard. After registration accuracy verification, the high-precision multi-modal data set (RGB image, infrared image, ultraviolet image and point cloud data one-to-one corresponding) in the unified three-dimensional coordinate system is finally output, and the mapping error is less than 0.3mm.
[0081] Step S120, calling a multi-modal feature fusion network to extract multi-modal features from the multi-modal data set, and fusing the multi-modal features in the feature space to generate a defect candidate area of the prefabricated part.
[0082] Among them, the multi-modal features include surface texture features extracted from the RGB image, temperature difference distribution features extracted from the infrared image, crack penetration fluorescence features extracted from the ultraviolet image, and geometric deformation and size features extracted from the point cloud data.
[0083] In some embodiments, the AI-based prefabricated part multi-modal visual quality intelligent detection method provided by the application comprises the following steps:
[0084] Step S121, feature extraction of the RGB image is performed through a convolutional neural network to obtain surface texture features of the prefabricated component.
[0085] Step S122, infrared temperature difference and ultraviolet fluorescence are calculated based on the infrared image and the ultraviolet image respectively to obtain temperature difference distribution features and crack penetration fluorescence features.
[0086] Step S123, geometric deviation of the prefabricated component is calculated based on the point cloud data to obtain geometric deformation and size features.
[0087] Step S124, a feature weighting fusion formula is called to perform feature fusion on the surface texture features, the temperature difference distribution features, the crack penetration fluorescence features and the geometric deformation and size features to obtain fused features.
[0088] Step S125, a preset fused feature threshold is obtained, and when the fused features are not lower than the fused feature threshold, a prefabricated component detection area corresponding to the fused features is marked as a defect candidate area, and a spatial clustering algorithm is called to generate a region boundary of the defect candidate area.
[0089] In a specific embodiment, the AI-based prefabricated component multi-modal visual quality intelligent detection method provided by the present application includes the following steps:
[0090] Sub-step 2.1, RGB image surface texture feature extraction.
[0091] Specifically, an RGB image dataset is constructed based on the RGB images collected in step 1, each RGB image has a size of HxW pixels, and each pixel contains three color channels (i.e., R channel, G channel and B channel). A convolutional neural network (CNN) is used to perform convolution operation on the RGB image dataset to extract corresponding surface texture features and construct a corresponding feature matrix based on the extracted surface texture features, the feature matrix has a dimension of MxD, M is the number of RGB image sampling blocks, and D is the texture feature dimension of each sampling block (between 8 and 64).
[0092] Sub-step 2.2, infrared temperature difference feature and ultraviolet fluorescence feature extraction.
[0093] Specifically, based on the infrared image and the ultraviolet image collected in step 1, the infrared temperature difference feature and the ultraviolet fluorescence feature are calculated respectively, wherein the infrared temperature difference feature is composed of the thermal feature values corresponding to all sampling blocks, the thermal feature value is the ratio of the difference between the average temperature of the defect candidate region and the reference temperature of the normal region to the reference temperature of the normal region, and each sampling block has a corresponding thermal feature value; the ultraviolet fluorescence feature is composed of the fluorescence feature values corresponding to all sampling blocks, the fluorescence feature value is the ratio of the difference between the maximum and minimum values of the fluorescence intensity in the sampling block to the maximum value of the fluorescence intensity, and each sampling block has a corresponding fluorescence intensity value.
[0094] Substep 2.3, three-dimensional point cloud geometry feature extraction.
[0095] Specifically, based on the point cloud data collected in step 1, a three-dimensional point cloud dataset is constructed, which is composed of N points, and the geometric deviation between the three-dimensional coordinates of the actual point and the three-dimensional coordinates of the corresponding point in the design model is calculated based on the three-dimensional coordinates of each point The expression is:
[0096]
[0097] In the formula, represents the three-dimensional coordinates of the mth actual point, represents the dth geometric deviation feature value of the mth actual point.
[0098] Substep 2.4, feature fusion and defect candidate region generation.
[0099] Specifically, based on the feature matrices corresponding to the surface texture feature, the infrared temperature difference feature, the ultraviolet fluorescence feature and the geometric deviation feature output by substeps 2.1-2.3, and combining the weight coefficients corresponding to each feature matrix (determined by historical sample training, and all weight coefficients add up to 1), the feature matrices are weighted, and the final fused multi-modal feature values are obtained. Then, based on the preset fusion feature threshold, the sampling blocks corresponding to the multi-modal feature values not lower than the fusion feature threshold are marked as defect candidate regions, and the spatial clustering algorithm (such as DBSCAN) is used to generate the region boundary of the defect candidate region.
[0100] Step S130, using a deep learning detection and segmentation model to classify the defect type and segment the defect boundary of the defect candidate region, and calculating the quantitative index of the defect region in the segmentation boundary using three-dimensional point cloud and infrared image.
[0101] In some embodiments, the AI-based prefabricated multi-modal visual quality intelligent detection method provided by the present application specifically includes the following steps:
[0102] Step S131, input the fusion features corresponding to the defect candidate region into the multi-task convolutional neural network for classification to output the defect type probability corresponding to the defect candidate region, and select the maximum defect type probability as the defect type of the defect candidate region.
[0103] Step S132, the boundary of the defect candidate region of the determined defect type is segmented by the Mask R-CNN segmentation network to generate a binary mask matrix corresponding to the segmented defect region.
[0104] Step S133, based on the binary mask matrix, the segmentation boundary is calculated on the RGB image, infrared image and ultraviolet image respectively, and the multiple segmentation results obtained are weighted and fused to obtain a defect region set.
[0105] In some embodiments, the AI-based prefabricated part multi-modal visual quality intelligent detection method provided by the present application further comprises the following steps:
[0106] Step S134, map the corresponding segmentation boundary points in the binary mask matrix to the three-dimensional point cloud coordinate system corresponding to the point cloud data, calculate the maximum boundary length, the maximum boundary width, the average depth of the defect region and the defect volume, and determine the heat diffusion index in the defect region according to the temperature difference distribution characteristics of the defect region.
[0107] Step S135, map all segmentation boundary points in the defect region to a unified three-dimensional coordinate system, and integrate the maximum boundary length, the maximum boundary width, the average depth of the defect region, the defect volume and the heat diffusion index to obtain quantitative indexes.
[0108] In a specific embodiment, the AI-based prefabricated part multi-modal visual quality intelligent detection method provided by the present application, step 3, defect type recognition and quantitative analysis. The defect candidate region output by step 2 is input into the deep learning detection and segmentation model, and each region is classified (crack, honeycomb pitting, hole, exposed reinforcement, etc.) and the boundary is accurately segmented. At the same time, the length, width, depth, volume and other quantitative indexes of the defect are calculated by using three-dimensional point cloud and infrared image, and are labeled in the unified component three-dimensional coordinate system, which provides accurate defect quantitative information for quality grading and life prediction. Including the following sub-steps:
[0109] Sub-step 3.1, deep learning classification of defect candidate region.
[0110] Specifically, the multi-modal features in the defect candidate region set obtained in step 2 are input into the multi-task convolutional neural network model in the form of a vector for classification, and then the defect type probability corresponding to each defect candidate region is output. In this example, the total number of defect categories is 4.
[0111] Then, the class with the maximum probability is selected as the type label of the corresponding defect region, and the type label is attached to the attribute information of the defect region.
[0112] Sub-step 3.2, defect region accurate boundary segmentation.
[0113] Specifically, an improved Mask R-CNN segmentation network is used to perform fine segmentation in the classified defect region, and a corresponding binary mask matrix is generated, denoted as:
[0114] = {1, pixel belongs to the defect region; 0, otherwise}
[0115] In the formula, is the pixel coordinate, is the pixel value of the segmentation mask corresponding to the pixel coordinate.
[0116] Then, the boundaries are calculated on the multi-modal images respectively, and the multi-modal segmentation results are merged by using a weighted fusion method to reduce the interference of light and noise.
[0117] Sub-step 3.3, defect size and shape parameter calculation.
[0118] Specifically, the boundary points in the segmentation mask are mapped to the three-dimensional point cloud coordinate system, and the maximum boundary distance is calculated as the length and the maximum distance in the vertical direction as the width. If the prefabricated component has a hole or a honeycomb shape, the average depth of the defect is further calculated, which is the average value of the difference between the average height of the normal surface around the defect region and the height of the points inside the defect. Then, the volume of the defect region is calculated according to the point cloud grid area corresponding to each point. Finally, the heat diffusion influence range is estimated by the temperature difference and the area of the defect region, to reflect the scale of the internal defects of the component.
[0119] Sub-step 3.4, defect three-dimensional labeling and coordinate system binding.
[0120] Specifically, all boundary points of the defect region are mapped to the unified three-dimensional coordinate system of the multi-modal data set in step 1, and the length, width, depth, volume and other parameters obtained from sub-steps 3.1 to 3.3 are attached as the attributes of the corresponding defect region. Finally, the defect region information set with these parameters and three-dimensional position is output, which can be directly used for digital twin model.
[0121] Step S140, dynamically thresholding the defect region according to the defect type and quantitative index of the defect region to determine the defect severity level of the defect region, and outputting the quality evaluation result of the prefabricated component.
[0122] In some embodiments, the AI-based prefabricated component multi-modal visual quality intelligent detection method provided by the present application specifically comprises the following steps in step S140:
[0123] Step S141, obtain the production information of the prefabricated component, and splice the quantitative indicators and the production information to obtain the feature vector of each defect area.
[0124] Step S142, construct a feature vector matrix based on the feature vectors of all defect areas, combine the historical defect area evaluation results, call a reinforcement learning model based on a proximal policy optimization to obtain a dynamically adjusted defect judgment threshold, calculate the defect severity score of each defect area according to different defect judgment thresholds, and determine the defect severity level of each defect area according to the defect severity score.
[0125] In some embodiments, the AI-based prefabricated component multi-modal visual quality intelligent detection method provided by the present application specifically further comprises the following steps in step S140:
[0126] Step S143, based on the defect area, the defect severity score and the defect position weight, calculate the comprehensive quality score of the prefabricated component through a comprehensive quality scoring formula, and mark the prefabricated component as a scrap component when the comprehensive quality score does not reach a first threshold.
[0127] Step S144, estimate the defect development trend of the defect area according to the comprehensive quality score and the historical defect area evaluation results, and mark the prefabricated component through the defect severity level, the comprehensive quality score and the defect development trend to obtain a component quality level label.
[0128] Step S145, write the component quality level label into the metadata node of the component digital twin model, and map the three-dimensional coordinates and dimensions of each defect area to the model surface to output the quality evaluation result of the prefabricated component.
[0129] In a specific embodiment, the AI-based prefabricated component multi-modal visual quality intelligent detection method provided by the present application, step 4, dynamic threshold judgment and defect severity calculation. Through the dynamically adjusted judgment threshold, the defect parameters output in step 3 and the component production information (such as type, curing age, surface roughness) are judged for defects, and the reinforcement learning method is used to adaptively adjust the defect judgment standard according to the component category and the production stage to calculate the defect severity level. For example, for the microcracks in the early curing stage, the judgment standard is appropriately relaxed, and for the deep cracks in the load-bearing part, the judgment standard is improved. Step 4 reduces the misjudgment caused by environmental and process differences, and improves the robustness of defect detection, including the following sub-steps:
[0130] Sub-step 4.1, defect and production information feature vector construction.
[0131] Specifically, the length, width, depth, volume, defect type (number), component maintenance age, and component surface roughness of the defect area are spliced to form a defect feature vector corresponding to a single defect area. Finally, all defect feature vectors corresponding to the defect areas are stacked and integrated by rows to obtain a final unified feature vector matrix, which has a dimension of KxD, K is the number of defect areas, and D is the feature dimension number (between 8 and 15) of each defect area.
[0132] Sub-step 4.2, reinforcement learning dynamic threshold adjustment.
[0133] Specifically, a reinforcement learning model based on proximal policy optimization (PPO) is used to adjust the threshold, the reward function is set as the weighted sum between the accuracy rate and the missed detection rate, the policy network outputs the threshold adjustment amount according to the defect feature vector, and the current defect judgment threshold is dynamically adjusted according to the output threshold adjustment amount to obtain a dynamic judgment threshold.
[0134] Sub-step 4.3, defect severity score calculation.
[0135] Specifically, the severity score is calculated by the main dimension of the defect area, the dynamic judgment threshold ratio, and the position weight, and the expression is:
[0136]
[0137] In the formula, is the main dimension of the defect area (length for cracks, diameter for holes, and square root of area for honeycomb surface); is the dynamic judgment threshold corresponding to the defect area; is the position weight coefficient of the defect area (determined by the load-bearing part of the component). When ≥100, it indicates that the defect area has exceeded the allowed defect range.
[0138] Sub-step 4.4, defect grade division and output binding.
[0139] Specifically, the hierarchical rule is used to divide the defect grade: <50, slight defect; 50≤ <80, moderate defect; 80≤ <100, serious defect; ≥100, scrap. After dividing the defect grade of all defect areas, the corresponding defect grade is bound to the corresponding defect area, and finally the defect set with bound defect grade is output.
[0140] Step 5, quality grading evaluation and result generation.
[0141] The defect severity level output from step 4 is input into the quality grading model, combined with factors such as defect number, location, development trend, etc., to calculate the comprehensive quality score of the component and output the automatic grading result (qualified, repaired, scrapped). When the defect distribution and severity reach the scrapping condition, it is directly judged as scrapped; if the defects can be eliminated by repair, it is judged as repaired, and finally the grading result and defect spatial distribution data are bound to the component digital model. Step 5 provides clear quality judgment basis for production process, completes the decision output of detection closed loop, realizes intelligent detection of prefabricated components, including the following sub-steps:
[0142] Sub-step 5.1, comprehensive quality score calculation.
[0143] Specifically, according to the severity score (0-100) of the defect area, combined with the weight coefficient of the defect area location and the total number of defect areas, the comprehensive quality score of the prefabricated component is estimated, and when the comprehensive quality score ≤0, the prefabricated component is directly judged as scrapped.
[0144] Sub-step 5.2, defect development trend evaluation.
[0145] Specifically, based on the defect severity score of the current detection and the total number of defects of the current detection, combined with the defect severity score of the last detection, the defect deterioration rate is estimated, and if the estimated result is not less than 0.2, it is considered that the prefabricated component is rapidly deteriorating.
[0146] Sub-step 5.3, automatic quality grading judgment.
[0147] Specifically, when the grade label of any defect is "scraped" or the comprehensive quality score is less than 60, the prefabricated component is directly judged as scrapped; when the comprehensive quality score is greater than or equal to 60 and less than 80, and the defect deterioration rate estimation result is not less than 0.2, the prefabricated component is judged as needing repair; if the comprehensive quality score is greater than or equal to 80 and all defect levels are marked as "mild" or "moderate", the prefabricated component is judged as qualified.
[0148] Sub-step 5.4, quality grading and digital model binding.
[0149] Specifically, the component quality level evaluation result obtained from sub-step 5.3 is written as a global attribute to the metadata node of the component digital twin model, and the three-dimensional coordinates and size of each defect are mapped to the model surface, combined with the defect spatial distribution data obtained in step 3, to update the parameter structure of the component on the component digital twin model.
[0150] In the description of the application, the terms "first", "second", "third", etc. are used only for descriptive purposes and do not connote or imply relative importance or a specific order. Thus, features identified as "first", "second", etc. can implicitly or explicitly include one or more of the features identified with those terms. The term "plurality" means two or more, unless expressly specified otherwise.
[0151] In the present application, unless specifically defined otherwise, the terms "mounting", "connected", "connecting", "fixed", "fixedly connected", etc. should be construed broadly, for example, they can be fixed connection, detachable connection, or integral; they can be mechanical connection, electrical connection, or both; they can be direct connection, or indirect connection via an intermediate medium; they can be internal connection between two elements, or interaction between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0152] In the present application, unless specifically defined otherwise, the first feature "on" or "under" the second feature can be direct contact between the first and second features, or indirect contact between the first and second features through an intermediate medium. Moreover, the first feature "above", "over" and "on" the second feature can be directly above or obliquely above the second feature, or simply indicate that the first feature is higher than the second feature in horizontal height. The first feature "below", "under" and "under" the second feature can be directly below or obliquely below the second feature, or simply indicate that the first feature is lower than the second feature in horizontal height.
[0153] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in the present application and the features of the different embodiments or examples, without contradiction.
[0154] Any processes or methods described in the flowcharts or otherwise described herein can be understood as representing modules, segments, or portions of code that include one or more executable instructions for implementing specific logic functions or steps, and alternate implementations are possible that include structure that is not shown or described herein, including implementations that use different terminology, structures, or approaches to achieve the same results. The scope of preferred embodiments of the present application includes any implementation that performs the functions described herein, whether explicitly discussed or not.
[0155] Logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be embodied in computer-readable medium, which can be any device or apparatus that can store, communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device. Computer readable medium can include any suitable type of non-transitory storage device, including hard disks, floppy disks, CD-ROMs, DVD-ROMs, Blu-ray discs, RAM, ROM, EEPROM, and the like. Computer readable medium can also include any suitable type of connections, including electrical connections, optical connections, and the like. Examples of a non-transitory computer-readable medium can include a hard disk, a CD-ROM, a Blu-ray disc, a RAM, a ROM, EEPROM, a portable computer diskette, a cache, a flash memory card, etc. For example, a computer program can be stored / distributed on a computer-readable medium such as one or more of the aforementioned types of medium. The computer program can be uploaded to, and executed by, a machine which can be a server or a client machine in a network environment. A mobile computing device also can include computer-readable media in the form of RAM, ROM, flash memory, one or more types of computer-readable media, etc., and the like. A basic hardware implementation of a mobile computing device can include a processing element and input / output elements. The
[0156] It should be understood that aspects of the present application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, or combinations thereof, can be used: a discrete logic circuit having logic gates for implementing logic functions upon data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), and / or the like.
[0157] Those skilled in the art can understand that all or part of the steps of the method carried out by the above-mentioned embodiments can be instructed by a program to related hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it includes one of the steps of the method embodiments or a combination thereof.
[0158] In addition, each functional unit in each embodiment of the present application can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.
[0159] Although the embodiments of the present application have been shown and described above, it should be understood that the above-mentioned embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-mentioned embodiments within the scope of the present application.
Claims
1. An AI-based intelligent inspection method for multimodal visual quality of prefabricated components, characterized in that, The method includes: Multimodal visual data of the detection area of prefabricated components is acquired by acquiring components from multiple perspectives, and the multimodal visual data is calibrated to construct a multimodal dataset with a unified three-dimensional coordinate system. A multimodal feature fusion network is invoked to extract multimodal features from the multimodal dataset, and the multimodal features are fused in the feature space to generate candidate defect regions for prefabricated components; The defect candidate region is classified into defect types and segmented into defect boundaries using a deep learning detection and segmentation model. Quantitative indicators of the defect region within the segmentation boundary are calculated using 3D point cloud and infrared images. Based on the defect type and quantitative indicators of the defect area, a dynamic threshold determination is performed on the defect area to determine the defect severity level of the defect area, and the quality assessment result of the precast component is output. The multimodal features include surface texture features extracted from RGB images, temperature difference distribution features extracted from infrared images, crack penetration fluorescence features extracted from ultraviolet images, and geometric deformation and size features extracted from point cloud data. The step of calling the multimodal feature fusion network to extract multimodal features from the multimodal dataset and fusing the multimodal features in the feature space to generate defect candidate regions for prefabricated components includes: The surface texture features of the prefabricated component are obtained by extracting features from the RGB image using a convolutional neural network. Infrared temperature difference and ultraviolet fluorescence are calculated based on the infrared and ultraviolet images, respectively, to obtain the temperature difference distribution characteristics and crack penetration fluorescence characteristics. The geometric deviation of the prefabricated component is calculated based on the point cloud data to obtain the geometric deformation and dimensional characteristics; The surface texture features, temperature difference distribution features, crack penetration fluorescence features, and geometric deformation and size features are fused using a feature weighted fusion formula to obtain fused features. A preset fusion feature threshold is obtained, and when the fusion feature is not lower than the fusion feature threshold, the detection area of the prefabricated component corresponding to the fusion feature is marked as a defect candidate area. At the same time, a spatial clustering algorithm is called to generate the region boundary of the defect candidate area. The process of using a deep learning detection and segmentation model to classify defect types and segment defect boundaries in the candidate defect regions, and calculating quantitative indicators of the defect regions within the segmentation boundaries using 3D point clouds and infrared images, includes: The fused features corresponding to the defect candidate region are input into a multi-task convolutional neural network for classification, so as to output the defect type probability corresponding to the defect candidate region, and the largest defect type probability is selected as the defect type of the defect candidate region. The Mask R-CNN segmentation network is used to segment the boundaries of multiple defect candidate regions to determine the defect type, and a binary mask matrix corresponding to the segmented defect region is generated. Based on the binary mask matrix, segmentation boundaries are calculated on the RGB image, infrared image and ultraviolet image respectively, and the multiple segmentation results are weighted and fused to obtain a set of defect regions; The step of dynamically thresholding the defect area based on the defect type and quantitative indicators to determine the defect severity level and outputting the quality assessment result of the precast component includes: The production information of the prefabricated components is obtained, and the quantitative indicators and production information are spliced together to obtain the feature vector of each defect area; A feature vector matrix is constructed based on the feature vectors of all defective regions. Combined with the historical defective region evaluation results, a reinforcement learning model based on near-end policy optimization is called to obtain the dynamically adjusted defect judgment threshold. Based on different defect judgment thresholds, the defect severity score of each defective region is calculated, and the defect severity level of each defective region is determined according to the defect severity score.
2. The AI-based intelligent detection method for multimodal visual quality of prefabricated components according to claim 1, characterized in that, The multi-view acquisition component is installed at the end of the robotic arm and consists of a multispectral imaging device and a three-dimensional structure scanner. The multispectral imaging device is used to acquire RGB images, infrared images, and ultraviolet images of the detection area of the prefabricated component through an RGB camera, an infrared thermal imager, and an ultraviolet fluorescence camera, respectively. The three-dimensional structure scanner is used to acquire point cloud data of the detection area of the prefabricated component. The process of acquiring multimodal visual data of the detection area of prefabricated components through multi-view acquisition and calibrating the multimodal visual data to construct a multimodal dataset with a unified three-dimensional coordinate system includes: Based on the RGB image, infrared image, ultraviolet image, point cloud data and data acquisition timestamp of the detection area of the prefabricated component, combined with the external dimensions of the prefabricated component and the kinematic parameters of the robotic arm, a multimodal raw dataset is constructed. The multimodal raw dataset is time-synchronized using the same clock source, and the RGB images, infrared images, ultraviolet images, and point cloud data are converted to a unified three-dimensional coordinate system using Zhang's calibration method to construct the multimodal dataset.
3. The AI-based intelligent detection method for multimodal visual quality of prefabricated components according to claim 1, characterized in that, The method of classifying defect types and segmenting defect boundaries using a deep learning detection and segmentation model, and calculating quantitative indicators of the defect region within the segmentation boundary using 3D point cloud and infrared images, further includes: Map the corresponding segmentation boundary points in the binary mask matrix to the three-dimensional point cloud coordinate system corresponding to the point cloud data, calculate the maximum boundary length, maximum boundary width, average depth and defect volume of the defect region, and determine the heat diffusion index within the defect region based on the temperature difference distribution characteristics of the defect region. All segmentation boundary points within the defect area are mapped to a unified three-dimensional coordinate system, and the maximum boundary length, maximum boundary width, average depth and volume of the defect area, and thermal diffusion index are integrated to obtain the quantitative index.
4. The AI-based intelligent detection method for multimodal visual quality of prefabricated components according to claim 1, characterized in that, The step of dynamically thresholding the defect area based on the defect type and quantitative indicators to determine the defect severity level and outputting the quality assessment result of the precast component further includes: Based on the defect area, defect severity score, and defect location weight, the comprehensive quality score of the prefabricated component is calculated using a comprehensive quality scoring formula, and the prefabricated component is marked as a scrapped component when the comprehensive quality score does not reach a first threshold. The defect development trend of the defect area is estimated based on the comprehensive quality score and historical defect area assessment results. The precast component is then marked using the defect severity level, comprehensive quality score, and defect development trend to obtain a component quality level label. The component quality grade label is written into the metadata node of the component digital twin model, and the three-dimensional coordinates and dimensions of each defect area are mapped to the model surface to output the quality assessment result of the prefabricated component.
5. An AI-based intelligent inspection system for multimodal visual quality of prefabricated components, used to perform the method described in any one of claims 1-4, characterized in that, The system includes: A multimodal visual data acquisition unit is used to acquire multimodal visual data of the detection area of the prefabricated component and to calibrate the multimodal visual data to construct a multimodal dataset. The multi-view acquisition component consists of a multispectral imaging device and a three-dimensional structure scanner. The feature fusion and candidate region extraction unit is used to input the multimodal dataset into the multimodal feature fusion network to extract multimodal features and fuse the multimodal features in the feature space to generate defect candidate regions for prefabricated components. The multimodal features include surface texture features extracted from RGB images, temperature difference distribution features extracted from infrared images, crack penetration fluorescence features extracted from ultraviolet images, and geometric deformation and size features extracted from point cloud data. The defect type identification and quantitative analysis unit is used to call a deep learning detection and segmentation model to classify the defect type and segment the defect boundary of the defect candidate area, and to calculate the quantitative index of the defect area within the segmentation boundary using three-dimensional point cloud and infrared image. The defect level discrimination and component quality assessment unit is used to dynamically determine the defect severity level of the defect area based on the defect type and quantitative indicators of the defect area, and output the quality assessment results of the prefabricated component.
6. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
Industrial defect detection method, system and device and storage medium
CN118967672A
Defect detection method for semiconductor packaging material based on deep learning
CN120525859A