Intelligent identification method and system for surface defects of autoclaved aerated concrete member

Through multi-view camera array and spatial transformation matrix technology, combined with deep learning and image processing methods, the comprehensive, blind spotless and high accuracy detection of surface defects of autoclaved aerated concrete components is achieved, solving the problems of low efficiency and inability to quantitatively evaluate traditional detection methods, and meeting the quality control needs of modern large-scale production lines.

CN120182253AActive Publication Date: 2025-06-20SHAANXI NEW FASHION CONSTR & INSTALLATION ENG CO LTD +1

Patent Information

Application Number
CN202510642916.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The surface defect detection of traditional autoclaved aerated concrete components relies on manual visual inspection, with low efficiency, different standards, and no quantitative evaluation, making it difficult to meet the quality control needs of modern large-scale production lines.

Method used

Multi-view camera array and spatial transformation matrix optimization technology are used to achieve all-round shooting and blind spot detection of the surface of autoclaved aerated concrete components. Combined with light equalization, adaptive contrast enhancement and variable-scale bilateral filtering technology, improve the visibility and edge clarity of defects. Use depth-separable convolution and multi-scale texture coding structures to extract texture characteristics of different types of defects, and realize effective detection of multi-scale defects through deformable convolution, multi-scale spatial pooling and attention mechanisms.

Benefits of technology

The completeness and accuracy detection of surface defects of autoclaved aerated concrete components is achieved, the accuracy and robustness of defect identification is improved, and it can adapt in complex environments, and it provides a scientific basis for production process optimization and quality control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182253A_ABST
    Figure CN120182253A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of surface defect recognition, and discloses an intelligent recognition method and system for surface defects of an autoclaved aerated concrete member. The method comprises the following steps: carrying out omnibearing shooting and spatial transformation matrix mapping on the surface of the autoclaved aerated concrete member to obtain a preprocessed image; performing texture feature extraction on the preprocessed image to obtain texture feature data; performing spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relation feature data; performing feature importance weighting and bidirectional cross fusion processing on the texture feature data and the spatial relation feature data to obtain comprehensive defect feature representation; and performing surface defect analysis on the autoclaved aerated concrete member based on the comprehensive defect feature representation, and outputting a surface quality comprehensive score and grade division result. According to the method, the problems of dead angles and image deformation existing in traditional single-view-angle detection are solved, and the completeness and accuracy of defect detection are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of surface defect recognition, and particularly to an intelligent recognition method and system for surface defects of autoclaved aerated concrete components. Background Art

[0002] As a lightweight and high-strength building material, autoclaved aerated concrete is widely used in the construction industry due to its excellent heat insulation, fire resistance, and construction convenience. However, during the production process, due to fluctuations in process parameters such as cement ratio, forming pressure, fine aggregate content, curing time, and autoclaving temperature, various types of defects such as pores, cracks, and depressions often appear on the surface of components. These defects not only affect the appearance quality of components but also have a significant impact on their structural performance. In severe cases, they can lead to reduced strength, decreased durability, and even structural failure during use. Traditional detection of surface defects of autoclaved aerated concrete components mostly relies on manual visual inspection, which has problems such as low efficiency, inconsistent standards, inability to quantitatively evaluate, and difficulty in correlating production process parameters, and it is difficult to meet the quality control requirements of modern large-scale production lines.

[0003] Existing automated vision detection technologies face many challenges when applied to the recognition of surface defects of autoclaved aerated concrete components. The surface of autoclaved aerated concrete components has characteristics of grayish-white color and low contrast, resulting in low distinguishability between defects and the background, and conventional image processing methods are difficult to effectively extract defect features. Secondly, the types of surface defects of autoclaved aerated concrete are complex, with variable shapes and large size spans, from tiny pores to large-area peeling, and it is difficult to comprehensively capture them with a single feature extraction method. Thirdly, traditional single-view detection methods have dead angles and image deformation problems, and it is impossible to perform dead-angle-free detection on the entire surface of components. Fourthly, there is a lack of a systematic method for correlating defect formation with production process parameters, and it is difficult to provide quantitative guidance for production process optimization. Summary of the Invention

[0004] The present invention provides an intelligent recognition method and system for surface defects of autoclaved aerated concrete components. The present invention solves the dead-angle problem and image deformation problem existing in traditional single-view detection, and ensures the integrity and accuracy of defect detection.

[0005] In a first aspect, the present invention provides an intelligent recognition method for surface defects of autoclaved aerated concrete components, and the intelligent recognition method for surface defects of autoclaved aerated concrete components includes: Performing all-round shooting and spatial transformation matrix mapping on the surface of autoclaved aerated concrete components to obtain a preprocessed image; Performing texture feature extraction on the preprocessed image to obtain texture feature data; Performing spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data; Perform feature importance weighting and two-way cross-fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation; Based on the comprehensive defect feature representation, perform surface defect analysis on the autoclaved aerated concrete component, and output the comprehensive surface quality score and the result of grade classification.

[0006] In a second aspect, the present invention provides an intelligent surface defect recognition system for autoclaved aerated concrete components, and the intelligent surface defect recognition system for autoclaved aerated concrete components includes: A mapping module, configured to perform all-round shooting and spatial transformation matrix mapping on the surface of the autoclaved aerated concrete component to obtain a preprocessed image; A feature extraction module, configured to extract texture feature data from the preprocessed image; A feature analysis module, configured to perform spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data; A cross-fusion module, configured to perform feature importance weighting and two-way cross-fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation; An output module, configured to perform surface defect analysis on the autoclaved aerated concrete component based on the comprehensive defect feature representation, and output the comprehensive surface quality score and the result of grade classification.

[0007] In the technical solution provided by the present invention, through the multi-view camera array and spatial transformation matrix optimization technology, the full coverage shooting of each surface of the autoclaved aerated concrete component is realized, solving the dead angle problem and image deformation problem existing in traditional single-view detection, and ensuring the integrity and accuracy of defect detection. Aiming at the characteristics of the grayish-white color and low contrast of the autoclaved aerated concrete surface, technologies such as illumination equalization, adaptive contrast enhancement, and variable-scale bilateral filtering are adopted, significantly improving the visibility of defects and the edge sharpness. The depthwise separable convolution and multi-scale texture coding structure are adopted to specifically optimize the texture characteristics of different types of defects such as pores, cracks, and depressions on the autoclaved aerated concrete surface, while greatly reducing the number of parameters and computational complexity, and maintaining a high feature extraction ability. Through the combination of deformable convolution, multi-scale spatial pooling, and attention mechanism, the effective detection of defects of different scales from tiny pores to large-area peeling is realized, solving the problem that traditional methods are difficult to adapt to the recognition of multi-scale defects at the same time. Through feature importance weighting and bidirectional cross-attention mechanism, the complementary advantages of texture features and spatial features are fully integrated, and the scientific knowledge constraints of autoclaved aerated concrete materials are introduced to improve the accuracy and robustness of defect recognition, especially the adaptability under complex environmental conditions. A multi-level quality evaluation system integrating defect classification, location, and process parameters is constructed, and a quantitative relationship model between defect formation and process parameters such as cement ratio, forming pressure, and fine aggregate content is established, providing a scientific basis for production process optimization and quality control, and realizing the transformation from simple defect detection to comprehensive quality evaluation and process guidance. Through model lightweight design and structural optimization, the system can run in real time on the edge equipment of the production line, meeting the real-time detection requirements in the industrial production environment, and providing a reliable real-time quality monitoring means for the production of autoclaved aerated concrete components. Description of the Drawings

[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0009] Figure 1 It is a schematic diagram of an embodiment of the intelligent recognition method for surface defects of autoclaved aerated concrete components in an embodiment of the present invention; Figure 2 It is a schematic diagram of an embodiment of the intelligent recognition system for surface defects of autoclaved aerated concrete components in an embodiment of the present invention. Detailed Embodiments

[0010] An embodiment of the present invention provides a method and system for intelligent identification of surface defects of autoclaved aerated concrete components. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0011] For ease of understanding, the specific process of the embodiment of the present invention will be described below. Please refer to Figure 1 , an embodiment of the method for intelligent identification of surface defects of autoclaved aerated concrete components in the embodiment of the present invention includes: Step S101, perform omnidirectional shooting and spatial transformation matrix mapping on the surface of the autoclaved aerated concrete component to obtain a preprocessed image; It can be understood that the execution subject of the present invention can be an intelligent identification system for surface defects of autoclaved aerated concrete components, or a terminal or a server. Specifically, it is not limited here. The embodiment of the present invention takes the server as the execution subject as an example for illustration.

[0012] Specifically, when constructing the visual perception platform, the main perspective camera is vertically installed directly above the autoclaved aerated concrete component conveyor belt to capture the entire top surface of the component. At the same time, N auxiliary perspective cameras are evenly arranged along the outer contour around the component. These cameras form an equiangular circular arrangement around the component, and high-resolution industrial-grade CCD devices are used to ensure seamless coverage of the component surface in space. The multi-perspective camera array constitutes the core structure of the system's image acquisition. By establishing a high-speed communication link with the central processing and control unit, synchronous triggering of all cameras and parallel acquisition of image data are achieved. After completing the synchronous image capture, an original image set containing component images from multiple different perspectives is obtained. Independent internal parameter calibration operations are performed on each camera to obtain core parameters such as the focal length, principal point position, radial distortion coefficient, and tangential distortion coefficient of each camera through a standard checkerboard calibration board. Based on the internal parameter, the spatial geometric relationship between each camera and the unified world coordinate system is established, and the spatial position and attitude of each camera are described by calculating its external parameter matrix, realizing the transformation mapping relationship between the component's three-dimensional coordinate system and each camera's coordinate system. Based on the positions of the matching feature points common to the adjacent perspective images in the original image set, a set of homography matrices between the cameras is constructed through an image registration algorithm to obtain the first set of spatial transformation matrices. This matrix set describes how the images are projected onto the same geometric plane through perspective transformation between adjacent perspectives. Using the three-dimensional geometric constraints constructed by the external parameters of each camera, an optimization strategy of minimizing the reprojection error is used to iteratively solve the first set of spatial transformation matrices. Under the guidance of the gradient descent algorithm, each set of transformation parameters is gradually adjusted until the error change in two consecutive rounds is lower than the set threshold, and the optimal second set of spatial transformation matrices is converged. Based on the second set of spatial transformation matrices, all original images are subjected to mapping transformation processing, so that the images from different perspectives are fused and reconstructed in a unified spatial framework, forming a corrected image set with no distortion, geometric alignment, and full coverage of the component surface. Image enhancement processing is performed on each corrected image, including illumination equalization, contrast enhancement, and denoising operations. In the illumination equalization stage, a brightness background model is constructed through Gaussian filtering, and brightness normalization is combined with the target average brightness value to solve the bright and dark area deviation caused by uneven light sources; then, adaptive histogram specification technology is used to perform contrast enhancement to improve the weak visibility of details caused by light and light material colors; a variable-scale bilateral filter is used to achieve local area adaptive smoothing, suppressing random noise interference while retaining edge details, and obtaining a preprocessed image.

[0013] In this embodiment, based on the internal parameters obtained by each camera during the calibration phase, including focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient, a geometric distortion repair operation is performed on the surface full coverage correction image. By introducing a distortion correction function, these parameters are embedded in the imaging geometry model, and the spatial distortion caused by lens distortion is corrected pixel by pixel to obtain a standard component surface image. The standard component surface image is subjected to an illumination equalization operation. The image is blurred by using a Gaussian filtering method. This process performs a low-pass filtering operation on the image in the spatial domain, effectively removes high-frequency detail information, and only retains the background brightness distribution trend to generate a background brightness reference image. Using this background image as a reference, the ratio between the target brightness value and the current brightness value of each pixel is calculated to construct a set of two-dimensional correction coefficient matrices. The correction coefficient matrix performs a pixel-by-pixel multiplication transformation within the entire image range, thereby realizing the normalization adjustment of the image brightness and outputting an illumination equalization image. On the basis of completing illumination equalization, in order to strengthen the weak contrast characteristics of the gray-white material on the surface of the component in the image, a histogram specification technology is introduced to perform contrast enhancement processing on the image. Perform pixel-level statistical analysis on the illumination-equalized image, calculate its grayscale histogram, and construct a cumulative distribution function based on the cumulative frequency of grayscale values. Then map and match the function with the target distribution function preset by the system, and adjust the original pixel grayscale distribution to the target distribution state through the inverse transformation of the function, so as to effectively stretch the grayscale dynamic range of the original image, so that the surface micro-depression, cracks and surrounding substrates form a more obvious distinction in the grayscale domain, and obtain a contrast-enhanced image. Perform structure-preserving noise reduction processing. Construct a variable-scale bilateral filtering mechanism to perform adaptive parameter adjustment according to the local gradient characteristics of each area of ​​the image. Calculate the gradient response at the image pixel level to evaluate the edge complexity or texture change amplitude of the current area, and then dynamically adjust the spatial kernel size and range parameters of the bilateral filter based on this gradient value, that is, control its spatial weight and grayscale weight of the neighboring pixels respectively. In the edge area, the filter enhances the edge retention ability by reducing the spatial kernel, and expands the kernel range in the flat area to enhance the smoothing effect. The variable-scale bilateral filter under adaptive parameter configuration is applied to the contrast-enhanced image to ensure that the defect edge structure is completely preserved while effectively filtering out the background noise introduced in the image due to uneven lighting, material reflection, and equipment jitter to obtain a preprocessed image.

[0014] Step S102, extracting texture features from the preprocessed image to obtain texture feature data; Specifically, the preprocessed image is input into the first depthwise separable convolutional layer, and the initial texture feature information is extracted through depthwise convolution operations with a relatively large convolutional kernel size to obtain the first-layer texture features. Then, these features are continued to be input into the second depthwise separable convolutional layer, where more refined edge and texture structures are extracted in a deeper receptive field, and the second-layer texture features are output. The second-layer features are continued to be input into the third depthwise separable convolutional structure to extract higher-order texture semantic patterns, forming the third-layer texture features. Subsequently, the final deep texture features are extracted through the fourth depthwise separable convolution, and the fourth-layer texture features with rich context information and multi-scale structure expression ability are formed. The entire process constructs a hierarchical texture representation network by stacking layers, gradually accumulating local information such as edges and corners at the low level to the texture combination patterns at the high level, strengthening the model's ability to analyze complex defects on the component surface. To improve the model's ability to distinguish texture features of different-scale defects (such as microcracks, holes, and spalling areas) after obtaining the above four levels of texture features, a parallel dilated convolution mechanism is introduced. The four levels of texture features are input into four groups of dilated convolution networks in parallel, and each group is configured with a different dilation rate to expand the receptive field without increasing the number of parameters, thereby extracting the response information of each layer of features at different spatial scales and obtaining a group of multi-scale texture feature groups with scale differentiation. At the same time, to highlight the region most sensitive to edge changes in the texture feature map, a bidirectional gradient enhancement mechanism is designed and specifically applied to the fourth-layer texture feature map. This mechanism calculates the gray difference between adjacent pixels in the horizontal direction of the fourth-layer texture feature map to obtain the first-direction gradient, calculates the gray change in the vertical direction to obtain the second-direction gradient, explicitly extracts pixel-level edge mutation information through an absolute difference operation, and then fuses the gradient feature maps in the horizontal and vertical directions through a group of lightweight convolutional networks to generate a bidirectional gradient feature map, thereby enhancing the expression ability of microstructural features such as crack boundaries and edge jumps. The multi-scale texture feature groups obtained by dilated convolution are weighted and fused in the channel dimension. During the fusion process, a channel attention mechanism is introduced to assign different weights to the feature maps at different scales to adapt to the differences in scale sensitivity of different defect types. The fused texture feature groups and the bidirectional gradient features are concatenated along the channel dimension to obtain the texture feature data.

[0015] Step S103: Perform spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data; Specifically, a fusion modeling of the preprocessed image and the texture feature data is performed through a spatial feature analysis mechanism. This mechanism is based on the lightweight MobileNetV3 backbone network. Through its depth convolution structure, the preprocessed image is feature-extracted to obtain an initial spatial feature map, which reflects low-level spatial information such as the geometric distribution, texture density, light response, and edge contour of the component surface in different regions. The initial spatial feature map and the texture feature data are fused and connected in the channel dimension, and a 1×1 convolution operation is introduced to perform information compression and redistribution on the connected multi-channel features, forming a set of target fusion feature maps with the joint characteristics of spatial position and texture structure. To improve the model's perception ability for different-scale defect morphologies (such as micro-cracks, large-area spalling, irregularly distributed pores), a parallel convolution processing path is applied to the target fusion features. Each convolution structure is configured with different receptive fields and dilation rates to extract spatial semantic features at different scale levels, forming a set of multi-scale spatial feature maps; at the same time, the target fusion features are fed into the global average pooling module to compress their spatial dimensions, obtaining a one-dimensional channel vector, and the channel weight structure is reconstructed through a 1×1 convolution to obtain global context features. Considering that the spatial importance of different regions is not consistent, a spatial attention mechanism is introduced to construct a spatial attention map based on the target fusion features. This process is carried out by cascading the feature map after performing maximum pooling and average pooling operations respectively, and then fusing them through a 7×7 large receptive field convolution kernel. Subsequently, the sigmoid function is used for normalization to generate the spatial attention map, which reflects the weight levels of each position in the current image in the discrimination task. The spatial attention map and the target fusion features are subjected to an element-wise multiplication operation to complete the attention weighting operation, and a spatially weighted feature map is output. The multi-scale spatial features, global context features, and spatially weighted features are concatenated in the channel dimension to comprehensively converge their complementary information, and a standard convolution operation is introduced to perform feature fusion and dimension unification on the concatenated multi-channel features, generating spatial relationship feature data. Step S104: Perform feature importance weighting and bidirectional cross-fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation; Specifically, global average pooling operations are respectively applied to the texture feature data and the spatial relationship feature data to compress their high-dimensional feature maps into a set of channel-level statistical vectors, thereby retaining the response distribution under the global context semantic background. The channel vectors are input into a non-linear feature encoder composed of a multi-layer perceptron, and feature mapping and activation normalization processing are performed on them, respectively outputting the texture feature importance score and the spatial feature importance score. Based on the feature importance evaluation results, a weighting operation is performed on the original texture feature data and the spatial relationship feature data, and each channel is proportionally scaled using its corresponding importance score as the weight coefficient to obtain the initial weighted fusion feature. A bidirectional cross-attention mechanism is introduced to construct an explicit semantic interaction path between the texture and the space. This mechanism regards the texture feature data as the query item, and the spatial relationship feature data as the set of keys and values, calculates the similarity matrix between the two in the channel dimension and spatial position, measures the correlation between each texture feature channel and the spatial feature through dot product operation, and normalizes the similarity matrix through the Softmax function to obtain the attention map from texture to space; subsequently, the mapping path from space to texture is constructed in reverse, using the spatial feature as the query item and the texture feature as the key-value pair, and similarity calculation and normalization are performed again to generate the attention map from space to texture. The original features are weighted and aggregated based on the two attention maps: the spatial feature is weighted and summed using the attention map from texture to space to obtain the enhanced result of the spatial feature guided by the texture; the texture feature is weighted and summed using the attention map from space to texture to generate the enhanced result of the texture feature guided by the space. These two results are concatenated and fused to form the bidirectional cross-attention feature, capturing the semantic resonance and complementary responses between cross-modalities. On this basis, the initial weighted fusion feature and the bidirectional cross-attention feature are concatenated in the channel dimension, and the material constraint logic regarding the surface defects of autoclaved aerated concrete components is introduced as a prior regulation mechanism. This material constraint consists of the physical form features reflected by the component defect types, including the high aspect ratio of cracks, the closed boundary characteristics of pores, the uniform texture and irregular edge structure of the spalling area, etc., and is transformed into a loss term through the morphological constraint function defined in the material rule library and embedded as a weight adjustment factor in the feature fusion stage. Under the integration of the deep neural network fusion module, the fused feature after channel concatenation and the material prior constraint are jointly mapped into a unified feature space, and a comprehensive defect feature representation with global judgment ability, local expression ability, and material consistency is output.

[0016] Step S105: Based on the comprehensive defect feature representation, perform surface defect analysis on the autoclaved aerated concrete component, and output the comprehensive surface quality score and the grade division result.

[0017] Specifically, the high-dimensional comprehensive defect feature map is restored to a predicted image of the same size as the original image. The feature map is upsampled using a layer-by-layer transposed convolution structure, and the resolution is gradually restored through reverse inference of convolution with a stride of 2. A Sigmoid function is applied in the output stage to compress the result into a probability map between 0 and 1, obtaining the probability of defect occurrence corresponding to each pixel point. Subsequently, the OTSU adaptive threshold algorithm is applied to perform binary processing on the probability map to form a defect binary map, marking the spatial positions and morphological contours of each potential defect area. For each connected region in the defect binary map, contour detection and regional feature extraction operations are performed, including multiple-dimensional indicators such as area, aspect ratio, boundary complexity, average gray value, and edge texture characteristics to form a regional feature vector. Subsequently, the feature vector is input into a pre-trained multi-class defect discrimination network, and each defect area is classified through a combination of convolution extraction and full connection discrimination, outputting classification results of specific defect types such as cracks, pores, and peeling. Based on the second set of spatial transformation matrices, the surface defects detected from each perspective of the autoclaved aerated concrete component are mapped to the three-dimensional coordinate system of the component to achieve unified spatial projection of the defect position information. To avoid repeated detection of the same defect from multiple perspectives, a defect fusion algorithm based on the joint judgment of intersection over union and confidence is used to eliminate and merge overlapping regions. That is, if the overlap degree of a certain defect area exceeds the threshold and the confidence is low from multiple perspectives, the area is automatically discarded to construct an all-round defect distribution map. Pixel-level area statistics are performed on the defect binary map, and the surface defect rate index is formed by calculating the proportion of the total surface area occupied by all defect pixels to characterize the overall defect severity of the component. The defect classification results and defect area data are combined for modeling, and corresponding weight coefficients are set according to the degree of damage of each type of defect to the mechanical properties to establish an initial surface quality index. According to the data such as cement ratio, forming pressure, and fine aggregate content recorded in the structural parameters, the initial quality index is weighted and corrected using a preset process influence coefficient model to obtain a weighted surface quality index. Regional density uniformity analysis is performed on the all-round defect distribution map. The surface of the component is divided into multiple regional units using a grid division strategy, and the defect density difference within each unit is calculated. Through variance standardization processing, a defect distribution uniformity index is formed. The lower this index, the more concentrated the defect distribution, and the higher it is, the more evenly the defects are diffused, thus providing risk warnings of different degrees to the structural stability. The surface defect rate, weighted surface quality index, and defect distribution uniformity index are jointly input into the surface quality scoring model. This model performs weighted fusion operations on the three indicators by setting a weight function and outputs a surface quality comprehensive score value between 0 and 100. The system matches the preset quality grading standard according to this score value, generates the component quality grade classification result, and binds the score result to the original component identifier for output.

[0018] In the embodiments of the present invention, through the multi-view camera array and spatial transformation matrix optimization technology, the full coverage shooting of each surface of autoclaved aerated concrete components is realized, solving the dead angle problem and image deformation problem existing in traditional single-view detection, and ensuring the integrity and accuracy of defect detection. Aiming at the characteristics of the grayish-white color and low contrast of the autoclaved aerated concrete surface, technologies such as illumination equalization, adaptive contrast enhancement, and variable-scale bilateral filtering are adopted, significantly improving the visibility of defects and the edge sharpness, laying a foundation for subsequent feature extraction. The depthwise separable convolution and multi-scale texture coding structure are adopted to specifically optimize the texture characteristics of different types of defects such as pores, cracks, and depressions on the autoclaved aerated concrete surface, while greatly reducing the number of parameters and computational complexity, and maintaining a high feature extraction ability. Through the combination of deformable convolution, multi-scale spatial pooling, and attention mechanism, the effective detection of different-scale defects from tiny pores to large-area peeling is realized, solving the problem that traditional methods are difficult to adapt to the recognition of multi-scale defects simultaneously. Through feature importance weighting and bidirectional cross-attention mechanism, the complementary advantages of texture features and spatial features are fully integrated, and the scientific knowledge constraints of autoclaved aerated concrete materials are introduced, improving the accuracy and robustness of defect recognition, especially the adaptability under complex environmental conditions. A multi-level quality assessment system integrating defect classification, localization, and process parameters is constructed, and a quantitative relationship model between defect formation and process parameters such as cement ratio, forming pressure, and fine aggregate content is established, providing a scientific basis for production process optimization and quality control, and realizing the transformation from simple defect detection to comprehensive quality assessment and process guidance. Through model lightweight design and structure optimization, the system can run in real time on the edge equipment of the production line, meeting the real-time detection requirements in the industrial production environment, and providing a reliable real-time quality monitoring means for the production of autoclaved aerated concrete components.

[0019] In a specific embodiment, the process of executing step S101 may specifically include the following steps: Install the main-view camera directly above the conveyor belt of the autoclaved aerated concrete component, and distribute N auxiliary-view cameras in a ring around the autoclaved aerated concrete component to obtain a multi-view camera array; Start the multi-view camera array to synchronously shoot the autoclaved aerated concrete component to obtain an original image set, and perform individual internal parameter calibration on each camera in the multi-view camera array to obtain the focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient of each camera; Establish the conversion relationship between the component three-dimensional coordinate system and each camera coordinate system according to the focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient of each camera to obtain the external parameter matrix of each camera; Based on the position relationship of the matching feature points of each view image in the original image set, calculate the homography matrix between adjacent cameras to obtain the first set of spatial transformation matrices; Based on the external parameter matrices of each camera, the first set of spatial transformation matrices is optimized by minimizing the reprojection error to obtain the second set of spatial transformation matrices, and the original image set is mapped and transformed by the second set of spatial transformation matrices to obtain the surface full-coverage corrected images; The surface full-coverage corrected images are subjected to illumination equalization, contrast enhancement, and noise reduction processing to obtain the preprocessed images.

[0020] Specifically, the main perspective camera is vertically installed directly above the autoclaved aerated concrete component conveyor belt to cover the upper surface area of the component as the main imaging reference. At the same time, a circular path is set around the component, and N auxiliary perspective cameras are installed at equal intervals to form an annular closed structure in the horizontal direction, ensuring that the included angle between the optical axes of adjacent cameras is 360° / N. All cameras use industrial-grade CCD image sensors with a resolution of not less than 1920×1080 and a frame rate greater than 60fps to capture clear images under high-speed transmission conditions. They are connected to the central image processing and control unit through an Ethernet link and establish a communication mechanism with the production line PLC control module to achieve centralized control and synchronous triggering of the entire camera array. When the component moves to the detection area, the system receives the position signal from the PLC controller and starts the multi-perspective camera array to perform synchronous shooting operations on the current component, ensuring that all perspective images are exposed at the same moment to avoid image distortion or occlusion defects caused by time differences. After completing image acquisition, an original image set including the main perspective image and N auxiliary perspective images is obtained. To eliminate the geometric distortion introduced by the lens structure, each camera is individually calibrated for its internal parameters. The calibration method is based on a standard checkerboard target image. By taking calibration board images from multiple angles, the corner coordinates in each image are extracted, and an internal parameter solution equation is constructed according to the pinhole camera model to solve the focal length parameter, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient corresponding to the camera. These parameters are used to model and repair the non-linear distortion during the imaging process. After obtaining the internal parameters of the cameras, the images of each camera are projected and transformed into a unified three-dimensional component coordinate system. For this purpose, an external parameter matrix of the camera is established, which consists of the rotation matrix and translation vector of each camera relative to the component coordinate system, representing the transformation relationship between the camera coordinate system and the world coordinate system. The relative positions between the component and each camera are measured by jointly using the calibration target and the laser scale, and the external parameter matrix is solved in combination with the internal parameter results. Feature point matching operations are performed based on the overlapping areas of the perspective images in the original image set. Robust feature extraction algorithms such as SIFT or ORB are selected to extract key point information, and the FLANN or BF matcher is used to complete cross-image matching. After obtaining the corresponding relationship of feature points between each pair of adjacent cameras, a homography matrix solution equation is constructed according to the perspective projection model, and the RANSAC algorithm is used to eliminate incorrect matches and solve the homography transformation matrix between each pair of adjacent images, forming the first set of spatial transformation matrices. This matrix set is used to describe the two-dimensional spatial projection relationship from one perspective image to its adjacent image. Based on the external parameter matrices of each camera, the first set of spatial transformation matrices is optimized, and the reprojection error of feature points between different perspective images is minimized during the transformation process.Adopt the least - squares optimization strategy to construct an objective function which is the sum of the squares of the Euclidean distances between the corresponding positions of the matching points among all perspectives after mapping. Use the gradient - descent method for iterative calculation until the error change in two consecutive rounds is lower than the set threshold, and obtain a converged set of second - space transformation matrices. Based on this set of second - space transformation matrices, perform a unified mapping transformation on the original image set, project the images from each perspective onto the same plane or the coordinate domain of a three - dimensional component, and generate a corrected image set that covers the entire surface of the component, is seamlessly aligned, and geometrically consistent. Perform image enhancement processing on the corrected image set, and this process includes operations such as illumination equalization, contrast enhancement, and noise reduction filtering. In the illumination equalization stage, perform a Gaussian filtering operation on the image to obtain a background brightness distribution map, calculate the ratio of the target average brightness to the current brightness, and construct an illumination compensation coefficient matrix, which is applied to the original image to complete the normalization of the image illumination and solve problems such as shadows, vignetting, and high - light areas caused by the directionality of on - site light sources; perform contrast enhancement on the image based on the histogram specification technique. By mapping the cumulative distribution function of the image gray - level histogram to a preset target distribution function, expand the dynamic range of the image gray - level, thereby enhancing the visibility of fine cracks and low - contrast defects under a light - colored background; to suppress the high - frequency noise generated after enhancement and avoid the blurring and loss of edge details, introduce a variable - scale bilateral filter, adaptively adjust the spatial kernel size and range - domain weight of the filter according to the local gradient value of the image, perform strong smoothing in texture - flat areas, and retain image details in structural edge areas to achieve a refined improvement in the overall quality of the image. Through image acquisition, geometric correction, spatial reconstruction, and image enhancement processing, a pre - processed image set is obtained.

[0021] In a specific embodiment, the process of performing illumination equalization, contrast enhancement, and noise reduction on the surface - fully - covered corrected image to obtain the pre - processed image may specifically include the following steps: According to the focal length, principal - point coordinates, radial distortion coefficients, and tangential distortion coefficients of each camera, perform distortion correction on the surface - fully - covered corrected image to obtain a standard component surface image; Perform Gaussian filtering on the standard component surface image to obtain an image representing the background brightness distribution; Calculate a correction coefficient matrix based on the image representing the background brightness distribution, and apply the correction coefficient matrix to the standard component surface image to obtain an illumination - equalized image; Perform histogram statistical analysis on the illumination - equalized image, calculate the cumulative distribution function, and map the cumulative distribution function to a preset target distribution function to obtain a contrast - enhanced image; Calculate the local gradient value of the contrast - enhanced image, and adaptively adjust the filter kernel size and range - domain parameters of the variable - scale bilateral filter according to the local gradient value to obtain target filtering parameters; Perform variable-scale bilateral filtering on the contrast-enhanced image based on the target filtering parameters to obtain a preprocessed image.

[0022] Specifically, the original image is distorted and repaired based on the camera imaging model. Since the acquisition equipment of each perspective image is an industrial-grade CCD camera, there is radial distortion caused by the lens curvature structure and tangential distortion caused by the lens assembly error in the imaging process. Therefore, based on the internal parameters such as focal length, principal point coordinates, radial distortion coefficient and tangential distortion coefficient obtained in the previous camera calibration stage, a distortion correction mapping model is established. According to the pinhole imaging and distortion correction model, the position of each pixel in the image is remapped, and accurately back-projected to the theoretical distortion-free plane to generate a standard component surface image. The brightness modeling of the standard component surface image is carried out to solve the problems of uneven illumination, local overexposure and shadow occlusion in the factory site environment. The Gaussian filtering operation is introduced to perform spatial smoothing on the standard image. In this process, a Gaussian kernel of a fixed scale is selected to perform a convolution operation on the image, thereby eliminating high-frequency changes in details and retaining the global brightness change trend. An image representing the background brightness distribution is obtained, which represents the brightness reference image when the illumination on the component surface should be uniform under ideal conditions. According to the ratio between the background brightness distribution map and the expected target brightness, the illumination correction coefficient matrix is ​​calculated. The value at each pixel position of the matrix represents the ratio factor between the original pixel brightness and the target brightness. By applying the coefficient matrix pixel by pixel to the standard image, that is, performing the pixel-level brightness normalization operation, the illumination equalization processing is completed to obtain the illumination equalization image. The illumination equalization image is subjected to histogram statistical analysis to construct the grayscale histogram of the original image, and the cumulative distribution function is calculated based on the histogram to reflect the cumulative probability distribution of each grayscale value in the entire image. The cumulative distribution function is mapped to the target distribution function preset by the system, which is usually a uniform distribution or an S-shaped distribution with enhanced intermediate grayscale response capability, and the grayscale of each pixel of the original image is adjusted through reverse mapping of the cumulative distribution function, thereby redistributing the grayscale range of the image, so that the low-gray area is stretched and the high-gray area is compressed, forming a contrast-enhanced image with enhanced grayscale difference and more sensitive texture response. A variable-scale bilateral filtering algorithm with adaptive capability is used to optimize the image. By performing local gradient analysis on the contrast-enhanced image, the brightness change values ​​in the horizontal and vertical directions are calculated pixel by pixel to form a local gradient map; then the filter kernel size and weight distribution are dynamically adjusted according to the gradient value at each pixel location, where the spatial domain filter kernel size is positively correlated with the gradient intensity to ensure that more detailed kernel operations are used in areas with complex textures, and the value domain weight parameter is negatively correlated with the gradient, that is, when the local changes are drastic, the weight range is reduced to avoid blurred boundaries, and when the local gradient is smooth, the fusion range is increased to improve the noise reduction effect. Through this step, the kernel function parameters of the bilateral filter are adaptively set according to the structural characteristics of the image itself, so that the image can smooth texture interference and random noise while maintaining the clarity of the edge contour, effectively improving the overall signal-to-noise ratio of the image. Through the above steps, a preprocessed image set is finally generated.

[0023] In a specific embodiment, the process of executing step S102 may specifically include the following steps: Input the preprocessed image into the first depthwise separable convolutional layer for processing to obtain the first-layer texture features; input the first-layer texture features into the second depthwise separable convolutional layer for processing to obtain the second-layer texture features; input the second-layer texture features into the third depthwise separable convolutional layer for processing to obtain the third-layer texture features; input the third-layer texture features into the fourth depthwise separable convolutional layer for processing to obtain the fourth-layer texture features; Input the first-layer texture features, the second-layer texture features, the third-layer texture features, and the fourth-layer texture features in parallel into four dilated convolutional layers for processing to obtain a multi-scale texture feature group; Calculate the first gradient difference in the horizontal direction and the second gradient difference in the vertical direction of the fourth-layer texture features respectively, and perform feature enhancement on the fourth-layer texture features according to the first gradient difference and the second gradient difference to obtain bidirectional gradient features; Perform weighted fusion on the multi-scale texture feature group, connect it with the bidirectional gradient features in the channel dimension, and then perform integration to obtain texture feature data.

[0024] Specifically, the preprocessed image is input into the first depthwise separable convolutional layer, which is composed of a depthwise convolution operation in the spatial dimension and a pointwise convolution across channels. The depthwise convolution extracts local patterns along the spatial range within each channel, while the pointwise convolution linearly fuses the features between channels. After passing through the non-linear activation function PReLU and batch normalization, the first-layer texture features are output. These features mainly represent the rough texture structure and edge responses on the component surface and belong to shallow perception representations. The first-layer texture features are input into the second depthwise separable convolutional layer, which expands the receptive field while retaining the local texture of the previous layer. By using a 7×7 or 5×5 depthwise convolution kernel, it enhances the ability to model the texture continuity between adjacent regions and extracts middle-layer texture features containing microscopic repetitive textures and weak crack edges, forming the second feature map. The second feature map is passed to the third depthwise separable convolutional structure, where the convolution kernel size is reduced to 3×3 to focus on finer-grained structural edges, and at the same time, the network depth is increased to enable it to represent complex texture overlaps such as intersecting cracks, dense pores, and blurred peeling boundaries. The output third-layer feature map has strong structural understanding ability and mid-term semantic perception ability of the image. The third-layer features are input into the fourth depthwise separable convolutional layer, which emphasizes feature integration more than the previous layer, extracts high-order semantic texture features, and completes the deep-mode abstraction of the complex defect morphology on the component surface, obtaining the fourth-layer texture feature map. The first to fourth-layer texture feature maps are respectively input into four dilated convolutional (atrous convolutional) layers with different dilation rates for parallel processing. The kernel sizes of these dilated convolutions are uniformly set to 3×3, and the dilation rates are respectively set to 1, 2, 4, and 8, so that the receptive field range extends from the local neighborhood to distant pixels, realizing the expansion of the perception range without increasing the number of convolution parameters, and effectively capturing the structural texture features of different types of defects such as punctate pores, linear cracks, and flaky peeling, obtaining a multi-scale texture feature group. The high-level texture feature maps are structurally enhanced, especially for the fourth-layer texture feature map, to further enhance its edge recognition ability while maintaining its semantic integrity. The gray-scale differences in the horizontal and vertical directions of the feature map are calculated respectively, that is, the absolute differences between each pixel position and its left and right adjacent pixels are obtained to get the first gradient response, and then the differences between the upper and lower adjacent pixels are obtained to get the second gradient response. The gradient differences in these two directions respectively represent the structural change intensities in the horizontal and vertical directions of the image. A lightweight convolution operation is used to perform edge mapping transformation on the gradient maps in both directions and connect them in the channel dimension to form a bidirectional gradient feature map, thereby explicitly enhancing the expression ability of the potential edge changes and structural fracture regions on the component surface, restoring and enhancing the edge morphology that is easily weakened by the smoothing operation in the deep convolution, obtaining a multi-scale texture feature group composed of the outputs of dilated convolutions covering multiple layers of semantics and a group of bidirectional gradient feature maps with clear boundary direction responses. In order to form a unified and strongly expressive texture representation, these features are fused and integrated.A channel weighting mechanism is introduced for the multi-scale texture feature group, and learnable weight parameters are introduced on each channel. Automatic weighted summation processing is performed according to the contribution degree of each scale feature to the final task to form a group of main texture feature maps. The fused texture feature map is concatenated with the previously constructed bidirectional gradient feature map along the channel dimension, so that the texture feature and the gradient structure are in a fused state of simultaneous cooperation in the feature tensor structure. A 1×1 convolution is introduced in the last texture integration module to complete feature compression and channel interaction, and texture feature data is output.

[0025] In a specific embodiment, the process of executing step S103 may specifically include the following steps: The preprocessed image is input into the MobileNetV3 backbone network for processing to obtain initial spatial features; The initial spatial features and the texture feature data are connected in the channel dimension and feature fusion is performed through a 1×1 convolutional layer to obtain target fusion features; Parallel convolution processing is performed on the target fusion features to obtain multi-scale spatial features, and the target fusion features are processed through global average pooling and 1×1 convolution to obtain global context features; A spatial attention map is calculated based on the target fusion features, and an element-wise multiplication operation is performed on the spatial attention map and the target fusion features to obtain spatially weighted features; The multi-scale spatial features, the global context features, and the spatially weighted features are connected in the channel dimension and integrated through a convolutional layer to obtain spatial relationship feature data.

[0026] Specifically, the preprocessed image is input into the backbone network structure of the lightweight neural network architecture MobileNetV3. Through the depthwise separable convolution module, asymmetric channel attention mechanism, and bottleneck residual structure inside the network, the global spatial layout and geometric construction pattern of the component surface are extracted on the premise of ensuring high efficiency, and an initial spatial feature map is obtained. This feature map retains the structural semantic information of the preprocessed image and covers spatial patterns such as edge orientation, region size, position distribution, and texture aggregation regions. The initial spatial features and texture feature data are concatenated tensorially in the channel dimension to construct a multimodal fusion feature map containing spatial distribution patterns and texture semantic structures. This concatenation operation unifies the features from different perceptual paths into the same representation space and completes the correlation mapping of the two types of features at the spatial distribution level through structural alignment. To achieve information compression and channel interaction, a 1×1 convolutional layer is used to perform convolution processing on the concatenated fusion features, and the output target fusion feature map is regulated through linear mapping and activation functions in the channel dimension. To capture the performance characteristics of component surface defects at different scales, a parallel multi-scale convolutional path structure is designed to process the target fusion features. Each path uses dilated convolutional kernels with different dilation rates (such as 6, 12, 18, 24) to expand the receptive field on the premise of keeping the number of convolutional parameters unchanged. These parallel paths extract spatial features from four levels: local neighborhood, small region blocks, large structural contours, and even the overall component layout, generating a group of multi-scale spatial feature maps with different perceptual levels; the feature maps output by the convolutional paths can capture the context scales required for different types of defects, such as dense pore regions, small crack clusters, large-scale spalling blocks, and geometric structure fractures, and reflect their semantic levels in the form of feature channels. While performing parallel convolutions, a global average pooling operation is applied to the target fusion feature map to compress the spatial dimension and extract the overall statistical semantics of the image, and then the channel dimension is restored through 1×1 convolution processing to form a global context feature map. This feature map serves as the context prior in the full-component scenario and guides the subsequent local feature judgment by combining structural information such as the overall component morphology and common defect positions. After feature fusion, a spatial attention mechanism module is introduced. Through this mechanism, the spatial positions that contribute most to the defect recognition task in the current image are automatically identified and highlighted. The specific operation is to perform maximum pooling and average pooling operations on the target fusion feature map respectively to retain different statistical features, and then the two are concatenated along the channel dimension and input into a standard convolutional layer with a receptive field of 7×7. The output two-dimensional spatial attention map is normalized through the sigmoid function. Each pixel value in this attention map represents the importance score of the corresponding spatial position. The map is multiplied pixel by pixel with the original target fusion feature map, that is, the channel value at each position is multiplied by the attention weight at the corresponding position, to enhance the significant regions and suppress the non-critical regions, obtaining a spatially weighted feature map.The multi-scale spatial features, global context features, and spatial weighted features are concatenated and integrated in the channel dimension, compressing features at multiple semantic levels into a unified feature space. To avoid the problems of feature redundancy and channel dimension mismatch caused by direct splicing, a 3×3 convolutional layer is introduced after integration to perform convolutional fusion. This layer jointly models the multi-source features and enhances the feature interaction ability through a non-linear activation function, enabling the fused feature map to have higher discriminative expressiveness, and finally obtaining the spatial relationship feature data.

[0027] In a specific embodiment, the process of performing step S104 may specifically include the following steps: The texture feature data and the spatial relationship feature data are respectively processed through global average pooling and a multi-layer perceptron to obtain the texture feature importance score and the spatial feature importance score; Based on the texture feature importance score and the spatial feature importance score, a weighting operation is performed on the texture feature data and the spatial relationship feature data to obtain the initial weighted fusion feature; Taking the texture feature data as the query item and the spatial relationship feature data as the key-value pair, calculating the similarity matrix between the features and performing normalization processing to obtain the texture-to-space attention map; Taking the spatial relationship feature data as the query item and the texture feature data as the key-value pair, calculating the similarity matrix between the features and performing normalization processing to obtain the space-to-texture attention map; Based on the texture-to-space attention map, weighted summation is performed on the spatial relationship feature data, and at the same time, based on the space-to-texture attention map, weighted summation is performed on the texture feature data to obtain the bidirectional cross-attention feature; The initial weighted fusion feature and the bidirectional cross-attention feature are connected in the channel dimension, and at the same time, autoclaved aerated concrete material constraints are introduced to obtain the comprehensive defect feature representation.

[0028] Specifically, the texture feature data and the spatial relationship feature data are input into the global average pooling module, and the mean values of the pixel values in the spatial dimension are calculated for each channel respectively, compressing the original three-dimensional feature map (height × width × channel) into a one-dimensional channel vector. This vector represents the global response intensity of each channel on the entire image and reflects the global sensitivity of the features. Subsequently, this one-dimensional vector is input into a non-linear mapping network composed of multi-layer perceptrons. After being processed by two linear transformation layers and activation functions, a normalized importance vector is output. This vector is mapped to between 0 and 1 through the sigmoid function, corresponding to the texture feature importance score and the spatial feature importance score respectively, representing the relative importance of this type of feature for the recognition task in the current component sample. The original texture feature data and spatial relationship feature data are subjected to channel weighting operations using the above two importance score vectors. Multiply the original features of each channel by the corresponding importance score to highlight the high-weight channels and suppress the low-response channels. The weighted texture features and spatial features are concatenated and integrated in the channel dimension to form an initial weighted fusion feature. On this basis, to establish a cross-semantic interaction channel between the texture and spatial features, a bidirectional cross-attention mechanism is designed to achieve explicit relationship modeling and mutual guidance enhancement between modalities. This mechanism takes the texture features as query terms, and the spatial features as keys and values to form a triple. The similarity relationship between each texture channel position and spatial channel position is calculated through tensor dot product, constructing a high-dimensional feature similarity matrix, and this matrix is normalized through the Softmax function to form an attention map from texture to space. Each element in the map represents the dependence degree or attention weight of a certain texture feature position on a certain position of the spatial feature. The original spatial feature map is subjected to a weighted summation operation using this attention map, that is, all spatial channels are recombined under attention weighting to output a texture-guided spatial reconstruction feature map. The reverse construction path takes the spatial features as query terms and the texture features as key-value pairs, and again constructs a similarity matrix from space to texture through the dot product method, and obtains an attention map from space to texture through normalization. This attention map is used for attention-weighted fusion of the texture feature map to generate a space-guided texture enhancement map. Through the dual-channel attention mechanism, the context relationship between the two modalities is effectively captured, especially showing higher discriminative ability in fuzzy regions such as the unclear boundary between the crack boundary and the background, and the fusion of air holes and surface gray patterns, obtaining a bidirectional cross-attention feature representation. The initial weighted fusion feature and the bidirectional cross-attention feature are connected in the channel dimension, and a deeper semantic expression space with a higher dimension is constructed through stacking. A material constraint mechanism is introduced in the fusion stage, embedding the material property prior in the form of a loss function into the fusion strategy.These include the continuity constraint of the high aspect ratio of cracks, the circular boundary and the closure constraint of pores, the irregular shape of the peeling area and the internal uniform gray-scale distribution constraint, etc. The boundary shape, gray-scale gradient, area distribution, etc. of the result features are adjusted in the fusion network through the morphological loss function respectively, and the comprehensive defect feature representation is obtained after the fusion convolution compression and normalization processing.

[0029] In a specific embodiment, the process of executing step S105 may specifically include the following steps: Restore the feature map size of the comprehensive defect feature representation to obtain a defect binary map; Extract the region feature vector for each connected region in the defect binary map, and perform multi-class defect discrimination on the region feature vector to obtain the surface defect category result; Map the surface defects detected from each perspective to the three-dimensional coordinate system of the autoclaved aerated concrete member based on the second set of spatial transformation matrices, and fuse the defects in the overlapping regions of adjacent perspectives through the intersection over union and confidence to obtain an all-round defect distribution map; Perform area statistical analysis on the defect binary map to obtain the surface defect rate index; Calculate the surface quality index based on the surface defect category result and the defect area, and correct it by combining the influence coefficients of the cement ratio, forming pressure, and fine aggregate content to obtain the weighted surface quality index; Calculate the regional defect density uniformity of the all-round defect distribution map to obtain the defect distribution uniformity index, and calculate the comprehensive surface quality score based on the surface defect rate index, weighted surface quality index, and defect distribution uniformity index, and generate the corresponding grade division result according to the comprehensive surface quality score.

[0030] Specifically, perform size reduction processing on the comprehensive defect feature representation in the spatial dimension to restore it from the downsampled state to the resolution of the original input image. In the implementation process, use a transposed convolutional structure layer by layer to upsample the feature map. Each layer of transposed convolution uses a setting with a stride of 2 and a padding of 1. At the same time, nest the BatchNorm and ReLU activation functions to enhance the non-linear expression and training stability. After three consecutive rounds of upsampling, the feature map is restored to the original size of the preprocessed image. And introduce a 1×1 convolution and a Sigmoid normalization activation function on the output channel to output the confidence value of each pixel point in the defect prediction map, that is, generate a defect probability map. Subsequently, apply the OTSU adaptive threshold algorithm to perform pixel-level binary segmentation on this probability map, determine the positions with confidence higher than the threshold as defect pixel points, and construct a defect binary map. Divide the defect binary map into several connected regions, and extract region feature vectors representing their geometric shapes and texture information for each connected region respectively. The extraction content includes multiple dimensions such as area, aspect ratio, perimeter, edge complexity, average gray value, edge gradient statistics, and direction consistency. These region features are used as classification features to input a pre-trained multi-class defect classifier. The classifier uses a shallow convolutional structure or an MLP classification network internally, and outputs the defect type probability distribution of each region through the Softmax layer, and takes the maximum probability term as the final classification result, outputting the surface defect category results including labels such as "crack", "porosity", "spalling", and "impurity". Based on the second set of spatial transformation matrices, uniformly map the defect regions detected from all viewpoints to the three-dimensional coordinate system of the component body. The mapping process performs back-projection transformation from the image pixel space to the physical space through the known internal parameter matrix and external parameter matrix to ensure spatial alignment and consistency between multiple viewpoints. Since there is a field of view overlap among multiple cameras, to avoid the same defect region being counted multiple times, use the intersection over union (IoU) as the region coincidence degree judgment standard, and set the IoU threshold to 0.3. When the intersection over union of defect regions from two different viewpoints exceeds this threshold, fuse the decisions by referring to their confidence values at the same time, retain the region with higher confidence or perform weighted merging on it, to achieve non-redundant integration of multi-view defect information and form an all-round defect distribution map. Perform pixel-level area statistics on the defect binary map, calculate the ratio of the total area of all defect regions to the area of the entire component surface image, and output the surface defect rate index, which serves as the most intuitive basis for surface quality measurement. At the same time, combine the defect category information obtained above, separately count the area occupied by different types of defects and assign preset weight coefficients, and calculate the surface quality index through weighted summation. This index reflects the severity of the impact of defects on the material properties. Introduce correction factors, including three key parameters: cement ratio, forming pressure, and fine aggregate content.By constructing an empirical model, the influence factors of cement ratio, forming pressure, and fine aggregate are solved respectively, and these factors are introduced into the correction formula of the surface quality index to calculate a weighted surface quality index that is more relevant to actual production. Perform regional defect density uniformity analysis on the all-round defect distribution map, divide the component surface into multiple spatial sub-blocks, count the defect density of each sub-block, then calculate the mean and standard deviation of the defect density of all sub-blocks, and use the standard deviation normalization index to represent the degree of uniformity. The higher this index, the more concentrated or extreme the defect distribution, and the lower it is, the more evenly the defects are dispersed, reflecting the surface structure consistency and local processing stability of the component. Calculate the comprehensive surface quality score based on the surface defect rate index, weighted surface quality index, and defect distribution uniformity index, define the quality grade according to the score value range, and finally output the score result and grade division, and bind them to the production identification or quality label of the component.

[0031] The intelligent identification method for surface defects of autoclaved aerated concrete components in the above embodiments of the present invention has been described. Next, the intelligent identification system for surface defects of autoclaved aerated concrete components in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the intelligent identification system for surface defects of autoclaved aerated concrete components in the embodiments of the present invention includes: A mapping module 201, configured to perform all-round shooting and spatial transformation matrix mapping on the surface of the autoclaved aerated concrete component to obtain a preprocessed image; A feature extraction module 202, configured to extract texture features from the preprocessed image to obtain texture feature data; A feature analysis module 203, configured to perform spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data; A cross-fusion module 204, configured to perform feature importance weighting and two-way cross-fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation; An output module 205, configured to perform surface defect analysis on the autoclaved aerated concrete component based on the comprehensive defect feature representation, and output the comprehensive surface quality score and the grade division result.

[0032] Through the collaborative cooperation of the above-mentioned various components, through the multi-view camera array and spatial transformation matrix optimization technology, the full coverage shooting of each surface of autoclaved aerated concrete components is achieved, solving the dead angle problem and image deformation problem existing in traditional single-view detection, and ensuring the integrity and accuracy of defect detection. Aiming at the characteristics of the grayish-white color and low contrast of the autoclaved aerated concrete surface, technologies such as illumination equalization, adaptive contrast enhancement, and variable-scale bilateral filtering are adopted, significantly improving the visibility and edge sharpness of defects, and laying a foundation for subsequent feature extraction. The depthwise separable convolution and multi-scale texture coding structure are adopted, which are specifically optimized for the texture characteristics of different types of defects such as pores, cracks, and depressions on the autoclaved aerated concrete surface. While greatly reducing the number of parameters and computational complexity, it maintains a high feature extraction ability. Through the combination of deformable convolution, multi-scale spatial pooling, and attention mechanism, the effective detection of different-scale defects from tiny pores to large-area peeling is realized, solving the problem that traditional methods are difficult to adapt to the recognition of multi-scale defects simultaneously. Through feature importance weighting and bidirectional cross-attention mechanism, the complementary advantages of texture features and spatial features are fully integrated, and the knowledge constraints of autoclaved aerated concrete material science are introduced, improving the accuracy and robustness of defect recognition, especially the adaptability under complex environmental conditions. A multi-level quality assessment system integrating defect classification, location, and process parameters is constructed, and a quantitative relationship model between defect formation and process parameters such as cement ratio, forming pressure, and fine aggregate content is established, providing a scientific basis for production process optimization and quality control, and realizing the transformation from simple defect detection to comprehensive quality assessment and process guidance. Through model lightweight design and structural optimization, the system can run in real time on the edge devices of the production line, meeting the real-time detection requirements in the industrial production environment, and providing a reliable real-time quality monitoring means for the production of autoclaved aerated concrete components.

[0033] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0034] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an intelligent identification device for surface defects of autoclaved aerated concrete components (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0035] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. An intelligent method for identifying surface defects of autoclaved aerated concrete components, characterized in that: include: The surface of the autoclaved aerated concrete component is photographed in all directions and the spatial transformation matrix is ​​mapped to obtain a pre-processed image; Extracting texture features from the preprocessed image to obtain texture feature data; Performing spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data; Performing feature importance weighting and bidirectional cross-fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation; Based on the comprehensive defect feature representation, surface defect analysis of the autoclaved aerated concrete component is performed, and a comprehensive surface quality score and grade classification result are output.

2. The method for intelligently identifying surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The process of performing all-round photographing and space transformation matrix mapping on the surface of the autoclaved aerated concrete component to obtain a pre-processed image includes: The main view camera is installed just above the conveyor belt of the autoclaved aerated concrete component, and N auxiliary view cameras are distributed in a ring shape around the autoclaved aerated concrete component to obtain a multi-view camera array; The multi-view camera array is started to synchronously shoot the autoclaved aerated concrete component to obtain an original image set, and each camera in the multi-view camera array is individually calibrated with internal parameters to obtain the focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient of each camera; According to the focal length, principal point coordinates, radial distortion coefficient and tangential distortion coefficient of each camera, the conversion relationship between the component three-dimensional coordinate system and each camera coordinate system is established to obtain the external parameter matrix of each camera; Based on the positional relationship of the matching feature points of the images of each viewing angle in the original image set, the homography matrix between adjacent cameras is calculated to obtain a first spatial transformation matrix set; Based on the extrinsic matrix of each camera, the first spatial transformation matrix set is optimized by minimizing the reprojection error to obtain a second spatial transformation matrix set, and the original image set is mapped and transformed by the second spatial transformation matrix set to obtain a surface full coverage correction image; The surface full coverage correction image is subjected to illumination equalization, contrast enhancement and noise reduction processing to obtain a preprocessed image.

3. The method for intelligently identifying surface defects of autoclaved aerated concrete components according to claim 2, characterized in that: The step of performing illumination equalization, contrast enhancement and noise reduction on the surface full coverage correction image to obtain a preprocessed image includes: According to the focal length, principal point coordinates, radial distortion coefficient and tangential distortion coefficient of each camera, the surface full coverage correction image is subjected to distortion correction to obtain a standard component surface image; Performing Gaussian filtering on the surface image of the standard component to obtain an image representing background brightness distribution; Calculating a correction coefficient matrix according to the image representing the background brightness distribution, and applying the correction coefficient matrix to the standard component surface image to obtain a lighting equalization image; Performing histogram statistical analysis on the illumination equalization image, calculating a cumulative distribution function, and mapping the cumulative distribution function to a preset target distribution function to obtain a contrast enhanced image; Calculating the local gradient value of the contrast enhanced image, and adaptively adjusting the filter kernel size and range parameters of the variable scale bilateral filter according to the local gradient value to obtain target filter parameters; The contrast enhanced image is subjected to variable scale bilateral filtering based on the target filtering parameters to obtain a preprocessed image.

4. The method for intelligently identifying surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The step of extracting texture features from the preprocessed image to obtain texture feature data comprises: Input the preprocessed image into a first depth-separable convolutional layer for processing to obtain a first layer of texture features; input the first layer of texture features into a second depth-separable convolutional layer for processing to obtain a second layer of texture features; input the second layer of texture features into a third depth-separable convolutional layer for processing to obtain a third layer of texture features; input the third layer of texture features into a fourth depth-separable convolutional layer for processing to obtain a fourth layer of texture features; Inputting the first layer of texture features, the second layer of texture features, the third layer of texture features and the fourth layer of texture features into four dilated convolutional layers in parallel for processing to obtain a multi-scale texture feature group; Respectively calculating a first gradient difference in a horizontal direction and a second gradient difference in a vertical direction of the fourth layer texture feature, and performing feature enhancement on the fourth layer texture feature according to the first gradient difference and the second gradient difference to obtain a bidirectional gradient feature; The multi-scale texture feature group is weightedly fused and connected with the bidirectional gradient feature in the channel dimension and then integrated to obtain texture feature data.

5. The method for intelligently identifying surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The performing spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data includes: Input the preprocessed image into the MobileNetV3 backbone network for processing to obtain initial spatial features; Connecting the initial spatial features with the texture feature data in the channel dimension, and performing feature fusion through a 1×1 convolution layer to obtain a target fusion feature; Performing parallel convolution processing on the target fusion features to obtain multi-scale spatial features, and performing global average pooling and 1×1 convolution processing on the target fusion features to obtain global context features; Calculating a spatial attention map based on the target fusion feature, and performing element-wise multiplication operation on the spatial attention map and the target fusion feature to obtain a spatial weighted feature; The multi-scale spatial features, the global context features and the spatial weighted features are connected in the channel dimension and integrated through a convolutional layer to obtain spatial relationship feature data.

6. The method for intelligently identifying surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The step of performing feature importance weighting and bidirectional cross-fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation includes: The texture feature data and the spatial relationship feature data are processed respectively by global average pooling and a multi-layer perceptron to obtain a texture feature importance score and a spatial feature importance score; Based on the texture feature importance score and the spatial feature importance score, weighting operation is performed on the texture feature data and the spatial relationship feature data to obtain an initial weighted fusion feature; The texture feature data is used as a query item and the spatial relationship feature data is used as a key-value pair, a similarity matrix between features is calculated and normalized, and a texture-to-space attention map is obtained; The spatial relationship feature data is used as a query item and the texture feature data is used as a key-value pair, a similarity matrix between features is calculated and normalized, and a space-to-texture attention map is obtained; Performing weighted summation on the spatial relationship feature data based on the texture-to-space attention map, and performing weighted summation on the texture feature data based on the space-to-texture attention map to obtain a bidirectional cross-attention feature; The initial weighted fusion feature is connected with the bidirectional cross attention feature in the channel dimension, and the autoclaved aerated concrete material constraint is introduced to obtain a comprehensive defect feature representation.

7. The method for intelligently identifying surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The surface defect analysis of the autoclaved aerated concrete component is performed based on the comprehensive defect feature representation, and the comprehensive surface quality score and grade classification results are output, including: Performing feature map size restoration on the comprehensive defect feature representation to obtain a defect binary map; Extracting a regional feature vector from each connected region in the defect binary map, and performing multi-category defect discrimination on the regional feature vector to obtain a surface defect category result; Based on the second spatial transformation matrix set, the surface defects of the autoclaved aerated concrete components detected at each viewing angle are mapped to the component-component three-dimensional coordinate system, and the defects of the overlapping areas of adjacent viewing angles are fused by intersection-over-union ratio and confidence to obtain a full-range defect distribution map; Performing area statistical analysis on the defect binary image to obtain a surface defect rate index; Calculating the surface quality index based on the surface defect category results and defect area, and correcting it in combination with the influence coefficient of cement ratio, molding pressure and fine aggregate content to obtain a weighted surface quality index; The regional defect density uniformity is calculated for the omnidirectional defect distribution map to obtain a defect distribution uniformity index, and a surface quality comprehensive score is calculated based on the surface defect rate index, the weighted surface quality index, and the defect distribution uniformity index, and a corresponding grade classification result is generated according to the surface quality comprehensive score.

8. An intelligent recognition system for surface defects of autoclaved aerated concrete components, characterized in that: Used to implement the method for intelligently identifying surface defects of autoclaved aerated concrete components according to any one of claims 1 to 7, the intelligent system for intelligently identifying surface defects of autoclaved aerated concrete components comprises: A mapping module is used to perform all-round photography and spatial transformation matrix mapping of the surface of the autoclaved aerated concrete component to obtain a pre-processed image; A feature extraction module, used for extracting texture features from the preprocessed image to obtain texture feature data; A feature analysis module, used for performing spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data; A cross fusion module, used for performing feature importance weighting and bidirectional cross fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation; The output module is used to perform surface defect analysis on the autoclaved aerated concrete component based on the comprehensive defect feature representation, and output a comprehensive surface quality score and grade classification result.

Citation Information

Patent Citations

  • Texture surface defect detection method and system

    CN110969606A

  • Lead screw module surface defect online detection method and device based on machine vision

    CN119090862A

  • Detection method, system and equipment for intelligent welding of pipe pile splicing and medium

    CN119559382A

  • Concrete surface intelligent construction method and system based on multi-modal sensing

    CN119784326A

  • Surface defect inspection method and surface defect inspection apparatus

    US20200025690A1

Cited By

  • Method for detecting porosity of autoclaved aerated concrete block

    CN120563500A

  • Paper roll raised strip detection system and method

    CN121259305A

  • Concrete member surface defect detection method and system based on image segmentation

    CN121280435A

  • Machine vision-based automatic identification method for assembly quality of assembly type station component

    CN121366386A

  • Method for automatically identifying assembly quality of prefabricated station components based on machine vision

    CN121366386B