Intelligent identification method and system for surface defects of autoclaved aerated concrete components
Through multi-view camera array and spatial transformation matrix optimization technology, combined with illumination equalization and depth-separable convolution, the blind spot and image deformation problems in the surface inspection of autoclaved aerated concrete components are solved, full coverage inspection and quality assessment are achieved, and the real-time inspection needs of industrial production are met.
Patent Information
- Application Number
- CN202510642916.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-19
AI Technical Summary
Traditional surface defect detection for autoclaved aerated concrete components suffers from low efficiency, inconsistent standards, inability to conduct quantitative evaluation, and difficulty correlating with production process parameters. Furthermore, traditional single-view detection methods suffer from blind spots and image deformation, making it impossible to perform full-scale surface inspection of components without blind spots.
By adopting multi-view camera array and spatial transformation matrix optimization technology, combined with illumination equalization, adaptive contrast enhancement and variable-scale bilateral filtering technology, through deep separable convolution and multi-scale texture coding structure, combined with deformable convolution, multi-scale spatial pooling and attention mechanism, full coverage shooting and feature extraction of the surface of autoclaved aerated concrete components are achieved, and a quantitative relationship model between defect formation and process parameters is constructed.
It achieves full coverage detection of the surface of autoclaved aerated concrete components, improves the integrity and accuracy of defect detection, can monitor quality in real time under industrial production environments, and provide a scientific basis for production process optimization.
Smart Images

Figure CN120182253B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of surface defect recognition, and in particular to a method and system for intelligently recognizing surface defects of autoclaved aerated concrete components. Background Art
[0002] As a lightweight and high-strength building material, autoclaved aerated concrete is widely used in the construction industry due to its excellent thermal insulation, fireproof properties, and ease of construction. However, during the production process, due to fluctuations in process parameters such as cement ratio, molding pressure, fine aggregate content, curing time, and autoclaving temperature, various types of defects such as pores, cracks, and dents often appear on the surface of components. These defects not only affect the appearance quality of the components, but also have a significant impact on their structural performance. In severe cases, they can lead to reduced strength, decreased durability, and even structural failure during use. Traditional surface defect detection of autoclaved aerated concrete components mostly relies on manual visual inspection, which has problems such as low efficiency, inconsistent standards, inability to quantitatively evaluate, and difficulty in correlating with production process parameters. It is difficult to meet the quality control needs of modern large-scale production lines.
[0003] Existing automated visual inspection technologies face numerous challenges when applied to surface defect identification in autoclaved aerated concrete components. The off-white, low-contrast surface of autoclaved aerated concrete components makes it difficult to distinguish defects from the background, making it difficult for conventional image processing methods to effectively extract defect features. Furthermore, surface defects in autoclaved aerated concrete are complex, morphologically diverse, and vary in size, ranging from tiny pores to large-scale detachments, making them difficult to fully capture using a single feature extraction method. Third, traditional single-view inspection methods suffer from blind spots and image distortion, making it impossible to inspect all surfaces without blind spots. Fourth, the lack of a systematic method to correlate defect formation with production process parameters makes it difficult to provide quantitative guidance for production process optimization. Summary of the Invention
[0004] The present invention provides a method and system for intelligently identifying surface defects of autoclaved aerated concrete components, which solves the blind spot and image deformation problems existing in traditional single-view detection and ensures the integrity and accuracy of defect detection.
[0005] In a first aspect, the present invention provides a method for intelligently identifying surface defects of autoclaved aerated concrete components, the method comprising:
[0006] The surface of the autoclaved aerated concrete component is photographed in all directions and mapped by spatial transformation matrix to obtain a pre-processed image;
[0007] Extracting texture features from the preprocessed image to obtain texture feature data;
[0008] Performing spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data;
[0009] Performing feature importance weighting and bidirectional cross-fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation;
[0010] Based on the comprehensive defect feature representation, surface defect analysis of the autoclaved aerated concrete component is performed, and a comprehensive surface quality score and grade classification result are output.
[0011] In a second aspect, the present invention provides an intelligent system for identifying surface defects of autoclaved aerated concrete components, the intelligent system for identifying surface defects of autoclaved aerated concrete components comprising:
[0012] A mapping module is used to perform omnidirectional photography and spatial transformation matrix mapping of the surface of the autoclaved aerated concrete component to obtain a pre-processed image;
[0013] A feature extraction module is used to extract texture features from the preprocessed image to obtain texture feature data;
[0014] A feature analysis module, configured to perform spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data;
[0015] A cross fusion module, configured to perform feature importance weighting and bidirectional cross fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation;
[0016] The output module is used to perform surface defect analysis on the autoclaved aerated concrete component based on the comprehensive defect feature representation, and output a comprehensive surface quality score and grade classification result.
[0017] In the technical solution provided by the present invention, full coverage shooting of all surfaces of autoclaved aerated concrete components is achieved through multi-view camera arrays and spatial transformation matrix optimization technology, solving the blind spot and image deformation problems existing in traditional single-view detection, and ensuring the integrity and accuracy of defect detection. In view of the grayish-white and low-contrast characteristics of the autoclaved aerated concrete surface, illumination equalization, adaptive contrast enhancement and variable-scale bilateral filtering technologies are used to significantly improve the visibility and edge clarity of defects. Deep separable convolution and multi-scale texture coding structures are used to optimize the texture characteristics of different types of defects such as pores, cracks, and depressions on the surface of autoclaved aerated concrete, while significantly reducing the number of parameters and computational complexity while maintaining a high feature extraction capability. Through the combination of deformable convolution, multi-scale spatial pooling and attention mechanism, effective detection of defects of different scales, from tiny pores to large-scale detachment, is achieved, solving the problem that traditional methods are difficult to adapt to multi-scale defect recognition at the same time. By using feature importance weighting and a bidirectional cross-attention mechanism, the complementary advantages of texture and spatial features are fully integrated, and constraints based on autoclaved aerated concrete material science knowledge are introduced to improve the accuracy and robustness of defect identification, especially its adaptability under complex environmental conditions. A multi-level quality assessment system integrating defect classification, location, and process parameters was constructed, and a quantitative relationship model between defect formation and process parameters such as cement ratio, molding pressure, and fine aggregate content was established. This provides a scientific basis for production process optimization and quality control, and realizes the transition from simple defect detection to comprehensive quality assessment and process guidance. Through lightweight model design and structural optimization, the system can run in real time on edge equipment on the production line, meeting the real-time detection needs in industrial production environments and providing a reliable real-time quality monitoring method for the production of autoclaved aerated concrete components. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 A schematic diagram of an embodiment of an intelligent method for identifying surface defects of autoclaved aerated concrete components according to an embodiment of the present invention;
[0020] Figure 2 Schematic diagram of an embodiment of an intelligent system for identifying surface defects of autoclaved aerated concrete components according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] Embodiments of the present invention provide a method and system for intelligently identifying surface defects in autoclaved aerated concrete components. The terms "first," "second," "third," "fourth," and so forth (if any) in the specification and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions; for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units expressly listed, but may include other steps or units not expressly listed or inherent to such process, method, product, or apparatus.
[0022] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 An embodiment of the method for intelligently identifying surface defects of autoclaved aerated concrete components according to the present invention includes:
[0023] Step S101: Perform omnidirectional photography and spatial transformation matrix mapping on the surface of the autoclaved aerated concrete component to obtain a pre-processed image;
[0024] It is understood that the execution subject of the present invention can be an intelligent recognition system for surface defects of autoclaved aerated concrete components, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0025] Specifically, when constructing the visual perception platform, a primary-view camera is mounted vertically above the conveyor belt of autoclaved aerated concrete components, capturing the entire top surface of the components. Simultaneously, N auxiliary-view cameras are evenly distributed around the components, forming a circular pattern at equal angles. These cameras utilize high-resolution, industrial-grade CCD devices to ensure comprehensive coverage of the component surface. The multi-view camera array forms the system's core image acquisition architecture. A high-speed communication link is established with the central processing unit to enable synchronized triggering of all cameras and parallel image data acquisition. After synchronized image capture, a raw image set containing component images from multiple different perspectives is obtained. Each camera undergoes independent intrinsic calibration, using a standard checkerboard calibration grid to determine key camera parameters such as focal length, principal point position, radial distortion coefficient, and tangential distortion coefficient. Based on these intrinsic parameters, a spatial geometric relationship is established between each camera and a unified world coordinate system. The spatial position and pose of each camera are described by calculating its extrinsic parameter matrix, enabling a transformation mapping between the component's three-dimensional coordinate system and the coordinate system of each camera. Based on the shared matching feature points between adjacent viewpoints in the original image set, an image registration algorithm is used to construct a set of inter-camera homography matrices, resulting in a first set of spatial transformation matrices. This set of matrices describes how images from adjacent viewpoints are projected onto the same geometric plane through perspective transformation. Leveraging the 3D geometric constraints constructed from the extrinsic parameters of each camera, an optimization strategy designed to minimize reprojection error is employed to iteratively solve the first set of spatial transformation matrices. Guided by a gradient descent algorithm, each set of transformation parameters is gradually adjusted until the error change over two consecutive rounds falls below a set threshold, converging to the optimal second set of spatial transformation matrices. Based on this second set of spatial transformation matrices, all original images are subjected to a mapping transformation, allowing images from different viewpoints to be fused and reconstructed within a unified spatial framework, forming a distortion-free, geometrically aligned, and fully covered corrected image set. Image enhancement processing, including illumination equalization, contrast enhancement, and denoising, is performed on each corrected image. In the illumination equalization stage, a brightness background model is constructed through Gaussian filtering, and the brightness is normalized based on the target average brightness value to solve the deviation between bright and dark areas caused by uneven light sources. Subsequently, adaptive histogram normalization technology is used to perform contrast enhancement to improve the problem of weak visual details caused by light material colors. A variable scale bilateral filter is used to achieve local area adaptive smoothing, which suppresses random noise interference while retaining edge details to obtain a preprocessed image.
[0026] In this embodiment, a geometric distortion repair operation is performed on the fully covered surface corrected image based on the internal parameters of each camera obtained during the calibration phase, including focal length, principal point coordinates, radial distortion coefficients, and tangential distortion coefficients. By introducing a distortion correction function and embedding these parameters into the imaging geometry model, spatial distortions caused by lens distortion are corrected pixel by pixel, resulting in a standard component surface image. Lighting equalization is then performed on the standard component surface image. Gaussian filtering is used to blur the image. This process applies a low-pass filter to the image in the spatial domain, effectively removing high-frequency details and retaining only the background brightness distribution trend, generating a background brightness reference image. Using this background image as a reference, the ratio between the target brightness and the current brightness value of each pixel is calculated to construct a two-dimensional correction coefficient matrix. This correction coefficient matrix is then multiplied pixel by pixel across the entire image to achieve normalized image brightness and output a lighting-balanced image. After achieving lighting equalization, histogram normalization is used to enhance the image contrast by enhancing the subtle contrast characteristics of the grayish-white surface material in the image. The illumination-balanced image undergoes pixel-level statistical analysis, calculating its grayscale histogram and constructing a cumulative distribution function based on the cumulative frequency of grayscale values. This function is then mapped to a pre-set target distribution function. Through an inverse transformation of the function, the original pixel grayscale distribution is adjusted to the target distribution, effectively extending the original image's grayscale dynamic range. This allows for more distinct grayscale distinction between surface indentations and cracks and the surrounding substrate, resulting in a contrast-enhanced image. Structure-preserving noise reduction is then performed. A variable-scale bilateral filtering mechanism is constructed, adaptively adjusting its parameters based on the local gradient characteristics of each image region. The gradient response is calculated at the pixel level to assess the edge complexity or texture variation of the current region. This gradient value is then used to dynamically adjust the bilateral filter's spatial kernel size and range parameters, controlling the spatial and grayscale weights of neighboring pixels, respectively. In edge regions, the filter reduces the spatial kernel size to enhance edge preservation, while in flat regions, the kernel range is expanded to enhance smoothing. A variable-scale bilateral filter with adaptive parameter configuration is applied to the contrast-enhanced image. While ensuring the complete preservation of the defect edge structure, the background noise introduced by uneven lighting, material reflection and equipment jitter in the image is effectively filtered out to obtain a preprocessed image.
[0027] Step S102: extracting texture features from the pre-processed image to obtain texture feature data;
[0028] Specifically, the preprocessed image is fed into the first depthwise separable convolutional layer, where initial texture feature information is extracted through deep convolution operations with larger kernel sizes, resulting in the first layer of texture features. This is then fed into the second depthwise separable convolutional layer, where finer edges and texture structures are extracted within a deeper receptive field, outputting the second layer of texture features. The second layer of features is then fed into the third depthwise separable convolutional layer, extracting higher-order texture semantic patterns to form the third layer of texture features. The final deep texture features are then extracted through the fourth depthwise separable convolution layer, forming the fourth layer of texture features with rich contextual information and multi-scale structural representation capabilities. The entire process constructs a hierarchical texture representation network through layer-by-layer stacking, gradually accumulating local information such as edges and corners at the lower level to higher-level texture combination patterns, thereby enhancing the model's ability to resolve complex surface defects on components. To improve the model's ability to resolve texture features of defects at different scales (such as microcracks, holes, and spalling areas) after obtaining these four levels of texture features, a parallel dilated convolution mechanism is introduced. These four layers of texture features are fed in parallel into four dilated convolutional networks, each configured with a different dilation rate. This expands the receptive field without increasing the number of parameters, thereby extracting information about the response of each layer at different spatial scales. This results in a set of scale-differentiated multi-scale texture features. Furthermore, to highlight the regions of the texture feature map most sensitive to edge changes, a bidirectional gradient enhancement mechanism is designed specifically for the fourth-layer texture feature map. This mechanism calculates the grayscale difference between the fourth-layer texture feature map and its adjacent pixels in the horizontal direction to obtain the first-direction gradient, and the grayscale change in the vertical direction to obtain the second-direction gradient. This mechanism explicitly extracts pixel-level edge mutation information through an absolute difference operation. The horizontal and vertical gradient feature maps are then fused through a set of lightweight convolutional networks to generate a bidirectional gradient feature map, enhancing the representation of microstructural features such as crack boundaries and edge transitions. The multi-scale texture feature map generated by the dilated convolutions is then subjected to a channel-weighted fusion. A channel-attention mechanism is introduced during the fusion process to assign different weights to feature maps at different scales, accommodating the varying scale sensitivity of different defect types. The fused texture feature map is then concatenated with the bidirectional gradient feature along the channel dimension to generate texture feature data.
[0029] Step S103: performing spatial feature analysis on the pre-processed image and texture feature data to obtain spatial relationship feature data;
[0030] Specifically, a spatial feature analysis mechanism is used to fuse and model the preprocessed image and texture feature data. This mechanism, based on a lightweight MobileNetV3 backbone network, extracts features from the preprocessed image through its deep convolutional structure, generating an initial spatial feature map. This map reflects low-level spatial information on the component surface, such as geometric distribution, texture density, illumination response, and edge contours in different regions. The initial spatial feature map is then fused with the texture feature data in the channel dimension, and a 1×1 convolution operation is introduced to compress and redistribute the concatenated multi-channel features, forming a target fused feature map with joint characteristics of spatial position and texture structure. To improve the model's sensitivity to defects of varying scales (such as microcracks, large-scale spalling, and irregularly distributed pores), a parallel convolutional processing path is applied to the target fused features. Each convolutional path is configured with a different receptive field and dilation rate to extract spatial semantic features at different scales, forming a set of multi-scale spatial feature maps. Simultaneously, the target fused features are fed into a global average pooling module to compress their spatial dimensions, resulting in a one-dimensional channel vector. The channel weight structure is then reconstructed through a 1×1 convolution to generate global contextual features. Taking into account that the spatial importance of different regions is not consistent, a spatial attention mechanism is introduced to construct a spatial attention map based on the target fusion feature. This process performs maximum pooling and average pooling operations on the feature map, cascades them, and fuses them through a 7×7 large receptive field convolution kernel. The sigmoid function is then used for normalization to generate a spatial attention map, which reflects the weight level of each position in the current image in the discrimination task. The spatial attention map is multiplied element-by-element with the target fusion feature to complete the attention weighting operation and output a spatial weighted feature map. The multi-scale spatial features, global context features and spatial weighted features are cascaded and spliced in the channel dimension to comprehensively aggregate their complementary information, and a layer of standard convolution operation is introduced to perform feature fusion and dimensional unification on the spliced multi-channel features to generate spatial relationship feature data. Step S104, feature importance weighting and bidirectional cross-fusion processing are performed on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation;
[0031] Specifically, global average pooling operations are applied to the texture feature data and spatial relationship feature data respectively, compressing their high-dimensional feature maps into a set of channel-level statistical vectors, thereby preserving the response distribution in the global contextual semantic background. The channel vectors are input into a set of nonlinear feature encoders composed of multi-layer perceptrons, which are subjected to feature mapping and activation normalization, and output texture feature importance scores and spatial feature importance scores respectively. Based on the feature importance evaluation results, a weighted operation is performed on the original texture feature data and spatial relationship feature data, and each channel is scaled using its corresponding importance score as the weight coefficient to obtain the initial weighted fusion feature. A bidirectional cross-attention mechanism is introduced to construct an explicit semantic interaction path between texture and space. This mechanism treats texture feature data as query terms and spatial relational feature data as a set of keys and values. Similarity matrices are calculated for both channels and spatial positions. The correlation between each texture feature channel and spatial feature is measured using a dot product operation. The similarity matrix is normalized using a softmax function to generate a texture-to-spatial attention map. The spatial-to-texture mapping path is then constructed in reverse, with spatial features as query terms and texture features as key-value pairs. Similarity calculation and normalization are performed again to generate a spatial-to-texture attention map. The original features are weightedly aggregated based on the two attention maps: The spatial features are weighted summed using the texture-to-spatial attention map to generate a texture-guided spatial feature enhancement result; and the texture features are weighted summed using the space-to-texture attention map to generate a space-guided texture feature enhancement result. These two results are cascaded and fused to form a bidirectional cross-attention feature, capturing semantic resonance and complementary responses across modalities. Furthermore, the initial weighted fusion feature and the bidirectional cross-attention feature are concatenated in the channel dimension. Material constraint logic for surface defects in autoclaved aerated concrete components is introduced as a priori control mechanism. This material constraint is composed of the physical morphological characteristics of the component defect type, including the high aspect ratio of cracks, the closed boundary characteristics of pores, and the uniform texture and irregular edge structure of spalled areas. This constraint is converted into a loss term using the morphological constraint function defined in the material rule library and embedded as a weight adjustment factor during the feature fusion stage. Integrating a deep neural network fusion module, the fused features after channel splicing and the material prior constraints are mapped into a unified feature space, outputting a comprehensive defect feature representation with global judgment capabilities, local expression capabilities, and material consistency.
[0032] Step S105: Analyze the surface defects of the autoclaved aerated concrete component based on the comprehensive defect feature representation, and output a comprehensive surface quality score and grade classification result.
[0033] Specifically, the high-dimensional comprehensive defect feature map is restored to a predicted image of the same size as the original image. The feature map is upsampled using a layer-by-layer transposed convolution structure. The resolution is restored step by step through convolution reverse reasoning with a step size of 2. The Sigmoid function is applied at the output stage to compress the result into a probability map between 0 and 1, and the probability of defect occurrence corresponding to each pixel is obtained. The OTSU adaptive threshold algorithm is then applied to binarize the probability map to form a defect binary map, marking the spatial position and morphological contours of each potential defect area. Contour detection and regional feature extraction operations are performed on each connected area in the defect binary map, including area, aspect ratio, boundary complexity, grayscale mean, edge texture characteristics and other dimensional indicators to form a regional feature vector. The feature vector is then input into a pre-trained multi-category defect discrimination network. By combining convolution extraction with full connection discrimination, each defect area is classified and the classification results of specific defect types such as cracks, pores, and detachment are output. Based on the second spatial transformation matrix set, surface defects detected from various viewpoints in autoclaved aerated concrete components are mapped to the component-to-component three-dimensional coordinate system, achieving a unified spatial projection of defect location information. To avoid duplicate detection of the same defect from multiple viewpoints, a defect fusion algorithm based on the intersection-over-union ratio and confidence level is employed to remove and merge overlapping regions. Specifically, if the overlap of a defect region from multiple viewpoints exceeds a threshold and the confidence level is low, the region is automatically discarded to construct a comprehensive defect distribution map. Pixel-level area statistics are performed on the binary defect map, and the surface defect rate index is calculated by calculating the proportion of all defect pixels to the total surface area. This index is used to characterize the overall defect severity of the component. The defect classification results are combined with the defect area data to create a model. A weighted coefficient is assigned to each defect type based on its impact on mechanical properties to establish an initial surface quality index. Based on the structural parameters, such as cement ratio, forming pressure, and fine aggregate content, a pre-defined process influence coefficient model is used to weight the initial quality index to obtain a weighted surface quality index. A regional density uniformity analysis is performed on the omnidirectional defect distribution map. A gridding strategy is used to divide the component surface into multiple regional units. The difference in defect density within each unit is calculated, and a defect distribution uniformity index is formed through variance normalization. A lower index indicates a more concentrated defect distribution, while a higher index indicates a more uniform distribution of defects, posing varying degrees of risk warning to structural stability. The surface defect rate, weighted surface quality index, and defect distribution uniformity index are input into the surface quality scoring model. The model performs a weighted fusion operation on these three indicators using a set weighting function, outputting a comprehensive surface quality score between 0 and 100. The system matches this score to the preset quality grading standard, generates a component quality grading result, and outputs the score result in conjunction with the original component identification.
[0034] In an embodiment of the present invention, a multi-view camera array and spatial transformation matrix optimization technology are used to achieve full coverage of all surfaces of autoclaved aerated concrete components, solving the blind spot and image deformation problems existing in traditional single-view detection and ensuring the integrity and accuracy of defect detection. In view of the grayish-white and low-contrast characteristics of the autoclaved aerated concrete surface, illumination equalization, adaptive contrast enhancement, and variable-scale bilateral filtering technologies are used to significantly improve the visibility and edge clarity of defects, laying the foundation for subsequent feature extraction. Deep separable convolution and multi-scale texture coding structures are used to optimize the texture characteristics of different types of defects such as pores, cracks, and depressions on the surface of autoclaved aerated concrete, significantly reducing the number of parameters and computational complexity while maintaining a high feature extraction capability. Through the combination of deformable convolution, multi-scale spatial pooling, and attention mechanism, effective detection of defects of different scales, from tiny pores to large-scale detachment, is achieved, solving the problem that traditional methods are difficult to adapt to multi-scale defect recognition at the same time. By using feature importance weighting and a bidirectional cross-attention mechanism, the complementary advantages of texture and spatial features are fully integrated, and constraints based on autoclaved aerated concrete material science knowledge are introduced to improve the accuracy and robustness of defect identification, especially its adaptability under complex environmental conditions. A multi-level quality assessment system integrating defect classification, location, and process parameters was constructed, and a quantitative relationship model between defect formation and process parameters such as cement ratio, molding pressure, and fine aggregate content was established. This provides a scientific basis for production process optimization and quality control, and realizes the transition from simple defect detection to comprehensive quality assessment and process guidance. Through lightweight model design and structural optimization, the system can run in real time on edge equipment on the production line, meeting the real-time detection needs in industrial production environments and providing a reliable real-time quality monitoring method for the production of autoclaved aerated concrete components.
[0035] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0036] The main view camera is installed directly above the conveyor belt of the autoclaved aerated concrete component, and N auxiliary view cameras are distributed in a ring around the autoclaved aerated concrete component to obtain a multi-view camera array;
[0037] A multi-view camera array is started to synchronously shoot the autoclaved aerated concrete component to obtain the original image set. Each camera in the multi-view camera array is individually calibrated with internal parameters to obtain the focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient of each camera.
[0038] According to the focal length, principal point coordinates, radial distortion coefficient and tangential distortion coefficient of each camera, the conversion relationship between the component three-dimensional coordinate system and the coordinate system of each camera is established to obtain the external parameter matrix of each camera;
[0039] Based on the positional relationship of the matching feature points of each viewpoint image in the original image set, the homography matrix between adjacent cameras is calculated to obtain a first spatial transformation matrix set;
[0040] Based on the extrinsic matrix of each camera, the first spatial transformation matrix set is optimized to minimize the reprojection error to obtain the second spatial transformation matrix set, and the original image set is mapped and transformed by the second spatial transformation matrix set to obtain a surface full coverage correction image;
[0041] The surface full coverage correction image is subjected to illumination equalization, contrast enhancement and noise reduction processing to obtain a preprocessed image.
[0042] Specifically, a primary-view camera is mounted vertically above the conveyor belt of autoclaved aerated concrete components, covering the component's upper surface area as the primary imaging reference. A circular path is defined around the component, and N auxiliary-view cameras are installed at equal intervals to form a closed ring in the horizontal direction. The angle between the optical axes of adjacent cameras is 360° / N. All cameras use industrial-grade CCD image sensors with a resolution of at least 1920 × 1080 and a frame rate greater than 60 fps, enabling clear images to be captured under high-speed transmission conditions. They are connected to a central image processing control unit via an Ethernet link and establish communication with the production line's PLC control module, enabling centralized control and synchronized triggering of the entire camera array. As a component moves into the inspection area, the system receives position signals from the PLC controller and activates the multi-view camera array to synchronously capture the component. This ensures that all viewpoints are exposed at the same time, avoiding image distortion or occlusion caused by time differences. After image acquisition, a raw image set consisting of the primary viewpoint image and N auxiliary viewpoint images is obtained. To eliminate geometric distortion introduced by the lens structure, each camera is individually calibrated with internal parameters. The calibration method is based on a standard checkerboard target pattern. Images of the calibration plate are captured from multiple angles, and the coordinates of the corner points in each image are extracted. Based on the pinhole camera model, an internal parameter solution equation is constructed to solve for the corresponding camera focal length parameters, principal point coordinates, radial distortion coefficients, and tangential distortion coefficients. These parameters are used to model and repair nonlinear distortion during imaging. After obtaining the camera intrinsic parameters, the images from each camera are projected into a unified component coordinate system. To achieve this, a camera extrinsic parameter matrix is established. This matrix consists of the rotation matrix and translation vector of each camera relative to the component coordinate system, representing the transformation relationship between the camera coordinate system and the world coordinate system. The relative position between the component and each camera is measured using a combination of a calibration target and a laser ruler, and the extrinsic parameter matrix is solved in combination with the internal parameter results. Feature point matching is performed based on the overlapping areas of the images from each viewpoint in the original image set. Key point information is extracted using robust feature extraction algorithms such as SIFT or ORB, and cross-image matching is performed using a FLANN or BF matcher. After obtaining the correspondence between the feature points of each pair of adjacent cameras, a homography matrix solution equation is constructed based on the perspective projection model. The RANSAC algorithm is used to eliminate false matches and solve the homography transformation matrix between each pair of adjacent images, forming a first spatial transformation matrix set. This matrix set describes the two-dimensional spatial projection relationship between one viewpoint image and its adjacent image. This first spatial transformation matrix set is optimized based on the extrinsic parameter matrices of each camera, minimizing the reprojection error of feature points between images from different viewpoints during the transformation process.Using a least-squares optimization strategy, the objective function is constructed as the sum of the squared Euclidean distances between all matching points between viewpoints after mapping. Gradient descent is then used to iteratively calculate the result until the error change over two consecutive rounds falls below a set threshold, resulting in a converged second spatial transformation matrix set. Based on this second spatial transformation matrix set, the original image set is uniformly mapped and transformed, and the images from each viewpoint are reconstructed and projected onto the same plane or three-dimensional component coordinate domain, generating a set of geometrically consistent corrected images that cover the entire component surface. Image enhancement processing is then performed on the corrected image set, which includes illumination equalization, contrast enhancement, and noise reduction filtering. During the illumination equalization stage, a Gaussian filter is applied to the image to obtain a background brightness distribution map. The ratio of the target average brightness to the current brightness is calculated, and an illumination compensation coefficient matrix is constructed. This matrix is applied to the original image to normalize the image illumination and address issues such as shadows, dark corners, and highlights caused by the directionality of the on-site light source. Image contrast is enhanced based on histogram normalization. By mapping the cumulative distribution function of the image grayscale histogram to a preset target distribution function, the image grayscale dynamic range is expanded, thereby improving the visibility of fine cracks and low-contrast defects against light backgrounds. To suppress high-frequency noise generated by the enhancement and avoid blurring and loss of edge details, a variable-scale bilateral filter is introduced. The filter's spatial kernel size and range weights are adaptively adjusted based on the local gradient value of the image. This performs strong smoothing in flat texture areas and preserves image details at structural edges, achieving a refined improvement in overall image quality. A set of preprocessed images is obtained through image acquisition, geometric correction, spatial reconstruction, and image enhancement.
[0043] In a specific embodiment, the step of performing illumination equalization, contrast enhancement, and noise reduction on the surface full coverage corrected image to obtain a pre-processed image may specifically include the following steps:
[0044] According to the focal length, principal point coordinates, radial distortion coefficient and tangential distortion coefficient of each camera, the surface full coverage correction image is subjected to distortion correction to obtain the surface image of the standard component;
[0045] Perform Gaussian filtering on the surface image of the standard component to obtain an image representing the background brightness distribution;
[0046] Calculate the correction coefficient matrix based on the image representing the background brightness distribution, and apply the correction coefficient matrix to the standard component surface image to obtain the illumination balanced image;
[0047] Perform histogram statistical analysis on the illumination equalization image, calculate the cumulative distribution function, and map the cumulative distribution function to the preset target distribution function to obtain a contrast-enhanced image;
[0048] Calculate the local gradient value of the contrast enhanced image, and adaptively adjust the filter kernel size and range parameters of the scaled bilateral filter according to the local gradient value to obtain the target filtering parameters;
[0049] The contrast enhanced image is processed by variable scale bilateral filtering based on the target filtering parameters to obtain a preprocessed image.
[0050] Specifically, distortion correction is performed on the original image based on a camera imaging model. Since each viewpoint image is acquired using an industrial-grade CCD camera, the imaging process introduces radial distortion caused by the lens curvature structure and tangential distortion due to lens assembly errors. Therefore, a distortion correction mapping model is established based on internal parameters such as focal length, principal point coordinates, radial distortion coefficients, and tangential distortion coefficients acquired during the previous camera calibration phase. The position of each pixel in the image is remapped according to the pinhole imaging and distortion correction model, and accurately back-projected onto a theoretical distortion-free plane to generate a standard component surface image. Brightness modeling is performed on the standard component surface image to address issues such as uneven lighting, local overexposure, and shadow occlusion in the factory environment. A Gaussian filter is introduced to perform spatial smoothing on the standard image. During this process, a fixed-scale Gaussian kernel is used to convolve the image, eliminating high-frequency variations in details while preserving the global brightness trend. This results in an image representing the background brightness distribution, representing the ideal brightness baseline for uniform component surface illumination. Based on the ratio between the background brightness distribution map and the desired target brightness, an illumination correction coefficient matrix is calculated. The value at each pixel position in this matrix represents the ratio factor between the original pixel brightness and the target brightness. This coefficient matrix is applied pixel by pixel to the standard image, performing pixel-level brightness normalization to achieve illumination equalization and obtain an illumination-balanced image. Histogram statistical analysis is performed on the illumination-balanced image to construct the grayscale histogram of the original image. Based on this histogram, a cumulative distribution function is calculated, reflecting the cumulative probability distribution of each grayscale value in the entire image. The cumulative distribution function is mapped to a preset target distribution function, typically a uniform distribution or an S-shaped distribution with enhanced intermediate grayscale response. The grayscale of each pixel in the original image is adjusted by inverse mapping the cumulative distribution function, thereby redistributing the image's grayscale range. Low-gray areas are stretched and high-gray areas are compressed, resulting in a contrast-enhanced image with enhanced grayscale differences and more sensitive texture response. An adaptive variable-scale bilateral filtering algorithm is used to optimize the image. By performing local gradient analysis on the contrast-enhanced image, the horizontal and vertical brightness changes are calculated pixel by pixel to form a local gradient map. The filter kernel size and weight distribution are then dynamically adjusted based on the gradient value at each pixel. The spatial domain filter kernel size is positively correlated with the gradient intensity, ensuring more detailed kernel operations in areas with complex textures. The domain weight parameter is negatively correlated with the gradient. This means that when local changes are drastic, the weight range is reduced to avoid blurred boundaries, while when local gradients are smooth, the fusion range is increased to improve noise reduction. This step adaptively sets the kernel function parameters of the bilateral filter based on the image's structural characteristics, smoothing texture interference and random noise while maintaining edge clarity, effectively improving the image's overall signal-to-noise ratio. The above steps ultimately generate a set of preprocessed images.
[0051] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0052] The preprocessed image is input into the first depth-wise separable convolution layer for processing to obtain the first layer of texture features; the first layer of texture features is input into the second depth-wise separable convolution layer for processing to obtain the second layer of texture features; the second layer of texture features is input into the third depth-wise separable convolution layer for processing to obtain the third layer of texture features; the third layer of texture features is input into the fourth depth-wise separable convolution layer for processing to obtain the fourth layer of texture features;
[0053] The first layer texture features, the second layer texture features, the third layer texture features and the fourth layer texture features are input into four dilated convolutional layers in parallel for processing to obtain a multi-scale texture feature group;
[0054] Calculating the first gradient difference in the horizontal direction and the second gradient difference in the vertical direction of the fourth layer texture feature respectively, and performing feature enhancement on the fourth layer texture feature according to the first gradient difference and the second gradient difference to obtain a bidirectional gradient feature;
[0055] The multi-scale texture feature group is weightedly fused and connected with the bidirectional gradient feature in the channel dimension and then integrated to obtain the texture feature data.
[0056] Specifically, the preprocessed image is input into the first depthwise separable convolutional layer, which consists of a depthwise convolution operation in a spatial dimension combined with a cross-channel pointwise convolution. The depthwise convolution extracts local patterns along the spatial range within each channel, while the pointwise convolution linearly fuses the features between channels. After the nonlinear activation function PReLU and batch normalization, the output is the first layer of texture features. This feature is mainly composed of the rough texture structure and edge response of the component surface, which belongs to the shallow perceptual representation. The first layer of texture features is input into the second depthwise separable convolutional layer. This layer expands the receptive field while retaining the local texture of the previous layer. The layer strengthens the modeling ability of texture continuity between adjacent regions through 7×7 or 5×5 deep convolution kernels, extracting mid-level texture features containing microscopic repetitive textures and weak crack edges, forming the second layer feature map. The second-layer feature map is passed to the third depthwise separable convolutional structure, where the convolution kernel size is reduced to 3×3 to focus on finer-grained structural edges. The network depth is also increased to capture complex texture overlaps, such as interlaced cracks, dense pores, and blurred delamination boundaries. The resulting third-layer feature map exhibits strong structural understanding and mid-term image semantic perception. The third-layer features are then fed into the fourth depthwise separable convolutional layer, which emphasizes feature integration compared to the previous layers, extracting high-order semantic texture features and performing deep pattern abstraction of complex surface defects, resulting in the fourth-layer texture feature map. The texture feature maps from the first to fourth layers are then fed into four dilated convolutional (or atrous convolutional) layers with different dilation rates for parallel processing. The kernel size of these dilated convolutions is uniformly set to 3×3, and the dilation rates are set to 1, 2, 4, and 8, respectively. This extends the receptive field from the local neighborhood to distant pixels, achieving an expansion of the perceptual range without increasing the number of convolutional parameters. This effectively captures the structural texture features of different types of defects, such as point-like pores, linear cracks, and flaking, and generates a multi-scale texture feature set. Structural enhancement is performed on high-level texture feature maps, especially the fourth-layer texture feature map, to further enhance its edge recognition capabilities while maintaining its semantic integrity. The grayscale differences of this feature map are calculated in the horizontal and vertical directions. At each pixel, the absolute difference with its left and right neighbors is taken to obtain the first gradient response, and the difference with its upper and lower neighbors is taken to obtain the second gradient response. The gradient differences in these two directions represent the intensity of structural changes in the image in the horizontal and vertical directions, respectively. Lightweight convolution operations are used to perform edge mapping conversion on the gradient maps in both directions, and they are connected in the channel dimension to form a bidirectional gradient feature map, thereby explicitly enhancing the expression ability of potential edge changes and structural fracture areas on the component surface, so that the edge morphology that is easily weakened by smoothing operations in deep convolutions can be restored and enhanced. A multi-scale texture feature group consisting of a set of dilated convolution outputs covering multiple layers of semantics and a set of bidirectional gradient feature maps with clear boundary direction responses are obtained. In order to form a unified and expressive texture representation, these features are fused and integrated.A channel-weighted mechanism is introduced into the multi-scale texture feature set, with learnable weight parameters added to each channel. Based on the contribution of each scale feature to the final task, a weighted summation process is automatically performed to form a set of primary texture feature maps. This fused texture feature map is then concatenated with the previously constructed bidirectional gradient feature map along the channel dimension, resulting in a fusion of texture features and gradient structure in the feature tensor structure. A 1×1 convolution is introduced into the final texture integration module to achieve feature compression and channel interaction, outputting texture feature data.
[0057] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0058] The preprocessed image is input into the MobileNetV3 backbone network for processing to obtain the initial spatial features;
[0059] The initial spatial features and texture feature data are connected in the channel dimension, and feature fusion is performed through a 1×1 convolution layer to obtain the target fusion feature;
[0060] Perform parallel convolution processing on the target fusion features to obtain multi-scale spatial features, and then process the target fusion features through global average pooling and 1×1 convolution to obtain global context features;
[0061] Calculate the spatial attention map based on the target fusion feature, and perform element-wise multiplication of the spatial attention map and the target fusion feature to obtain the spatial weighted feature;
[0062] Multi-scale spatial features, global context features, and spatial weighted features are connected in the channel dimension and integrated through a convolutional layer to obtain spatial relationship feature data.
[0063] Specifically, the preprocessed image is fed into the backbone of the lightweight MobileNetV3 neural network. Through the network's depthwise separable convolutional modules, asymmetric channel attention mechanism, and bottleneck residual structure, the global spatial layout and geometric construction patterns of the component surface are extracted with high efficiency, generating an initial spatial feature map. This feature map preserves the structural semantic information of the preprocessed image, encompassing spatial patterns such as edge orientation, region size, location distribution, and texture clustering. The initial spatial features are concatenated with texture feature data at the channel level to construct a multimodal fused feature map that incorporates both spatial distribution patterns and texture semantic structure. This concatenation operation unifies features from different perceptual pathways into the same representation space and, through structural alignment, enables correlation mapping between the two feature types at the spatial distribution level. To achieve information compression and channel interaction, the concatenated fused features are convolved with a 1×1 convolutional layer. The fused target feature map is then output through linear mapping along the channel dimension and activation functions. To capture the characteristic features of component surface defects at different scales, a parallel multi-scale convolutional pathway structure is designed to process the target fused features. Each pathway uses dilated convolution kernels with different dilation rates (e.g., 6, 12, 18, and 24) to expand the receptive field while maintaining the same number of convolution parameters. These parallel pathways extract spatial features from four levels: local neighborhoods, small regional patches, large structural outlines, and even the overall component layout, generating a set of multi-scale spatial feature maps with different perceptual levels. The feature maps output by the convolutional pathways capture the contextual scales required for different defect types, such as dense pores, clusters of small cracks, large spalling patches, and geometric fractures, and represent their semantic hierarchy as feature channels. While performing the parallel convolutions, a global average pooling operation is applied to the target fused feature map to compress the spatial dimensions and extract the overall statistical semantics of the image. The channel dimensions are then restored through a 1×1 convolution to form a global contextual feature map. This feature map serves as a contextual prior for the entire component scenario, guiding subsequent local feature judgments by incorporating structural information such as the overall component morphology and common defect locations. After feature fusion, a spatial attention mechanism is introduced to automatically identify and highlight the spatial locations in the current image that contribute most to the defect recognition task. Specifically, the target fusion feature map is subjected to maximum pooling and average pooling to preserve different statistical features. The two are then concatenated along the channel dimension and fed into a standard convolutional layer with a 7×7 receptive field. The resulting two-dimensional spatial attention map is normalized using a sigmoid function. Each pixel value in this attention map represents the importance score of the corresponding spatial location. This map is then multiplied pixel by pixel with the original target fusion feature map. This multiplication multiplies the channel value at each location by the corresponding attention weight, enhancing salient areas and suppressing non-critical areas to produce a spatially weighted feature map.Multi-scale spatial features, global context features, and spatially weighted features are cascaded and integrated in the channel dimension, compressing features from multiple semantic levels into a unified feature space. To avoid feature redundancy and channel dimension mismatch caused by direct concatenation, a 3×3 convolutional layer is introduced after the integration to perform convolutional fusion. This layer jointly models multi-source features and enhances feature interaction through nonlinear activation functions, making the fused feature map more discriminative and expressive, ultimately generating spatial relationship feature data.
[0064] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0065] The texture feature data and spatial relationship feature data are processed by global average pooling and multi-layer perceptron respectively to obtain the texture feature importance score and spatial feature importance score;
[0066] Based on the texture feature importance score and the spatial feature importance score, the texture feature data and the spatial relationship feature data are weighted to obtain the initial weighted fusion feature;
[0067] The texture feature data is used as the query item and the spatial relationship feature data as the key-value pair. The similarity matrix between the features is calculated and normalized to obtain the texture-to-space attention map.
[0068] The spatial relationship feature data is used as the query item and the texture feature data as the key-value pair. The similarity matrix between the features is calculated and normalized to obtain the space-to-texture attention map.
[0069] Based on the texture-to-space attention map, the spatial relationship feature data is weighted and summed. At the same time, based on the space-to-texture attention map, the texture feature data is weighted and summed to obtain the bidirectional cross-attention feature.
[0070] The initial weighted fusion features are connected with the bidirectional cross-attention features in the channel dimension, and the autoclaved aerated concrete material constraint is introduced to obtain a comprehensive defect feature representation.
[0071] Specifically, the texture feature data and spatial relationship feature data are fed into a global average pooling module. The pixel values of each channel in the spatial dimension are averaged, compressing the original three-dimensional feature map (height × width × channel) into a one-dimensional channel vector. This vector represents the global response strength of each channel across the entire image, reflecting the global sensitivity of the feature. This one-dimensional vector is then fed into a nonlinear mapping network consisting of a multilayer perceptron. After two linear transformation layers and an activation function, it outputs a normalized importance vector. This vector is mapped to a value between 0 and 1 using a sigmoid function. These vectors correspond to the importance scores of the texture and spatial features, respectively, indicating the relative importance of each feature type for the recognition task within the component sample. These two importance score vectors are then used to perform a channel-wise weighting operation on the original texture and spatial relationship feature data. Each channel's original feature is multiplied by its corresponding importance score, emphasizing high-weight channels and suppressing low-response channels. The weighted texture and spatial features are then channel-wise concatenated to form the initial weighted fusion feature. Based on this, a bidirectional cross-attention mechanism is designed to establish a cross-semantic interaction channel between texture and spatial features, enabling explicit modeling of inter-modal connections and mutually guided enhancement. This mechanism uses texture features as query terms and spatial features as key-value pairs to form triplets. It calculates the similarity between each texture channel position and each spatial channel position via tensor dot products, constructing a high-dimensional feature similarity matrix. This matrix is normalized using a softmax function to form a texture-to-spatial attention map, where each element represents the degree of dependence of a texture feature position on a spatial feature position, or the attention weight. This attention map is then used to perform a weighted sum operation on the original spatial feature map, recombining all spatial channels under attention weights to output a texture-guided spatial reconstruction feature map. The reverse construction path uses spatial features as query terms and texture features as key-value pairs. A spatial-to-texture similarity matrix is again constructed via dot products, and then normalized to obtain a spatial-to-texture attention map. This attention map is then used to perform attention-weighted fusion on the texture feature map to generate a spatially guided texture enhancement map. The dual-channel attention mechanism effectively captures the contextual relationship between the two modalities, demonstrating enhanced discrimination in ambiguous areas, such as those where crack boundaries and backgrounds are unclear, or where pores merge with surface gray streaks. This results in a bidirectional cross-attention feature representation. The initial weighted fusion features are concatenated with the bidirectional cross-attention features along the channel dimension, and stacked to construct a higher-dimensional deep semantic representation space. A material constraint mechanism is introduced during the fusion phase, embedding material property priors into the fusion strategy in the form of a loss function.These include the continuity constraints of high aspect ratio cracks, the circular boundaries and closure constraints of pores, the irregular shape of the detached area and the uniform internal grayscale distribution constraints, etc. The boundary shape, grayscale gradient, area distribution, etc. of the result features are adjusted in the fusion network through the morphological loss function, and the comprehensive defect feature representation is obtained after fusion convolution compression and normalization processing.
[0072] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0073] Perform feature map size recovery on the comprehensive defect feature representation to obtain a defect binary map;
[0074] Extract regional feature vectors from each connected region in the defect binary map, perform multi-category defect discrimination on the regional feature vectors, and obtain surface defect classification results;
[0075] Based on the second spatial transformation matrix set, the surface defects of the autoclaved aerated concrete components detected at each viewing angle are mapped to the component-component three-dimensional coordinate system, and the defects of the overlapping areas of adjacent viewing angles are fused by using the intersection-over-union ratio and confidence level to obtain a full-scale defect distribution map;
[0076] Perform area statistical analysis on the defect binary image to obtain the surface defect rate index;
[0077] The surface quality index is calculated based on the surface defect category results and defect area, and is corrected by combining the influence coefficients of cement ratio, molding pressure and fine aggregate content to obtain a weighted surface quality index.
[0078] The regional defect density uniformity of the omnidirectional defect distribution map is calculated to obtain the defect distribution uniformity index. The surface quality comprehensive score is calculated based on the surface defect rate index, weighted surface quality index, and defect distribution uniformity index, and the corresponding grade classification results are generated according to the surface quality comprehensive score.
[0079] Specifically, the integrated defect feature representation is spatially resized, restoring it from its downsampled state to the resolution of the original input image. This process involves upsampling the feature map using a layer-by-layer transposed convolutional structure. Each layer of transposed convolution uses a stride of 2 and padding of 1. BatchNorm and ReLU activation functions are embedded to enhance nonlinear representation and training stability. After three rounds of upsampling, the feature map is restored to its original size. A 1×1 convolution and Sigmoid normalization activation function are introduced on the output channel to output the confidence value of each pixel in the defect prediction map, generating a defect probability map. The OTSU adaptive thresholding algorithm is then applied to this probability map for pixel-level binary segmentation. Locations with confidence values above a threshold are identified as defective pixels, constructing a defect binary map. The defect binary map is then partitioned into several connected regions, and regional feature vectors representing the geometric morphology and texture information of each connected region are extracted. This extraction includes multiple dimensions, such as area, aspect ratio, perimeter, edge complexity, grayscale mean, edge gradient statistics, and directional consistency. These regional features are fed into a pre-trained multi-category defect classifier as classification features. The classifier internally utilizes a shallow convolutional structure or MLP classification network. The softmax layer outputs a probability distribution of defect types for each region, with the maximum probability term used as the final classification result. This output includes surface defect classification results with labels such as "crack," "pore," "peeling," and "impurity." Based on a second set of spatial transformation matrices, defect regions detected from all perspectives are uniformly mapped to the component's three-dimensional coordinate system. This mapping process uses known intrinsic and extrinsic parameter matrices to perform a back-projection transformation from image pixel space to physical space, ensuring consistent spatial alignment across multiple perspectives. Due to overlapping fields of view across multiple cameras, to avoid the same defect region being counted multiple times, the intersection over union (IoU) is used as the criterion for regional overlap, with an IoU threshold of 0.3. When the IoU of defect regions from two different perspectives exceeds this threshold, their confidence values are simultaneously referenced for fusion decisions. Regions with higher confidence levels are retained or weighted merged, achieving redundancy-free integration of multi-perspective defect information and forming a comprehensive defect distribution map. Pixel-level area statistics are performed on the binary defect image, calculating the ratio of the total area of all defect regions to the area of the entire component surface image. This outputs a surface defect rate index, which serves as the most intuitive surface quality metric. Furthermore, combined with the previously obtained defect category information, the area occupied by different defect types is calculated and assigned a preset weight coefficient. The weighted summation results in a surface quality index, which reflects the severity of the defect's impact on material performance. Correction factors are introduced, including three key parameters: cement ratio, molding pressure, and fine aggregate content.By constructing an empirical model to solve the cement ratio influencing factor, the molding pressure influencing factor, and the fine aggregate influencing factor respectively, and introducing these factors into the correction formula of the surface quality index, a weighted surface quality index with more real production relevance is calculated. A regional defect density uniformity analysis is performed on the all-round defect distribution map. The component surface is divided into multiple spatial sub-blocks. The defect density of each sub-block is statistically analyzed, and then the mean and standard deviation of the defect density of all sub-blocks are calculated. The standard deviation normalization index is used to represent the degree of uniformity. The higher the index, the more concentrated or extreme the defect distribution, and the lower the index, the more uniformly dispersed the defects, reflecting the consistency of the component surface structure and the local processing stability. The comprehensive surface quality score is calculated based on the surface defect rate index, the weighted surface quality index, and the defect distribution uniformity index. The quality grade is defined according to the score value range. The final score result and grade division are output and bound to the production identification or quality label of the component.
[0080] The above describes the intelligent recognition method for surface defects of autoclaved aerated concrete components according to the embodiment of the present invention. The following describes the intelligent recognition system for surface defects of autoclaved aerated concrete components according to the embodiment of the present invention. Figure 2 In one embodiment of the present invention, an intelligent system for identifying surface defects of autoclaved aerated concrete components includes:
[0081] A mapping module 201 is used to perform omnidirectional photography and spatial transformation matrix mapping on the surface of the autoclaved aerated concrete component to obtain a pre-processed image;
[0082] The feature extraction module 202 is used to extract texture features from the pre-processed image to obtain texture feature data;
[0083] Feature analysis module 203, used to perform spatial feature analysis on the pre-processed image and texture feature data to obtain spatial relationship feature data;
[0084] Cross-fusion module 204, used to perform feature importance weighting and bidirectional cross-fusion processing on texture feature data and spatial relationship feature data to obtain a comprehensive defect feature representation;
[0085] The output module 205 is used to perform surface defect analysis on the autoclaved aerated concrete component based on the comprehensive defect feature representation, and output a comprehensive surface quality score and grade classification result.
[0086] Through the collaborative efforts of the aforementioned components, a multi-view camera array and spatial transformation matrix optimization technology achieve full coverage of all surfaces of autoclaved aerated concrete components, addressing the blind spot and image distortion issues inherent in traditional single-view detection and ensuring the completeness and accuracy of defect detection. Addressing the off-white, low-contrast nature of autoclaved aerated concrete surfaces, illumination equalization, adaptive contrast enhancement, and scaled bilateral filtering techniques significantly improve defect visibility and edge clarity, laying the foundation for subsequent feature extraction. Employing a depthwise separable convolution and multi-scale texture encoding architecture, the system specifically optimizes the texture characteristics of various defects, such as pores, cracks, and dents, on the autoclaved aerated concrete surface. This significantly reduces the number of parameters and computational complexity while maintaining high feature extraction capabilities. Through a combination of deformable convolution, multi-scale spatial pooling, and an attention mechanism, the system effectively detects defects of varying scales, from tiny pores to large-scale detachments, addressing the difficulty of traditional methods in simultaneously accommodating multi-scale defect recognition. By using feature importance weighting and a bidirectional cross-attention mechanism, the complementary advantages of texture and spatial features are fully integrated, and constraints based on autoclaved aerated concrete material science knowledge are introduced to improve the accuracy and robustness of defect identification, especially its adaptability under complex environmental conditions. A multi-level quality assessment system integrating defect classification, location, and process parameters was constructed, and a quantitative relationship model between defect formation and process parameters such as cement ratio, molding pressure, and fine aggregate content was established. This provides a scientific basis for production process optimization and quality control, and realizes the transition from simple defect detection to comprehensive quality assessment and process guidance. Through lightweight model design and structural optimization, the system can run in real time on edge equipment on the production line, meeting the real-time detection needs in industrial production environments and providing a reliable real-time quality monitoring method for the production of autoclaved aerated concrete components.
[0087] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0088] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling an intelligent surface defect identification device for autoclaved aerated concrete components (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0089] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent method for identifying surface defects of autoclaved aerated concrete components, characterized in that: include: The surface of the autoclaved aerated concrete component is photographed in all directions and mapped by spatial transformation matrix to obtain a pre-processed image; Extracting texture features from the preprocessed image to obtain texture feature data; Performing spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data; The texture feature data and the spatial relationship feature data are subjected to feature importance weighting and bidirectional cross fusion processing to obtain a comprehensive defect feature representation, including: processing the texture feature data and the spatial relationship feature data respectively through global average pooling and multi-layer perceptron to obtain a texture feature importance score and a spatial feature importance score; based on the texture feature importance score and the spatial feature importance score, performing a weighted operation on the texture feature data and the spatial relationship feature data to obtain an initial weighted fusion feature; using the texture feature data as a query item and the spatial relationship feature data as a key-value pair to calculate a similarity matrix between features and performing normalization processing to obtain a texture-to-space attention map; using the spatial relationship feature data as a query item and the texture feature data as a key-value pair, calculating the similarity matrix between features and performing normalization processing to obtain a space-to-texture attention map; performing weighted summation on the spatial relationship feature data based on the texture-to-space attention map, and simultaneously performing weighted summation on the texture feature data based on the space-to-texture attention map to obtain a bidirectional cross-attention feature; connecting the initial weighted fusion feature with the bidirectional cross-attention feature in the channel dimension, and introducing autoclaved aerated concrete material constraints to obtain a comprehensive defect feature representation; Based on the comprehensive defect feature representation, surface defect analysis of the autoclaved aerated concrete component is performed, and a comprehensive surface quality score and grade classification result are output.
2. The intelligent identification method for surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The omnidirectional shooting and spatial transformation matrix mapping of the surface of the autoclaved aerated concrete component to obtain a pre-processed image includes: The main view camera is installed directly above the conveyor belt of the autoclaved aerated concrete component, and N auxiliary view cameras are distributed in a ring around the autoclaved aerated concrete component to obtain a multi-view camera array; Starting the multi-view camera array to synchronously photograph the autoclaved aerated concrete component to obtain an original image set, and performing individual internal parameter calibration on each camera in the multi-view camera array to obtain a focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient of each camera; According to the focal length, principal point coordinates, radial distortion coefficient and tangential distortion coefficient of each camera, the conversion relationship between the component three-dimensional coordinate system and the coordinate system of each camera is established to obtain the external parameter matrix of each camera; Based on the positional relationship of the matching feature points of the images of each perspective in the original image set, calculating the homography matrix between adjacent cameras to obtain a first spatial transformation matrix set; Based on the extrinsic matrix of each camera, the first spatial transformation matrix set is optimized to minimize the reprojection error to obtain a second spatial transformation matrix set, and the original image set is mapped and transformed by the second spatial transformation matrix set to obtain a surface full coverage correction image; The surface full coverage correction image is subjected to illumination equalization, contrast enhancement and noise reduction processing to obtain a preprocessed image.
3. The intelligent identification method for surface defects of autoclaved aerated concrete components according to claim 2, characterized in that: The performing illumination equalization, contrast enhancement and noise reduction processing on the surface full coverage correction image to obtain a preprocessed image includes: Performing distortion correction on the surface full coverage correction image according to the focal length, principal point coordinates, radial distortion coefficient, and tangential distortion coefficient of each camera to obtain a standard component surface image; Performing Gaussian filtering on the surface image of the standard component to obtain an image representing background brightness distribution; Calculating a correction coefficient matrix based on the image representing the background brightness distribution, and applying the correction coefficient matrix to the standard component surface image to obtain a lighting balanced image; Performing histogram statistical analysis on the illumination-equalized image, calculating a cumulative distribution function, and mapping the cumulative distribution function to a preset target distribution function to obtain a contrast-enhanced image; Calculating a local gradient value of the contrast-enhanced image, and adaptively adjusting a filter kernel size and a range parameter of a scaled bilateral filter according to the local gradient value to obtain target filter parameters; The contrast-enhanced image is subjected to variable-scale bilateral filtering based on the target filtering parameters to obtain a preprocessed image.
4. The intelligent identification method for surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The extracting texture features of the pre-processed image to obtain texture feature data includes: Inputting the preprocessed image into a first depthwise separable convolutional layer for processing to obtain a first layer of texture features; Inputting the first layer of texture features into a second depthwise separable convolutional layer for processing to obtain a second layer of texture features; Inputting the second layer of texture features into the third depthwise separable convolutional layer for processing to obtain the third layer of texture features; Inputting the third layer texture features into a fourth depthwise separable convolutional layer for processing to obtain a fourth layer texture features; Inputting the first layer of texture features, the second layer of texture features, the third layer of texture features, and the fourth layer of texture features into four dilated convolutional layers in parallel for processing to obtain a multi-scale texture feature group; respectively calculating a first gradient difference in a horizontal direction and a second gradient difference in a vertical direction of the fourth layer texture feature, and performing feature enhancement on the fourth layer texture feature according to the first gradient difference and the second gradient difference to obtain a bidirectional gradient feature; The multi-scale texture feature group is weightedly fused and connected with the bidirectional gradient feature in the channel dimension and then integrated to obtain texture feature data.
5. The intelligent identification method for surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The performing spatial feature analysis on the pre-processed image and the texture feature data to obtain spatial relationship feature data includes: Input the preprocessed image into the MobileNetV3 backbone network for processing to obtain initial spatial features; connect the initial spatial features with the texture feature data in the channel dimension, and perform feature fusion through a 1×1 convolution layer to obtain target fusion features; Performing parallel convolution processing on the target fusion features to obtain multi-scale spatial features, and performing global average pooling and 1×1 convolution processing on the target fusion features to obtain global context features; Calculating a spatial attention map based on the target fusion feature, and performing element-wise multiplication operation on the spatial attention map and the target fusion feature to obtain a spatial weighted feature; The multi-scale spatial features, the global context features and the spatial weighted features are connected in the channel dimension and integrated through a convolutional layer to obtain spatial relationship feature data.
6. The intelligent identification method for surface defects of autoclaved aerated concrete components according to claim 1, characterized in that: The surface defect analysis of the autoclaved aerated concrete component based on the comprehensive defect feature representation and output of the surface quality comprehensive score and grade classification results include: Performing feature map size restoration on the comprehensive defect feature representation to obtain a defect binary map; Extracting a regional feature vector from each connected region in the defect binary map, and performing multi-category defect discrimination on the regional feature vector to obtain a surface defect classification result; Based on the second spatial transformation matrix set, the surface defects of the autoclaved aerated concrete components detected at each viewing angle are mapped to the component's three-dimensional coordinate system. The defects in the overlapping areas of adjacent viewing angles are fused using the intersection-over-union ratio and confidence level to obtain a comprehensive defect distribution map. Performing area statistical analysis on the defect binary image to obtain a surface defect rate index; calculating a surface quality index based on the surface defect category results and defect area, and performing correction based on the influence coefficients of cement ratio, molding pressure, and fine aggregate content to obtain a weighted surface quality index; The regional defect density uniformity of the omnidirectional defect distribution map is calculated to obtain a defect distribution uniformity index, and a surface quality comprehensive score is calculated based on the surface defect rate index, the weighted surface quality index, and the defect distribution uniformity index, and a corresponding grade classification result is generated according to the surface quality comprehensive score.
7. An intelligent system for identifying surface defects of autoclaved aerated concrete components, characterized in that: For implementing the method for intelligently identifying surface defects of autoclaved aerated concrete components according to any one of claims 1 to 6, the intelligent system for intelligently identifying surface defects of autoclaved aerated concrete components comprises: A mapping module is used to perform omnidirectional photography and spatial transformation matrix mapping of the surface of the autoclaved aerated concrete component to obtain a pre-processed image; A feature extraction module is used to extract texture features from the preprocessed image to obtain texture feature data; A feature analysis module, configured to perform spatial feature analysis on the preprocessed image and the texture feature data to obtain spatial relationship feature data; The cross fusion module is used to perform feature importance weighting and bidirectional cross fusion processing on the texture feature data and the spatial relationship feature data to obtain a comprehensive defect feature representation, including: processing the texture feature data and the spatial relationship feature data respectively through global average pooling and multi-layer perceptron to obtain a texture feature importance score and a spatial feature importance score; based on the texture feature importance score and the spatial feature importance score, performing a weighted operation on the texture feature data and the spatial relationship feature data to obtain an initial weighted fusion feature; using the texture feature data as a query item and the spatial relationship feature data as a key-value pair to calculate the correlation between features; The similarity matrix is normalized to obtain a texture-to-space attention map; the spatial relationship feature data is used as a query item and the texture feature data as a key-value pair, the similarity matrix between the features is calculated and normalized to obtain a space-to-texture attention map; the spatial relationship feature data is weighted summed based on the texture-to-space attention map, and the texture feature data is weighted summed based on the space-to-texture attention map to obtain a bidirectional cross-attention feature; the initial weighted fusion feature is connected with the bidirectional cross-attention feature in the channel dimension, and the autoclaved aerated concrete material constraint is introduced to obtain a comprehensive defect feature representation; The output module is used to perform surface defect analysis on the autoclaved aerated concrete component based on the comprehensive defect feature representation, and output a comprehensive surface quality score and grade classification result.
Citation Information
Patent Citations
Texture surface defect detection method and system
CN110969606A
Lead screw module surface defect online detection method and device based on machine vision
CN119090862A