Ceramic tile laying position detection system and method

By combining the improved residual network and the Conformer network, combined with adaptive convolution and sparse self-attention layer, the problems of low efficiency and low accuracy in tiles laying detection are solved, and efficient and real-time tiles laying error detection and correction are achieved, and construction quality is improved.

CN120235841AInactive Publication Date: 2025-07-01GUANGDONG BUILDING DECORATION GRP CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357211.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as low detection efficiency, low accuracy and difficulty in adapting to complex environments during the tiles laying process. In particular, traditional image processing methods have low accuracy under multi-angle and multi-light conditions, deep learning models are difficult to balance in computing complexity and real-time, and incremental learning is insufficient in large-scale data flows.

Method used

Combining the improved residual network and the improved Conformer network for feature extraction, the adaptive convolution layer and the sparse self-attention layer are used to process tiles of different sizes and details levels, and the network parameters are optimized through incremental learning, and the model is adjusted in real time to adapt to changes in the construction environment.

Benefits of technology

It realizes high-precision and real-time tiles laying inspection in complex construction environments, and can promptly detect and correct errors, improve construction quality and efficiency, and reduce calculation complexity and overfitting risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235841A_ABST
    Figure CN120235841A_ABST
Patent Text Reader

Abstract

The invention discloses a tile laying position detection system and method. The method comprises the following steps: S1, collecting a multi-angle image of a tile laying site; s2, the multi-angle image is preprocessed; s3, performing feature extraction by using a residual network to generate a multi-dimensional feature map; s4, inputting the multi-dimensional feature map into an improved Conformer network, extracting deep feature information of the ceramic tiles, and extracting position coordinates, angles and alignment state information of the ceramic tiles; s5, the deviation between the tile laying position and the preset requirement is judged, and an error value and a deviation type are generated; s6, outputting the error value and the deviation type to monitoring end equipment; s7, continuously collecting images of the tile laying site, and generating a real-time detection result; and S8, the improved Conform network is optimized by using the improved incremental learning. According to the method, the residual error network is combined, the Conform network and incremental learning are improved, accurate detection of the tile laying position is realized, and the method has the advantages of high precision, strong real-time performance and good adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and particularly to a detection system and method for the laying position of tiles. Background Art

[0002] With the rapid development of the modern construction industry, especially in tile laying projects, the requirements for laying quality are getting higher and higher. Most traditional tile laying techniques rely on manual experience and visual inspection for position alignment and inspection, but this method has many limitations. First of all, manual inspection is not only inefficient but also extremely vulnerable to subjective factors, resulting in unstable tile laying quality, with position deviations, angular errors, and inaccurate alignment. Secondly, due to differences in different construction environments and construction workers, errors in the tile laying process are often difficult to detect and correct in a timely manner, thus affecting the overall decoration quality. In order to improve the accuracy in the tile laying process, many construction companies have tried to introduce automated detection technologies, but most of the existing technologies can only meet some basic requirements and cannot handle complex construction scenarios.

[0003] In the prior art, the most common tile laying detection method is based on traditional image processing techniques. The on-site image is collected by a camera device, and then basic image recognition algorithms, such as edge detection, template matching, etc., are applied to locate the tiles in the image. However, these traditional technologies usually rely on manually designed feature extraction methods, and are prone to feature extraction failure or low accuracy when facing changes such as complex tile patterns, multiple angles, and multiple illuminations. For example, traditional edge detection methods often cannot effectively identify the edges and corners of tiles in an environment with low contrast or complex textures, thus affecting the overall positioning accuracy. In addition, traditional image-based detection methods usually cannot handle the tiny position errors that may occur in tile laying, resulting in difficulty in accurately judging and correcting errors even if there are errors through traditional algorithms.

[0004] In addition, deep learning techniques have made remarkable progress in the field of image recognition in recent years. In particular, convolutional neural networks and deep networks have demonstrated superior capabilities in feature extraction and pattern recognition. However, in the detection of tile laying positions, traditional deep learning models still have many problems. First, the training of deep neural networks requires a large amount of labeled data. Especially in the construction field, image annotation is very time-consuming and expensive, which limits the wide application of deep learning techniques. Second, existing deep learning networks often fail to provide stable performance in various complex environments. Although convolutional neural networks perform well in processing fixed image sizes, for tile laying images with multiple angles and scales, the setting of the size of the convolutional kernel and the receptive field is still difficult to handle images with different sizes and levels of detail. Furthermore, deep neural networks also have significant problems in terms of computational complexity. Especially when performing multi-layer feature extraction, the computational amount of the model increases sharply, making real-time processing a major challenge.

[0005] To overcome these problems, in recent years, technologies such as improved convolutional neural networks and self-attention mechanisms have gradually been applied to the detection of tile laying. Improved convolutional neural networks (such as residual networks) effectively alleviate the problems of gradient disappearance and network degradation in deeper networks by introducing residual block structures. However, traditional residual networks still have problems such as a large number of parameters and high computational complexity. Especially in large-scale construction sites where real-time requirements are higher, the balance between the computational speed and accuracy of the network has always been a difficult problem. In addition, although convolutional neural networks can extract high-dimensional features from images, they still show poor robustness when dealing with complex textures, lighting changes, and dynamic scenes.

[0006] In addition, in current technologies, incremental learning is considered an effective way to solve the problems of training and updating deep learning models. Incremental learning can continuously adjust and optimize the model when new data arrives by updating model parameters in real time, thereby improving the adaptability of the model in a changing environment. However, incremental learning faces multiple challenges in practical applications, such as the selection of training data, the stability of the model, and the prevention of overfitting. Existing incremental learning methods often fail to quickly adapt and make effective adjustments when facing large-scale data streams, which makes real-time tile laying detection difficult to achieve. Summary of the Invention

[0007] An object of the present invention is to propose a method for detecting tile laying positions. The present invention automatically extracts various feature information during the tile laying process, including position coordinates, angles, and alignment states, and performs effective detection by combining an improved residual network and an improved Conformer network. At the same time, the present invention uses an incremental learning method to continuously optimize the network, adjusts network parameters based on real-time feedback, and improves the accuracy and robustness of network detection.

[0008] A method for detecting the laying position of tiles according to an embodiment of the present invention includes the following steps:

[0009] S1. Collect multi-angle images of the tile laying site through an image acquisition device;

[0010] S2. Preprocess the multi-angle images to generate preprocessed image data;

[0011] S3. Use a residual network for feature extraction. The residual network extracts the edges, corners, and surface textures of the tiles from the preprocessed image data through a multi-layer residual block structure to generate a multi-dimensional feature map;

[0012] S4. Input the multi-dimensional feature map into an improved Conformer network to extract deep feature information of the tiles. The improved Conformer network combines an adaptive convolutional layer and a sparse self-attention layer, adjusts the size of the convolutional kernel according to the multi-dimensional feature map, processes tile images of different sizes and detail levels. The sparse self-attention layer limits the range of attention calculation. The multi-level feature fusion layer fuses the outputs of the adaptive convolutional layer and the sparse self-attention layer to generate a fused feature map, and extracts the position coordinates, angles, and alignment status information of the tiles;

[0013] S5. Perform error analysis on the fused feature map, and judge the deviation between the tile laying position and the predetermined requirements according to the position coordinates, angles, and alignment status information of the tiles to generate an error value and a deviation type;

[0014] S6. Output the error value and the deviation type to the monitoring terminal device to display the error position of the tile laying;

[0015] S7. Continuously collect images of the tile laying site through the image acquisition device, input the real-time collected images into the residual network and the improved Conformer network for feature map extraction, and perform tile laying position detection, update the tile laying status, and generate a real-time detection result;

[0016] S8. Continuously optimize the improved Conformer network using improved incremental learning, feedback each detection result to the training data set of the improved Conformer network, and introduce a regularization term into the objective function through the elastic weight consolidation method to constrain the parameters of the improved Conformer network and optimize the parameters of the improved Conformer network.

[0017] Optionally, the S3 specifically includes:

[0018] S31. Input the preprocessed image data into a residual network. The residual network adopts a multi-layer residual block structure, and each residual block includes multiple convolutional layers, followed by a batch normalization layer and a ReLU activation function layer.

[0019] S32. Let the input preprocessed image data be , and perform a convolution operation through the convolution kernel at the position to obtain a local feature map. The convolution operation uses convolution kernels of different sizes, and the size of the convolution kernel is , where represents the size of the convolution kernel. Each convolutional layer generates a set of local feature response values by performing a sliding window operation on the input image:

[0020] ;

[0021] where, represents the local feature map obtained after the convolution operation, represents the input preprocessed image data at the position in and represent the positions of the input preprocessed image data, represents the element of the convolution kernel, and represent the index values of the convolution operation, represents the convolution kernel, represents the size of the convolution kernel, represents the position of the input preprocessed image data;

[0022] S33. Normalize the output of each convolutional layer through a batch normalization layer:

[0023] ;

[0024] where, represents the local feature map passing through the batch normalization layer, represents the mean value of the output of the convolutional layer, represents the standard deviation of the output of the convolutional layer, represents a smoothing factor to prevent division by zero errors;

[0025] S34. Through the output of each residual block, perform a skip connection in the residual network. The output of each layer in the residual network is added to the input of the previous layer. In the output of the residual block, Indicates the number of the current residual block:

[0026] ;

[0027] Among them, represents the output of the layer in the residual network, represents performing a convolution operation on the output of the previous layer, represents the non-linear activation function, represents the output of the layer in the residual network, represents the weights of the convolutional layer of the

[0028] S35. After processing through several residual blocks, a multi-dimensional feature map is output, and in the multi-dimensional feature map represents the total number of layers of the residual network. The final feature map contains the tile edges, corners, and surface textures extracted from the input image, and the size of the feature map is where represents the number of channels of the feature map, and represent the height and width of the image respectively.

[0029] Optionally, the S4 specifically includes:

[0030] S41. Input the multi-dimensional feature map into the improved Conformer network, which includes an adaptive convolutional layer, a sparse self-attention layer, and a multi-level feature fusion layer;

[0031] S42. Process the multi-dimensional feature map through the adaptive convolutional layer, and dynamically adjust the size of the convolutional kernel according to the detail complexity of the input image :

[0032] ;

[0033] ;

[0034] Among them, represents the calculated complexity metric of the multi-dimensional feature map, represents the second-order gradient of the multi-dimensional feature map at the position, represents the non-linear adjustment factor, represents the adjusted hyperparameter, represents the natural exponential function, represents the size of the convolutional kernel;

[0035] When performing convolution processing on the input feature map, the convolution operation is adjusted according to the adaptive convolution kernel size as follows:

[0036] ;

[0037] where, represents the feature map output by the adaptive convolution layer, represents the adaptive convolution kernel size, represents the elements of the convolution kernel, and represent the index values of the convolution operation, represents the input feature map at position ;

[0038] S44. Input the feature map processed by the adaptive convolution layer into the sparse self-attention layer. The sparse self-attention layer uses the dynamic sparse update matrix to limit the calculation range of the attention mechanism through the adaptive sparsity strategy and calculate the attention matrix at time :

[0039] ;

[0040] where, represents the attention matrix at time , represents the normalization operation, represents the query vector, represents the transpose of the key vector, represents the dimension of the query and key vectors, represents the key vector;

[0041] S45. Through the sparsity strategy, only calculate the attention in the local area to generate the sparsification matrix:

[0042] ;

[0043] where, represents the sparsification matrix, represents the sparsification factor, which controls the intensity of the sparsity, represents the number of non-zero elements in the attention matrix ;

[0044] S46. Generate the feature map processed by the sparse self-attention layer according to the sparsification matrix and the value vector:

[0045] ;

[0046] where, Represents the feature map after being processed by the sparse self-attention layer, represents the value vector;

[0047] S47. Through the multi-level feature fusion layer, the output of the adaptive convolutional layer and the output of the sparse self-attention layer are weighted and fused:

[0048] ;

[0049] Among them, represents the feature map after weighted fusion, and respectively represent the weighted coefficients of the output of the adaptive convolutional layer and the output of the sparse self-attention layer:

[0050] ;

[0051] ;

[0052] S48. Analyze the fused feature map, and extract the position coordinates, angle, and alignment status information of the tiles from the fused feature map.

[0053] Optionally, the S5 specifically includes:

[0054] S51. Conduct error analysis on the position coordinates, angle, and alignment status information of the tiles extracted from the fused feature map. The fused feature map includes the position coordinates, angle, and alignment status information of the tiles. Compare the position error, angle error, and alignment error with the predetermined requirements through the actual position , angle and alignment status of the tiles;

[0055] S52. According to the actual position and the preset target position of the tiles, calculate the position error through the non-linear error metric function:

[0056] ;

[0057] Among them, represents the position error, represents the natural exponential function, and represent the adjustment parameters, controlling the non-linearity degree of the error function, represents the actual position of the tile, represents the preset target position of the tile;

[0058] S53. Calculate the angle error based on the actual angle of the tile and the preset target angle , and the calculation formula is:

[0059] ;

[0060] wherein, represents the angle error, represents the arctangent function, which is used to normalize the angle error, represents the preset target angle of the tile, represents the actual angle of the tile;

[0061] S54. Calculate the alignment error based on the actual alignment state of the tile and the preset alignment state , and the calculation formula is:

[0062] ;

[0063] wherein, represents the alignment error, represents the adjustment factor of the alignment error, represents the actual alignment state of the tile, represents the preset alignment state of the tile;

[0064] S55. Calculate the comprehensive error evaluation value based on the position error, angle error and alignment error. The total error evaluation value is the weighted sum of the position error, angle error and alignment error;

[0065] S56. Judge the quality of the tile laying according to the comprehensive error evaluation value. If the comprehensive error evaluation value is less than the preset threshold, it is considered that the tile laying position meets the standard. Otherwise, it is considered that there is an error and an error type is generated.

[0066] Optionally, the specific steps of S8 include:

[0067] S81. Continuously optimize the improved Conformer network through improved incremental learning, store the error information into the incremental training set , and store and update the incremental data. The goal of the improved incremental learning is to update the parameters of the improved Conformer network according to the real-time feedback data. The improvement points of the improved incremental learning include sampling the incremental data through the experience replay strategy and adopting the elastic weight consolidation method to prevent the improved Conformer network from overfitting during training:

[0068] ;

[0069] wherein, represents the set containing incremental data The updated training set, represents the incremental training set, represents "union", represents the incremental data, which is the error information of the real-time detection result;

[0070] S82. Sample the incremental data through the experience replay strategy, and the sampling is dynamically adjusted based on the incremental data as follows:

[0071] ;

[0072] Among them, represents the incremental data is the probability of being sampled, represents the natural exponential function, represents the adjustment factor, which controls the influence of the error magnitude on the sampling probability, represents the normalization constant, which ensures that the sum of all sampling probabilities is 1, represents the incremental data is the norm of;

[0073] S83. Input the sampled incremental data set into the improved Conformer network for training, and update the parameters of the improved Conformer network through the backpropagation algorithm:

[0074] ;

[0075] Among them, represents the updated parameter, represents the current parameter of the improved Conformer network, represents the learning rate, represents the gradient of the loss function with respect to the parameter of the improved Conformer network, represents the loss function;

[0076] S84. To prevent overfitting, the elastic weight consolidation method is adopted to constrain the parameters of the improved Conformer network by introducing a regularization term into the objective function:

[0077] ;

[0078] Among them, represents the updated parameter, represents the current parameter of the improved Conformer network, represents the learning rate, represents the gradient of the loss function with respect to the parameter The gradient, represents the loss function, represents the elastic weight consolidation coefficient, which controls the influence of the regularization term on the parameters, represents the L2 norm of the parameters of the improved Conformer network:

[0079] S85. After training with improved incremental learning and elastic weight consolidation, optimized parameters are obtained , and the optimized network parameters are used for detecting the real-time tile laying position.

[0080] A tile laying position detection system according to an embodiment of the present invention includes:

[0081] An image acquisition device for acquiring multi-angle images of the tile laying site;

[0082] A preprocessing module for denoising, color normalization, and image enhancement of the acquired multi-angle images to generate preprocessed image data;

[0083] A feature extraction module for extracting the edges, corners, and surface textures of the tiles from the preprocessed images through a residual network to generate a multi-dimensional feature map;

[0084] A deep feature extraction module for inputting the feature map into the improved Conformer network to extract the position coordinates, angles, and alignment status information of the tiles;

[0085] An error analysis module for analyzing the laying position, angles, and alignment status of the tiles, calculating the error value, and generating the deviation type;

[0086] An output module for outputting the error value and deviation type to the monitoring terminal device to display the error position of the tile laying;

[0087] A real-time detection module for continuously acquiring images and inputting them into the network for real-time detection, updating the tile laying status, and generating real-time detection results;

[0088] An incremental learning module for optimizing the parameters of the improved Conformer network through real-time detection feedback, improving the detection accuracy, and applying the elastic weight consolidation method to avoid overfitting in the training of the improved Conformer network.

[0089] The beneficial effects of the present invention are:

[0090] First, the present invention utilizes deep learning techniques, particularly residual networks and improved Conformer networks, to effectively overcome the problem in traditional image processing methods of being unable to accurately extract the details of tile laying. Through the multi-layer residual block structure in the residual network, it is able to extract detailed features such as the edges, corners, and surface textures of tiles at a deeper level, while avoiding the phenomena of gradient disappearance and network degradation. This enables the network to still extract accurate feature information even in complex and dynamic construction environments, providing a more precise basis for subsequent error analysis and detection.

[0091] Secondly, the application of the adaptive convolutional layer and the sparse self-attention layer in the improved Conformer network enables the system to dynamically adjust the size of the convolutional kernel according to the multi-dimensional feature map and process tile images of different sizes and detail levels. Through adaptive convolutional operations, the feature extraction strategy can be adjusted according to the complexity of the image, reducing the consumption of computing resources and improving the balance between real-time performance and computational complexity of the system. The sparse self-attention layer, on the other hand, effectively reduces the computational complexity by restricting the scope of attention calculation, while ensuring the high-efficiency adaptability of the model in changing construction environments. This technical improvement enables the system to accurately detect the laying position, angle, and alignment state of tiles under low latency and achieve real-time feedback.

[0092] Finally, the introduction of incremental learning enables the network to continuously optimize according to the feedback data during the real-time construction process. Each detection result can be fed back as incremental data to the training set, thereby continuously updating and improving the parameters of the network, making the model always maintain good performance under different lighting conditions, tile types, and construction environments. By combining with the elastic weight consolidation method, the situation of network overfitting is avoided, further improving the stability and generalization ability of the model in long-term use.

[0093] In summary, the present invention not only improves the accuracy and efficiency of tile laying position detection through the innovative combination of deep learning, adaptive convolution, sparse self-attention mechanism, and incremental learning, but also enables real-time optimization and feedback during the actual construction process. This technical solution effectively solves the deficiencies in traditional methods, especially the robustness in multi-angle, multi-scale, and dynamic environments, enabling the errors in the tile laying process to be detected and corrected in a timely manner, thereby ensuring the improvement of the final laying quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0095] Figure 1 is a flowchart of a method for detecting the laying position of tiles proposed by the present invention;

[0096] Figure 2 Schematic diagram of the structure of the improved Conformer network for the detection method of the laying position of tiles proposed by the present invention. Specific implementation mode

[0097] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0098] Reference Figure 1 And Figure 2 , a detection method for the laying position of tiles, comprising the following steps:

[0099] S1. Collect multi-angle images of the tile laying site through an image acquisition device;

[0100] S2. Preprocess the multi-angle images to generate preprocessed image data;

[0101] S3. Use a residual network for feature extraction. The residual network extracts the edges, corners, and surface textures of the tiles from the preprocessed image data through a multi-layer residual block structure to generate a multi-dimensional feature map;

[0102] S4. Input the multi-dimensional feature map into the improved Conformer network to extract the deep feature information of the tiles. The improved Conformer network combines an adaptive convolutional layer and a sparse self-attention layer, adjusts the size of the convolutional kernel according to the multi-dimensional feature map, processes tile images of different sizes and detail levels. The sparse self-attention layer limits the range of attention calculation, and the multi-level feature fusion layer fuses the outputs of the adaptive convolutional layer and the sparse self-attention layer to generate a fused feature map, and extracts the position coordinates, angles, and alignment status information of the tiles;

[0103] S5. Perform error analysis on the fused feature map, and judge the deviation between the tile laying position and the predetermined requirements according to the position coordinates, angles, and alignment status information of the tiles to generate an error value and a deviation type;

[0104] S6. Output the error value and the deviation type to the monitoring terminal device to display the error position of the tile laying;

[0105] S7. Continuously collect images of the tile laying site through the image acquisition device, input the real-time collected images into the residual network and the improved Conformer network for feature map extraction, and perform the detection of the tile laying position, update the tile laying status, and generate a real-time detection result;

[0106] S8. Continuously optimize the improved Conformer network using improved incremental learning, feedback each detection result to the training dataset of the improved Conformer network, and introduce a regularization term into the objective function to constrain the parameters of the improved Conformer network by the elastic weight consolidation method, thereby optimizing the parameters of the improved Conformer network.

[0107] In this embodiment, the specific steps of S3 include:

[0108] S31. Input the preprocessed image data into the residual network, which adopts a multi-layer residual block structure. Each residual block includes multiple convolutional layers, followed by a batch normalization layer and a ReLU activation function layer;

[0109] S32. Assume that the input preprocessed image data is , perform a convolution operation through the convolution kernel at the position to obtain a local feature map. The convolution operation uses convolution kernels of different sizes, and the size of the convolution kernel is , where represents the size of the convolution kernel. Each convolutional layer generates a set of local feature response values by performing a sliding window operation on the input image:

[0110] ;

[0111] Among them, represents the local feature map obtained after the convolution operation, represents the input preprocessed image data at the position of the pixel value, and represent the position of the input preprocessed image data, represents the element of the convolution kernel, and represent the index values of the convolution operation, represents the convolution kernel, represents the size of the convolution kernel, represents the position of the input preprocessed image data;

[0112] S33. The output of each convolutional layer is normalized through a batch normalization layer:

[0113] ;

[0114] Among them, represents the local feature map passing through the batch normalization layer, represents the mean value of the output of the convolutional layer, Represents the output of the convolutional layer , the standard deviation of represents a smoothing factor to prevent division-by-zero errors;

[0115] S34. Through the output of each residual block , perform skip connections in the residual network, and the output of each layer in the residual network is added to the input of the previous layer . Among the outputs of the residual block represents the number of the current residual block:

[0116] ;

[0117] Among them, represents the output of the layer in the residual network, represents performing a convolution operation on the output of the previous layer, represents a non-linear activation function, represents the output of the layer in the residual network, represents the weights of the convolutional layer of the

[0118] S35. After processing through several residual blocks, output a multi-dimensional feature map . Among the multi-dimensional feature map represents the total number of layers of the residual network. The final feature map contains the tile edges, corners, and surface textures extracted from the input image. The size of the feature map is , where represents the number of channels of the feature map, and represent the height and width of the image respectively.

[0119] In this embodiment, the S4 specifically includes:

[0120] S41. Input the multi-dimensional feature map into the improved Conformer network, and the improved Conformer network includes an adaptive convolutional layer, a sparse self-attention layer, and a multi-level feature fusion layer;

[0121] S42. Process the multi-dimensional feature map through the adaptive convolutional layer, and dynamically adjust the size of the convolutional kernel according to the detail complexity of the input image :

[0122] ;

[0123] ;

[0124] Among them, represents the calculated complexity metric of the multi-dimensional feature map, represents the second-order gradient of the multi-dimensional feature map at the position, represents the non-linear adjustment factor, represents the adjusted hyperparameter, represents the natural exponential function, represents the size of the convolutional kernel;

[0125] S43. When performing convolution processing on the input feature map, the convolution operation is adjusted according to the adaptive convolutional kernel size as follows:

[0126] ;

[0127] where represents the feature map output by the adaptive convolutional layer, represents the adaptive convolutional kernel size, represents the element of the convolutional kernel, and represent the index value of the convolution operation, represents at the position the input feature map;

[0128] S44. Input the feature map processed by the adaptive convolutional layer into the sparse self-attention layer. The sparse self-attention layer uses the dynamic sparse update matrix , restricts the calculation range of the attention mechanism through the adaptive sparsity strategy, and calculates the attention matrix at time as follows:

[0129] ;

[0130] where represents the attention matrix at time , represents the normalization operation, represents the query vector, represents the transpose of the key vector, represents the dimension of the query and key vectors, represents the key vector;

[0131] S45. Through the sparsity strategy, only calculate the attention in the local area to generate the sparsification matrix:

[0132] ;

[0133] where represents the sparsification matrix, represents the sparsification factor, which controls the intensity of the sparsity, Represents the attention matrix The number of non-zero elements in;

[0134] S46. Generate a feature map processed by the sparse self-attention layer according to the sparsified matrix and the value vector:

[0135] ;

[0136] Wherein, Represents the feature map processed by the sparse self-attention layer, Represents the value vector;

[0137] S47. Perform weighted fusion on the output of the adaptive convolutional layer And the output of the sparse self-attention layer :

[0138] ;

[0139] Wherein, Represents the feature map after weighted fusion, And Respectively represent the weighted coefficients of the output of the adaptive convolutional layer and the output of the sparse self-attention layer:

[0140] ;

[0141] ;

[0142] S48. Analyze the fused feature map and extract the position coordinates, angle, and alignment status information of the tile from the fused feature map.

[0143] In this embodiment, the S5 specifically includes:

[0144] S51. Perform error analysis on the position coordinates, angle, and alignment status information of the tile extracted from the fused feature map. The fused feature map Includes the position coordinates, angle, and alignment status information of the tile. Compare the position error, angle error, and alignment error with the predetermined requirements through the actual position , angle And alignment status ;

[0145] S52. Calculate the position error According to the actual position of the tile And the preset target position :

[0146] ;

[0147] Among them, represents the position error, represents the natural exponential function, and represents the adjustment parameter, which controls the nonlinear degree of the error function, represents the actual position of the tile, represents the preset target position of the tile;

[0148] S53. According to the actual angle and the preset target angle of the tile, calculate the angle error:

[0149] ;

[0150] Among them, represents the angle error, represents the arctangent function, which is used to normalize the angle error, represents the preset target angle of the tile, represents the actual angle of the tile;

[0151] S54. According to the actual alignment state and the preset alignment state of the tile, calculate the alignment error:

[0152] ;

[0153] Among them, represents the alignment error, represents the adjustment factor of the alignment error, represents the actual alignment state of the tile, represents the preset alignment state of the tile;

[0154] S55. According to the position error, angle error and alignment error, calculate the comprehensive error evaluation value, and the total error evaluation value is the weighted sum of the position error, angle error and alignment error;

[0155] S56. Judge the quality of tile laying according to the comprehensive error evaluation value. If the comprehensive error evaluation value is less than the preset threshold, it is considered that the tile laying position meets the standard. Otherwise, it is considered that there is an error, and an error type is generated.

[0156] In this embodiment, the specific content of S8 includes:

[0157] S81. Continuously optimize the improved Conformer network through improved incremental learning, and store the error information into the incremental training set , and store and update the incremental data. The goal of the improved incremental learning is to update the parameters of the improved Conformer network according to the real-time feedback data. The improvements of the improved incremental learning include sampling the incremental data through an experience replay strategy and adopting an elastic weight consolidation method to prevent the improved Conformer network from overfitting during training:

[0158] ;

[0159] Among them, represents the updated training set containing the incremental data . represents the incremental training set, represents "and", represents the incremental data, which is the error information of the real-time detection result;

[0160] S82. Sample the incremental data through an experience replay strategy, and the sampling is dynamically adjusted based on the incremental data :

[0161] ;

[0162] Among them, represents the probability that the incremental data is sampled, represents the natural exponential function, represents the adjustment factor that controls the influence of the error magnitude on the sampling probability, represents the normalization constant that ensures the sum of all sampling probabilities is 1, represents the incremental data .

[0163] S83. Input the sampled incremental data set into the improved Conformer network for training, and update the parameters of the improved Conformer network through the backpropagation algorithm:

[0164] ;

[0165] Among them, represents the updated parameters, represents the current parameters of the improved Conformer network, represents the learning rate, represents the gradient of the loss function with respect to the parameters of the improved Conformer network, represents the loss function;

[0166] S84. To prevent overfitting, the elastic weight consolidation method is adopted. By introducing a regularization term into the objective function, the parameters of the improved Conformer network are constrained:

[0167] ;

[0168] Among them, represents the updated parameter, represents the current parameter of the improved Conformer network, represents the learning rate, represents the gradient of the loss function with respect to the parameter of the improved Conformer network, represents the loss function, represents the elastic weight consolidation coefficient, which controls the influence of the regularization term on the parameters, represents the L2 norm of the parameters of the improved Conformer network:

[0169] S85. After training with improved incremental learning and elastic weight consolidation, the optimized parameter is obtained, and the optimized network parameters are used for the detection of the real-time tile laying position.

[0170] A tile laying position detection system includes:

[0171] An image acquisition device for acquiring multi-angle images of the tile laying site;

[0172] A preprocessing module for denoising, color normalization, and image enhancement of the acquired multi-angle images to generate preprocessed image data;

[0173] A feature extraction module for extracting the edges, corners, and surface textures of the tiles from the preprocessed images through a residual network to generate a multi-dimensional feature map;

[0174] A deep feature extraction module for inputting the feature map into the improved Conformer network to extract the position coordinates, angles, and alignment status information of the tiles;

[0175] An error analysis module for analyzing the laying position, angle, and alignment status of the tiles, calculating the error value, and generating the deviation type;

[0176] An output module for outputting the error value and deviation type to the monitoring terminal device to display the error position of the tile laying;

[0177] A real-time detection module for continuously acquiring images and inputting them into the network for real-time detection, updating the tile laying status, and generating real-time detection results;

[0178] An incremental learning module for optimizing and improving the parameters of the Conformer network through real-time detection feedback, enhancing the detection accuracy, and applying the elastic weight consolidation method to avoid overfitting in the training of the improved Conformer network.

[0179] Example 1:

[0180] To verify the feasibility of the present invention in implementation, the present invention is applied to the tile laying quality inspection project in an actual project of a construction company. This project is located in the decoration project of a newly built commercial office building in a certain city. The construction party is a well-known construction company in this area. The project is large in scale, involving multiple floors and a large number of tile laying tasks. Tile laying occupies an important position in construction, and its quality is directly related to the overall decoration effect and the maintenance cost in later use. Although the traditional manual laying inspection method is acceptable in small-scale projects, for large-scale projects, it is not only inefficient but also prone to serious quality problems due to human oversight.

[0181] At the initial stage of this project, the construction company adopted the traditional manual detection method, relying on workers to manually measure and visually confirm whether the tile laying is compliant. However, many problems have emerged in actual application. First, the manual inspection is inefficient. Usually, a worker can only inspect about 50 square meters of tile laying area per day. And due to visual fatigue and subjective factors, some small errors (such as the angular deviation and misalignment of tiles) are difficult to detect in a timely manner. Second, the errors during the laying process cannot be timely feedback and corrected, resulting in continuous problems in the overall tile laying effect during the project progress. Workers have to repeatedly rework, causing waste of time and cost. Therefore, in response to this problem, the construction company decided to try to use the tile laying position detection method in the present invention for improvement and expected to improve the quality inspection efficiency through intelligent detection.

[0182] In actual application, the project party first collected multi-angle on-site images of tile laying through a professional image acquisition device. This image acquisition device is an intelligent camera with a resolution of up to 4K, and it is combined with a three-dimensional scanning device to collect images at every corner of the construction site to ensure that the images cover all laying areas. The collected data is processed by the image preprocessing module for denoising, color normalization, and enhancement to ensure the image quality. Then, the image data is sent to the feature extraction module in the deep learning model, and feature extraction is performed through the residual network and the improved Conformer network, and then a high-dimensional feature map is generated. The feature map contains information such as the edges, corners, surface textures of the tiles, as well as the specific positions and angles of the laying, and this information will be used for subsequent error analysis and detection.

[0183] After a series of analysis and processing by deep learning algorithms, the present invention can generate real-time information on the laying status of each tile, including position coordinates, angles, and whether they are aligned. These data are compared with the preset standard requirements, and the detection system can accurately judge the deviation of tile laying. For example, due to construction environment problems, some tiles may have angular deviations or misalignments. The detection method of the present invention can detect these errors in real time during the tile laying process and generate error feedback, indicating the specific position and deviation type, thereby helping construction workers make immediate corrections and avoiding a large number of reworks.

[0184] During the implementation process, the incremental learning mechanism of the present invention also demonstrated its powerful advantages. In the initial stage, the deep learning model trained by big data can quickly adapt to various on-site environmental changes, including factors such as lighting and differences in tile surface textures. As the construction progress advances, the system performs incremental learning by continuously feeding back new detection data, optimizing the detection accuracy and making the detection results gradually tend to be accurate. Especially in complex construction sites, the present invention can adapt to various types of tiles and environmental conditions, ensuring the smooth progress of the construction.

[0185] Table 1 Experimental comparison data table

[0186] Inspection method Inspection time (days) Inspection area (square meters) Average position error (mm) Average angular error (degrees) Error correction feedback rate (%) Manual inspection 5 200 ±2 ±0.5 30% The method of the present invention 5 4500 ±1 ±0.2 90%

[0187] From the comparison of inspection methods, we can see that there is a significant gap in inspection efficiency between traditional manual inspection and the intelligent detection method of the present invention. The traditional manual inspection method can only inspect an area of about 50 square meters per worker per day, and due to the limitations of manual inspection, it may not be able to detect small errors in a timely manner, resulting in a large number of reworks. In contrast, the intelligent detection method of the present invention, through automated deep learning algorithms and real-time feedback mechanisms, greatly increases the daily inspection area to 4500 square meters and can generate detection results in real time, reducing the time of manual intervention.

[0188] In terms of the comparison of inspection time, the traditional manual inspection method takes 5 days to complete the inspection of a 200-square-meter area, while the intelligent detection method of the present invention can complete the detection of a large-scale area in the same time, improving the construction efficiency and avoiding human delays and waste.

[0189] From the perspective of error analysis, the traditional manual inspection method has significant limitations in the error detection of tile laying. The average position error of manual inspection is ±2 mm, and the angle error is ±0.5 degrees. Due to the limited accuracy of manual inspection, some minor errors are often difficult to detect, which also leads to a large amount of rework and correction. In contrast, the intelligent detection method of the present invention can reduce the position error to ±1 mm and the angle error to ±0.2 degrees, improving the detection accuracy and ensuring the laying precision.

[0190] In terms of the error correction feedback rate, the feedback rate of the intelligent detection method of the present invention is 90%, which means that it can promptly identify and feedback most of the laying errors, and the construction workers can immediately make corrections to avoid the further expansion of errors. In contrast, the error correction feedback rate of the traditional manual inspection is only 30%, indicating that the ability of manual inspection to promptly feedback and correct errors is limited and difficult to meet the requirements of large-scale construction projects.

[0191] In summary, the tile laying position detection method of the present invention demonstrates significant advantages in multiple aspects, especially the improvement in inspection efficiency, error precision, and real-time feedback, which enhances the construction quality and efficiency and reduces the labor and time costs during the construction process.

[0192] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.

Claims

1. A method for detecting the laying position of tiles, characterized in that: The steps include: S1, collecting multi-angle images of the tile laying site through an image acquisition device; S2, preprocessing the multi-angle image to generate preprocessed image data; S3, using a residual network to perform feature extraction, wherein the residual network extracts the edges, corners and surface textures of the tiles from the preprocessed image data through a multi-layer residual block structure to generate a multi-dimensional feature map; S4, inputting the multidimensional feature map into the improved Conformer network to extract the deep feature information of the tile, the improved Conformer network combines the adaptive convolution layer and the sparse self-attention layer, adjusts the size of the convolution kernel according to the multidimensional feature map, and processes tile images of different sizes and detail levels, the sparse self-attention layer limits the scope of attention calculation, and the multi-level feature fusion layer fuses the outputs of the adaptive convolution layer and the sparse self-attention layer to generate a fused feature map, and extracts the position coordinates, angle and alignment status information of the tile; S5, performing error analysis on the fused feature map, and judging the deviation between the tile laying position and the predetermined requirement according to the position coordinates, angle and alignment status information of the tile, and generating an error value and a deviation type; S6, outputting the error value and deviation type to the monitoring terminal device to display the tile laying error position; S7, continuously collecting images of the tile laying site through an image acquisition device, and inputting the real-time collected images into the residual network and the improved Conformer network to extract feature maps, detect the tile laying position, update the tile laying status, and generate real-time detection results; S8. Use improved incremental learning to continuously optimize the improved Conformer network, feed back each detection result to the training data set of the improved Conformer network, and introduce regularization terms in the objective function through the elastic weight solidification method to constrain the parameters of the improved Conformer network and optimize the parameters of the improved Conformer network.

2. A method for detecting the laying position of ceramic tiles according to claim 1, characterized in that: The S3 specifically includes: S31, inputting the preprocessed image data into a residual network, wherein the residual network adopts a multi-layer residual block structure, each residual block includes multiple convolutional layers, and the convolutional layers are followed by a batch normalization layer and a ReLU activation function layer; S32, set the input pre-processed image data to , through the convolution kernel In Location A convolution operation is performed to obtain a local feature map. The convolution operation uses convolution kernels of different sizes. The size of the convolution kernel is ,in Represents the size of the convolution kernel. Each convolution layer generates a set of local feature response values ​​by performing a sliding window operation on the input image: ; in, Represents the local feature map obtained after the convolution operation, Represents the input preprocessed image data Middle position The pixel value at and represents the location of the input preprocessed image data, represents the elements of the convolution kernel, and Represents the index value of the convolution operation, represents the convolution kernel, represents the size of the convolution kernel, Indicates the location of input preprocessed image data; S33. The output of each convolutional layer is normalized through a batch normalization layer: ; in, represents the local feature map passed through the batch normalization layer, Represents the convolutional layer output The mean of Represents the convolutional layer output The standard deviation of represents the smoothing factor to prevent division by zero errors; S34, through the output of each residual block , skip connections are made in the residual network, and each layer in the residual network outputs With the input of the previous layer Perform an addition operation, the output of the residual block Indicates the number of the current residual block: ; in, Represents the residual network The output of the layer, Indicates the convolution operation on the output of the previous layer. represents a nonlinear activation function, Represents the residual network The output of the layer, express The weights of the convolutional layers; S35. After being processed by several layers of residual blocks, a multi-dimensional feature map is output , in the multidimensional feature map Represents the total number of layers of the residual network. The final feature map contains the tile edges, corners, and surface textures extracted from the input image. The size of the feature map is ,in Indicates the number of channels of the feature map, and Represents the height and width of the image respectively.

3. A method for detecting the laying position of ceramic tiles according to claim 1, characterized in that: The S4 specifically includes: S41, inputting the multidimensional feature map into an improved Conformer network, wherein the improved Conformer network includes an adaptive convolution layer, a sparse self-attention layer, and a multi-level feature fusion layer; S42, process the multi-dimensional feature map through the adaptive convolution layer, and dynamically adjust the size of the convolution kernel according to the detail complexity of the input image : ; ; in, represents the calculated multi-dimensional feature graph complexity measure, Represents a multidimensional feature map in The second-order gradient of position, represents the nonlinear adjustment factor, represents the hyperparameters to be adjusted, represents the natural exponential function, Indicates the size of the convolution kernel; S43, when performing convolution processing on the input feature map, the convolution operation is based on the adaptive convolution kernel size Make adjustments: ; in, represents the feature map output by the adaptive convolution layer, represents the adaptive convolution kernel size, represents the elements of the convolution kernel, and Represents the index value of the convolution operation, Indicates at location Input feature map of S44, input the feature map processed by the adaptive convolution layer into the sparse self-attention layer, and the sparse self-attention layer uses a dynamic sparse update matrix , through the adaptive sparsity strategy to limit the calculation range of the attention mechanism, the calculation time The attention matrix: ; in, Indicates time The attention matrix, represents the normalization operation, represents the query vector, represents the transpose of the key vector, represents the dimensions of the query and key vectors, represents the key vector; S45. Through the sparse strategy, only the attention of the local area is calculated to generate a sparse matrix: ; in, represents the sparse matrix, represents the sparsification factor, which controls the strength of sparsity. Represents the attention matrix The number of non-zero elements in ; S46. Generate a feature map processed by the sparse self-attention layer based on the sparse matrix and value vector: ; in, represents the feature map after being processed by the sparse self-attention layer, represents a value vector; S47, output the adaptive convolution layer through the multi-level feature fusion layer and the sparse self-attention layer output Perform weighted fusion: ; in, represents the feature map after weighted fusion, and Represent the weight coefficients of the adaptive convolution layer output and the sparse self-attention layer output respectively: ; ; S48, analyzing the fused feature map, and extracting the position coordinates, angle and alignment status information of the tile from the fused feature map.

4. A method for detecting the laying position of ceramic tiles according to claim 1, characterized in that: The S5 specifically includes: S51, performing error analysis on the position coordinates, angles and alignment status information of the tiles extracted from the fused feature map, wherein the fused feature map Including the location coordinates, angle and alignment status information of the tile, through the actual position of the tile ,angle and alignment status , compare with the predetermined requirements to determine the position error, angle error and alignment error; S52, according to the actual position of the tile and preset target location , calculate the position error through the nonlinear error metric function : ; in, represents the position error, represents the natural exponential function, and Represents the adjustment parameter, which controls the nonlinearity of the error function. represents the actual position of the tile, Indicates the preset target position of the tile; S53, according to the actual angle of the tile and preset target angle , calculate the angle error: ; in, represents the angle error, represents the inverse tangent function, which is used to normalize the angle error. Indicates the preset target angle of the tile, Indicates the actual angle of the tile; S54, according to the actual alignment state of the tiles and preset alignment state , calculate the alignment error: ; in, represents the alignment error, represents the adjustment factor for the alignment error, Indicates the actual alignment status of the tile, Indicates the preset alignment state of the tile; S55, calculating a comprehensive error evaluation value according to the position error, the angle error and the alignment error, wherein the total error evaluation value is a weighted sum of the position error, the angle error and the alignment error; S56. The quality of the tile laying is determined according to the comprehensive error evaluation value. If the comprehensive error evaluation value is less than a preset threshold, it is considered that the tile laying position meets the standard. Otherwise, it is considered that an error exists and an error type is generated.

5. A method for detecting the laying position of ceramic tiles according to claim 1, characterized in that: The S8 specifically includes: S81. Continuously optimize the improved Conformer network by improving incremental learning and converting error information Store incremental training set , and store and update incremental data. The goal of the improved incremental learning is to update the parameters of the improved Conformer network according to real-time feedback data. The improvements of the improved incremental learning include sampling incremental data through the experience replay strategy and using the elastic weight solidification method to prevent the improved Conformer network from overfitting during the training process: ; in, Indicates that incremental data is included The updated training set, represents the incremental training set, Indicates "and", Indicates incremental data, which is the error information of real-time detection results; S82, sampling the incremental data through the experience replay strategy, the sampling is based on the incremental data To make dynamic adjustments: ; in, Indicates incremental data The probability of being sampled, represents the natural exponential function, Represents the adjustment factor, which controls the effect of the error on the sampling probability. represents the normalization constant, ensuring that the sum of all sampling probabilities is 1, Indicates incremental data The norm of S83, the sampled incremental data set Input to the improved Conformer network for training, and update the parameters of the improved Conformer network through the back propagation algorithm: ; in, represents the updated parameters, Represents the current parameters of the improved Conformer network, represents the learning rate, Represents the loss function for improving the parameters of the Conformer network The gradient of represents the loss function; S84. To prevent overfitting, the elastic weight solidification method is used to constrain the parameters of the improved Conformer network by introducing a regularization term in the objective function: ; in, represents the updated parameters, Represents the current parameters of the improved Conformer network, represents the learning rate, Represents the loss function for improving the parameters of the Conformer network The gradient of represents the loss function, Represents the elastic weight solidification coefficient, which controls the influence of the regularization term on the parameters. Represents the L2 norm of the parameters of the improved Conformer network: S85, after training with improved incremental learning and elastic weight solidification, the optimized parameters are obtained ,The optimized network parameters are used for real-time tile laying position detection.

6. A system for detecting the laying position of ceramic tiles, which executes the method for detecting the laying position of ceramic tiles according to any one of claims 1 to 5, characterized in that: include: Image acquisition equipment, used to collect multi-angle images of the tile laying site; A preprocessing module is used to perform denoising, color normalization and image enhancement on the collected multi-angle images to generate preprocessed image data; A feature extraction module is used to extract the edges, corners and surface textures of tiles from the preprocessed image through a residual network to generate a multi-dimensional feature map; The deep feature extraction module is used to input the feature map into the improved Conformer network to extract the position coordinates, angles and alignment status information of the tiles; The error analysis module is used to analyze the laying position, angle and alignment of the tiles, calculate the error value, and generate the deviation type; The output module is used to output the error value and deviation type to the monitoring terminal device to display the tile laying error position; Real-time detection module, used to continuously collect images and input them into the network for real-time detection, update the tile laying status, and generate real-time detection results; The incremental learning module is used to optimize and improve the parameters of the Conformer network through real-time detection feedback, improve detection accuracy, and apply the elastic weight solidification method to avoid overfitting of the improved Conformer network training.