Ceramic tile hollowing identification method based on multi-modal data acquired by wall-climbing robot

By collecting multimodal data using a wall-climbing robot and employing a multimodal fusion model to identify hollow tiles, the problems of subjectivity, low efficiency, and environmental interference in existing detection methods are solved, achieving high-precision and robust automated detection.

CN121994919APending Publication Date: 2026-05-08HANGZHOU YUENTROPY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU YUENTROPY TECHNOLOGY CO LTD
Filing Date
2026-01-09
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing methods for detecting hollow tiles are subjective, inefficient, and unsafe. They cannot adapt to efficient automated testing, are severely affected by environmental interference, and single-modal data are easily affected by noise, resulting in insufficient robustness.

Method used

A wall-climbing robot was used to collect multimodal data, including ultrasound, acoustic and visual image data. Hollow areas were identified through a multimodal fusion model, and discrimination was performed using a Transformer fusion module and a Softmax classifier. Real-time screening was achieved by combining a lightweight ultrasound pre-screening model on the edge side.

Benefits of technology

It achieves high security and reproducibility in complex scenarios for detecting hollow tiles, improves detection accuracy and robustness, is suitable for automated detection in multiple scenarios, and provides reliable quantitative analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121994919A_ABST
    Figure CN121994919A_ABST
Patent Text Reader

Abstract

The invention discloses a ceramic tile hollowing identification method based on multi-modal data collected by a wall-climbing robot, and the method comprises the steps: carrying out the synchronous collection of the multi-modal data through a wall-climbing robot, and enabling the multi-modal data to comprise ultrasonic modal data, acoustic modal data and visual image modal data; preprocessing the multi-modal data, and constructing a multi-modal data set for training a constructed multi-modal fusion model; the multi-modal fusion model comprises an ultrasonic branch, an acoustic branch, a visual image branch and a Transform fusion module; and after the Transform fusion module fuses the output of the ultrasonic branch, the output of the acoustic branch and the output of the visual image branch, hollowing / non-hollowing discrimination is carried out by using a Softmax classifier. According to the method, ultrasonic, acoustics and vision multi-modal data are fused, a cloud edge collaborative intelligent architecture is combined, high-precision and high-efficiency automatic hollowing detection is realized, the traditional artificial experience is converted into a standardized digital process, the quantification, reproduction and tracing of a detection result are realized, and reliable data support is provided for building safety operation and maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tile hollow detection technology, and in particular to a method for identifying tile hollowness based on multimodal data collected by a wall-climbing robot. Background Technology

[0002] Hollow tiles on building exterior walls are a common building quality and safety hazard. Current detection methods have the following defects: (1) High subjectivity, low efficiency and questionable safety: They rely on the experience and hearing of the inspectors for judgment, which inevitably leads to subjective randomness, making it difficult to quantify, reproduce and trace the detection results. In addition, for high-rise buildings, scaffolding or suspended baskets are even required, resulting in extremely low detection efficiency, which can no longer meet the stringent requirements of efficiency and safety in modern building operation and maintenance. (2) Severely constrained by the environment and have detection blind spots: Although infrared thermal imaging has achieved non-contact, large-area rapid scanning, its physical mechanism determines its strong environmental dependence. External factors such as sunlight, wind speed, rainfall, wall material and humidity can seriously interfere with the heat conduction characteristics of the wall, which can easily lead to false alarms or missed alarms. For example, shadows, wall stains or material differences may appear as hot spots similar to hollow tiles on thermal images, while deep, small-area hollow tiles or hollow tiles with insignificant temperature differences are difficult to be effectively captured. Therefore, this method can usually only be used as a preliminary screening method and is difficult to provide reliable and accurate quantitative diagnostic conclusions independently. (3) Incompatible with automation scenarios and limited data: Ultrasonic testing can identify internal defects by analyzing the propagation and reflection of stress waves in the medium, and has a theoretical advantage in detecting hollow areas. However, traditional ultrasonic testing relies on manual handheld point-by-point measurements. To achieve effective acoustic coupling, the probe needs to be perpendicular to the wall and a constant pressure needs to be applied. This operation mode is heavily reliant on manual labor, is time-consuming and labor-intensive, and cannot achieve automated data acquisition for large-scale exterior walls. More importantly, the ultrasonic signal at a single point is easily affected by the coupling interference of various factors such as sensor coupling state, wall roughness, and internal steel reinforcement, resulting in a low signal-to-noise ratio, difficulty in interpretation, and insufficient robustness. Summary of the Invention

[0003] Purpose of the invention: The purpose of this invention is to provide a method for identifying hollow tiles based on multimodal data collected by a wall-climbing robot, which achieves high safety, detection accuracy, and reproducibility in complex multi-scenario tile hollow tile detection.

[0004] Technical Solution: To achieve the above objectives, the present invention provides a method for identifying hollow tiles based on multimodal data collected by a wall-climbing robot, comprising: synchronously collecting multimodal data using a wall-climbing robot, including ultrasonic modal, acoustic modal, and visual image modal data; preprocessing the multimodal data to construct a multimodal dataset for training a constructed multimodal fusion model; the multimodal fusion model includes an ultrasonic branch, an acoustic branch, a visual image branch, and a Transformer fusion module; the Transformer fusion module fuses the outputs of the ultrasonic branch, acoustic branch, and visual image branch, and then uses a Softmax classifier to distinguish between hollow and non-hollow tiles;

[0005] The multimodal dataset includes an ultrasonic modal sample structure comprising "ultrasonic echo signal - time spectrum diagram", an acoustic modal sample structure comprising "discrete-time acoustic signal sequence - short-time energy / spectral amplitude / spectral entropy / Mel cepstral coefficient features", and a visual image modal sample structure comprising "RGB image - texture feature vector / high-level feature vector".

[0006] Preferably, the wall-climbing robot performs two-dimensional scanning along the wall with a set step distance, sampling interval, and frequency, triggering ultrasonic and acoustic detection, as well as image acquisition, and performs time alignment on the acquired modal data.

[0007] Preferably, in the multimodal data preprocessing, the preprocessing of the ultrasonic modal data includes: sequentially performing DC component removal, bandpass filtering, envelope detection, and STFT transformation on the acquired ultrasonic echo signal to generate a time-frequency diagram.

[0008] Preferably, the preprocessing of the multimodal data includes: performing analog-to-digital conversion on the acquired acoustic signals to obtain a discrete-time acoustic signal sequence; performing frame-by-frame processing on the discrete-time acoustic signal sequence; and using the framed signals to calculate the short-time energy, spectral amplitude, spectral entropy, and Mel-frequency cepstral coefficients of the sound wave.

[0009] Preferably, based on the short-time energy of the sound wave Spectral amplitude Spectral entropy Mel-frequency cepstral coefficients (MFCC) Methods for identifying hollow tiles include: when and When, it is marked as a high-confidence candidate point for hollowness; when ,and The higher-order components are greater than a preset multiple of the sum of their historical mean and corresponding standard deviations, or the spectral amplitude. If the energy percentage within the preset low-to-medium frequency band exceeds a threshold, the detection point is marked as a suspicious point. The suspicious point is then triggered by a wall-climbing robot for secondary sampling, recording the acoustic echo signal and combining the ultrasonic echo data for hollowness detection. , The short-time energy of all sampling points within the current detection area The arithmetic mean and standard deviation.

[0010] Preferably, in the multimodal data preprocessing, the preprocessing of visual image modal data includes: performing gridding processing on the acquired image; performing adaptive histogram equalization enhancement on the gridded image to enhance local contrast; performing local binary pattern texture description on the enhanced image and extracting texture feature vectors for each grid unit; extracting semantic features from the image through a convolutional neural network to obtain high-level feature vectors; and determining whether there is texture continuity disruption or semantic feature anomaly in the corresponding grid region based on the texture feature vectors and high-level feature vectors. When the anomaly meets a preset condition, the grid region is marked as a visually suspicious hollow region for subsequent multimodal fusion discrimination.

[0011] Preferably, the ultrasound branch uses a ResNet-18 network with the following hierarchical structure: 7×7 convolutional layers (stride 2) + batch normalization + ReLU; 4 residual blocks, each containing two 3×3 convolutional layers; and global average pooling to output the ultrasound feature vector. ;

[0012] The acoustic branch employs a hybrid 1D-CNN+LSTM structure: 1D-CNN extracts local acoustic spectrum features; pooling layers reduce temporal redundancy; LSTM units capture the evolution trend of acoustic energy over time; and fully connected layers output acoustic feature vectors. ;

[0013] The visual image branch uses a ResNet-18 backbone network to extract surface texture and crack information, outputting a visual image feature vector. .

[0014] Preferably, the Transformer fusion module outputs a fused feature vector as follows:

[0015] ,

[0016] in, , , These are trainable weight parameters.

[0017] Preferably, the multimodal fusion model uses the following joint loss function during the training phase:

[0018] ,

[0019] in, The cross-entropy loss is used for classification to learn how to distinguish between hollow and non-hollow areas; The multimodal consistency reconstruction loss constrains the reconstruction capability of spectrogram autoencoders or voiceprint features. For the physical consistency loss based on the constraints of the sound wave propagation equation, , It's a hyperparameter.

[0020] Preferably, it also includes a lightweight ultrasonic pre-screening model deployed on the edge side. The lightweight ultrasonic pre-screening model includes a three-layer lightweight one-dimensional convolutional neural network for extracting time-domain features from ultrasonic time-domain echo signals. The extracted time-domain features are compressed in dimension by global average pooling and then input to the Sigmoid output layer to generate a confidence score for the corresponding sampling point as hollow, so as to realize real-time pre-screening of hollow tiles under low computing power edge computing conditions.

[0021] Beneficial Effects: This invention has the following advantages: 1. It integrates ultrasonic, acoustic, and visual modal data, achieving information complementarity, improving detection accuracy and robustness, and effectively overcoming the limitations of single-modal methods and environmental interference; 2. It constructs a cloud-edge collaborative intelligent architecture, achieving dual optimization of detection efficiency and system resources, and is suitable for detecting hollow tiles in complex multi-scenario environments; 3. Through automated robotic data collection and algorithmic model discrimination, the detection process is transformed from relying on subjective human experience to a standardized process relying on machines and algorithms; 4. This invention transforms the detection process into a standardized digital process driven by "machine + algorithm," realizing quantitative analysis, accurate reproduction, and full traceability of detection results, providing reliable data support for building quality assessment and long-term operation and maintenance management. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0023] Figure 2 This is a schematic diagram illustrating the actual application of the method described in this invention;

[0024] Figure 3 This is a flowchart of the closed-loop resampling decision process. Detailed Implementation

[0025] The technical solution of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.

[0026] like Figure 1 , 2 As shown, this invention provides a method for identifying hollow tiles based on multimodal data collected by a wall-climbing robot, comprising the following steps:

[0027] S1, Multimodal Data Acquisition

[0028] Wall-climbing robots with step length A two-dimensional scan is performed along the wall, triggering ultrasonic and acoustic detection, as well as image acquisition, with each sampling interval... Among them, the speed of travel Number of sampling points With sampling rate Determined by both the bandwidth of the ultrasonic transducer and the signal attenuation depth, it satisfies the following:

[0029] ,

[0030] Therefore, step size With sampling rate pass Indirect correlation, meaning that the smaller the step size and the higher the temporal resolution, the more sampling points are required. It increases accordingly.

[0031] The acquired modal data are synchronized using the IEEE 1588 Precise Time Protocol (PTP) to generate timestamps:

[0032] ,

[0033] in, for Seconds The offset is on the order of microseconds. Multimodal signal time synchronization error. .

[0034] S2, Multimodal Data Preprocessing

[0035] S201. The ultrasonic detection data is preprocessed sequentially using DC component removal, bandpass filtering, envelope detection, and STFT transformation to generate a time-frequency graph. The specific process is as follows: First, the DC component is removed from the original ultrasonic echo signal to eliminate the bias of the robot acquisition system; second, a Butterworth bandpass filter with a preset frequency band is used to filter out low-frequency noise and high-frequency interference; then, the envelope signal is obtained through Hilbert transform to enhance the echo envelope characteristics; finally, an STFT transformation is performed on the envelope signal to generate a time-frequency graph, achieving a joint representation of time-domain and frequency-domain information.

[0036] S2011, Remove DC component

[0037] For the original ultrasonic echo signal In the observation window [0, Internal removal of DC:

[0038] ,

[0039] in, For the length of the window, The number of sampling points within the observation window. The sampling rate within the observation window; This represents the time-domain representation of the ultrasonic echo signal.

[0040] S2012, Bandpass Filter

[0041] A fourth-order Butterworth filter is used, with a passband of Matching the transducer's operating bandwidth, the frequency response function is:

[0042] ,

[0043] in: The imaginary unit; , These are the lower and upper cutoff frequencies of the filter, respectively. For frequency variables.

[0044] In the frequency domain, the signal after DC removal processing To obtain the output spectrum, bandpass filtering is performed first. Perform a Fourier transform:

[0045] ,

[0046] in, This represents the frequency domain of the ultrasonic echo signal. This is the Fourier transform operator.

[0047] Then, filtering operations are performed in the frequency domain:

[0048] ,

[0049] Finally, perform an inverse Fourier transform back to the time domain:

[0050] ,

[0051] in, This represents the inverse Fourier transform operator.

[0052] S2013, Envelope Detection

[0053] Envelope detection uses the Hilbert transform to construct an analytic signal:

[0054] ,

[0055] in, For the Hilbert operator, its frequency domain expression is:

[0056] ,

[0057] in, ,

[0058] The envelope signal is obtained by taking the modulus:

[0059] ,

[0060] in, It is the real signal after bandpass filtering (the time domain signal of ultrasonic echo); yes The Hilbert transform, also known as the quadrature component of a signal, is 90° out of phase with the original signal and is used to construct the imaginary part of a complex signal. It is a constructed complex analytic signal.

[0061] S2014, Time-Frequency Transform (STFT)

[0062] Perform STFT transformation on the envelope signal:

[0063] ,

[0064] Window function Select Blackman-Harris window:

[0065] ,

[0066] Window length When the overlap is 75%, the spectrogram for a 256×256 pixel image is generated:

[0067] ,

[0068] in, This is a time-frequency spectrum diagram; This is a short-time Fourier transform; For frequency; Indicates the window function; Indicates time.

[0069] S202. Perform short-time energy calculation, spectrum analysis, spectral entropy calculation, and MFCC feature extraction on the acoustic detection data to construct an acoustic feature vector.

[0070] S2021, Acquired analog acoustic signals After analog-to-digital conversion (ADC), a discrete-time acoustic signal sequence is obtained, denoted as... ,in, The sampling point index is an integer, and the sampling frequency is... .

[0071] right Short-time frame processing is used to process discrete signals. Window function After adding a window, we get the first Frame signal :

[0072] ,

[0073] in, It is a discrete-time acoustic signal sequence; n is the frame number; The frame length is the number of sampling points contained in each frame. Frame shift refers to the number of sampling points between the start points of two adjacent frames. ; This represents the index of the sampling point within the frame. In the current frame n, m ranges from 0 to L-1. This represents a window function of length L; This represents the amplitude of the m-th sampling point of the n-th frame of the signal. This represents the temporal sample value of the nth frame.

[0074] S2022, Short-time energy is used to reflect the instantaneous intensity of a sound wave and is defined as:

[0075] ,

[0076] in, This represents the short-time energy of the nth frame.

[0077] S2023, Spectrum Analysis

[0078] Each frame signal The spectral amplitude was obtained by Fast Fourier Transform:

[0079]

[0080] The spectral power density is:

[0081]

[0082] Where n is the frame number; m is the index of the sampling point within the frame; This represents the m-th sampling point (time-domain signal) of the n-th frame; L is the length of each frame; Let be the frequency domain representation of the signal in the nth frame; j is the imaginary unit. It is a complex exponential basis function; This represents the spectral power density of the nth frame; This represents the spectral amplitude.

[0083] S2024. Spectral entropy is used to characterize the uniformity of sound wave energy distribution in the frequency domain, and is defined as:

[0084]

[0085] in, Let K be the spectral entropy of the nth frame; K is the number of spectral components. For the normalized spectral energy density, b = 1, 2, ..., K; Let be the spectral amplitude of the nth frame signal at the b-th frequency point.

[0086] S2025, Mel-frequency cepstral coefficient (MFCC) extraction

[0087] spectral power The output power of the q-th filter is obtained using a Mel filter bank containing Q triangular bandpass filters. .

[0088] Calculate the m-th order Mel-frequency cepstral coefficients :

[0089]

[0090] in, For the first Mehr cepstral coefficients; Q represents the number of cepstral coefficients extracted (typically 12-20); Q represents the number of filters in the Mel filter bank. Let q be the output power of the q-th Mel filter.

[0091] MFCC features can serve as a supplement to the auditory features for identifying hollow areas, reflecting the resonant frequency characteristics of different structural interfaces.

[0092] S2026. As shown in Table 1, the comprehensive judgment logic for the four acoustic features is as follows:

[0093] 1) When and When the time is right, mark it as a high-confidence candidate point for hollowness;

[0094] In the formula, The short-time energy of all sampling points within the current detection area (a pre-divided detection unit). The arithmetic mean; The short-time energy of all sampling points within the corresponding area The standard deviation.

[0095] 2) When Normal but The higher-order components show drastic changes and are marked as suspicious points.

[0096] 3) Suspicious points are sampled a second time by the robot, and the acoustic echo signal is recorded and judged together with the ultrasonic echo data.

[0097] Final acoustic feature vector composition:

[0098] ,

[0099] in, This is the center frequency of the spectrum.

[0100] Table 1. Correlation and Judgment Logic of Four Acoustic Features

[0101] Module Feature indicators Main physical meaning Relationship with hollow drum Short-time energy Time-domain energy intensity Increased amplitude of sound wave reflection Energy peak rises at hollow areas Spectrum Analysis Frequency domain energy distribution Hollow reflections are concentrated in the mid-to-low frequencies. Bandwidth power concentration Spectral entropy Energy Concentration Reduced uniformity of spectral distribution For hollow drum MFCC Auditory characteristics Resonance frequency change High-frequency cepstral fluctuations

[0102] .

[0103] S203. For the acquired raw images Step-by-step preprocessing and multi-layer feature extraction are performed, including ROI cropping, gridded segmentation, CLAHE adaptive histogram enhancement, LBP local binary pattern texture description, and high-level semantic feature extraction by convolutional neural network (CNN) to improve the identifiability of hollow areas in tile surface images and the accuracy of algorithm discrimination.

[0104] S2031, ROI clipping and meshing

[0105] Based on the detection area of ​​the ultrasonic transducer or the field of view of the camera, extract the region of interest (ROI). Let the original image be... Then the ROI region can be defined as:

[0106]

[0107] in, The coordinates of the top left corner of the ROI. , The height and width of the ROI.

[0108] To improve the local responsiveness of feature extraction, the ROI region is divided into M×N grid cells, each grid cell... It can be represented independently as:

[0109] ,

[0110] in, , i and j are the row and column indices (row number and column number) of the grid, respectively.

[0111] This gridded segmentation helps to analyze subtle texture changes in the tiles at a local scale, thereby capturing structural anomalies in hollow areas.

[0112] S2032, Adaptive Histogram Equalization (CLAHE) Enhancement

[0113] To improve texture visibility under different illumination conditions, a contrast-limited adaptive histogram equalization algorithm is employed. Its basic principle is to independently calculate and equalize the gray-level histogram on each local grid cell, while simultaneously introducing a contrast-limiting threshold. This prevents excessive enhancement in localized areas.

[0114] Let the local gray-level probability density be... The mapping function of CLAHE can be expressed as:

[0115] ,

[0116] ,

[0117] in, The number of gray levels; Let be the grayscale probability density function of the input image; It is grayscale The number of pixels; It is the total number of pixels; A contrast limit threshold is used to prevent excessive enhancement in local areas; This represents the restricted probability density.

[0118] Enhanced image It can achieve a more uniform brightness distribution, making texture features such as hollow lines and cracks more prominent.

[0119] In actual detection, image signals are often affected by fluctuations in external lighting and sensor noise. Therefore, noise suppression algorithms are introduced before and after CLAHE enhancement, using Gaussian smoothing or bilateral filtering.

[0120] ,

[0121] in, and These are Gaussian kernel functions for the spatial domain and the grayscale domain, respectively. This is the normalization factor. This process effectively reduces the interference of random noise while maintaining the clarity of tile cracks and hollow edges.

[0122] S2033, Local Binary Mode (LBP) Texture Description

[0123] Texture pattern encoding is performed on the enhanced image, using rotation-invariant unified local binary mode (ROM). LBP operator. The LBP calculation formula is:

[0124] ,

[0125] in, The grayscale value of the center pixel; Let P be the gray value of the p-th neighboring pixel on radius R1; P is the number of sampling points; R1 is the sampling radius. This is a symbolic function used for binarization.

[0126] The operator compares the brightness difference between neighboring pixels and the center pixel to form a binary pattern code, thereby describing the local microstructure of the image (such as edges, corners, texture direction, etc.).

[0127] After LBP encoding, each grid cell obtains a set of texture feature vectors. This can be statistically analyzed using histograms, and expressed in the following form:

[0128] ,

[0129] in, For indicator functions, For the total number of patterns, This represents the total number of pixels.

[0130] S2034. Semantic Feature Extraction from Convolutional Neural Networks (CNN)

[0131] After completing the local texture encoding, the LBP feature mapping results are input into a convolutional neural network (CNN) to extract high-level semantic features. The CNN includes several convolutional layers, batch normalization layers, non-linear activation layers, and pooling layers. Through multiple layers of non-linear mapping, the CNN model can automatically learn the deep differences in texture, structure, and lighting features between hollow areas and normal areas of tiles.

[0132] Let the input be a local feature map. Then the high-level feature vector extracted by the CNN is:

[0133] ,

[0134] in, Represented as feature map height × feature map width × number of channels; It is a feature vector of length 512. Vectors have good discriminative power in feature space and can effectively distinguish between hollow and intact regions.

[0135] S3. Based on the multimodal data preprocessed in S2, a deep multimodal fusion model is constructed, including three feature extraction branches (ultrasound branch, acoustic branch, and visual image branch) and a Transformer fusion module. The three feature extraction branches correspond to the ultrasound modality, acoustic modality, and visual image modality, respectively.

[0136] S301, Ultrasonic Branch, training data is ultrasonic echo signal and the corresponding generated time spectrum The training set consists of labels, with the label structure being {label: hollow / non-hollow, coupling state:}. Source: Simulation / Real Measurement} This represents the "contact status of the robot probe at the tile interface".

[0137] The ultrasound branch uses a ResNet-18 network with the following hierarchical structure: 7×7 convolutional layers (stride 2) + batch normalization + ReLU; 4 residual blocks, each containing two 3×3 convolutional layers; and a 512-dimensional ultrasound feature vector output after global average pooling. , =[ This branch automatically learns the local energy distribution, harmonic modes, and time-frequency resonance characteristics of the spectrum through convolutional kernels, thereby extracting the ultrasonic response characteristics of the hollow area.

[0138] In this embodiment, to increase the number of ultrasound samples or modify the original ultrasound data, an ultrasound reflection physical model based on multilayer dielectric transmission line theory and frequency-dependent attenuation correction is constructed. This model serves as a mathematical model to simulate the "reflection signal spectrum under different physical conditions." The specific implementation process is as follows:

[0139] A multi-medium ultrasonic reflection model was established, treating the ceramic tile-adhesive layer-wall as three non-homogeneous media with acoustic impedances of... The ultrasonic reflection coefficient is modeled using a transmission line model:

[0140]

[0141] in, , where is the frequency-varying acoustic impedance of the i-th layer; The phase constant of the adhesive layer; This refers to the thickness of the adhesive layer.

[0142] In an ideal, lossless medium, the spectrum of the reflected echo can be expressed as:

[0143] ,

[0144] in, The spectrum of the transmitted signal; The reflection coefficient is determined by the difference in acoustic impedance at the interface.

[0145] However, in reality, ultrasound waves experience energy attenuation as they propagate from the transducer to the reflecting interface and back. The attenuation originates from factors including material absorption. Surface roughness scattering Additional losses caused by uneven coupling agent thickness and poor contact.

[0146] Therefore, the true reflected signal should be written as:

[0147] ,

[0148] in, It is a multiplicative attenuation factor that performs frequency-weighted correction on the reflection coefficient.

[0149] Considering the scattering attenuation caused by uneven coupling agent thickness and surface roughness, a frequency-dependent attenuation model is introduced:

[0150] ,

[0151] in, The standard deviation of surface roughness; Wave number; The propagation path length (the round trip distance from the probe to the interface and back); The material absorption coefficient, This is the empirical absorption coefficient constant; The initial amplitude calibration coefficient represents the emission characteristics of the probe and the coupling layer; the first term The attenuation caused by material absorption decreases exponentially with increasing propagation distance; the second term Due to roughness scattering attenuation, high-frequency components attenuate even faster.

[0152] In the attenuation model, "material" specifically refers to the combined medium of the tile layer, adhesive layer, and coupling agent layer; among them, the thickness and roughness of the adhesive layer are the most significant, so the attenuation is mainly reflected in the adhesive layer.

[0153] The import method is as follows:

[0154] In the frequency domain, the reflection coefficient is directly corrected:

[0155] ,

[0156] in, The actual reflection intensity after adding frequency weighting is closer to the spectral amplitude distribution of the true sample.

[0157] Then substitute it into the calculation of the reflected signal:

[0158] ,

[0159] Then, the inverse Fourier transform is taken to obtain the time-domain echo signal:

[0160] ,

[0161] This step is used to simulate echo waveforms under different coupling conditions, thereby adding more samples to the training set and improving the robustness of the model.

[0162] In ultrasonic testing, "coupling conditions" refer to the state of acoustic energy transfer between the probe and the surface being tested. In actual testing, it is affected by several factors, as shown in Table 2:

[0163] Table 2. Influencing Factors under Different Coupling Conditions

[0164] Parameter symbol Physical meaning Corresponding impact Surface roughness ( ) The larger the value, the stronger the scattering attenuation. Absorption coefficient constant The higher the value, the more energy the material absorbs. Effective transmission distance ( ) The greater the coupling agent thickness or the longer the sound path, the greater the signal attenuation. Adhesive layer thickness Changing the phase of reflection / transmission interference Emission Amplitude Calibration Different contact pressures between the probe and the coupling medium

[0165] These collectively define the physical characteristics of the "coupled state." Therefore, different coupling conditions can be combined in different ways:

[0166]

[0167] Each This represents a "robot probe-tile interface contact state".

[0168] S302, Acoustic Branch, training data consists of discrete-time acoustic signal sequences. and the corresponding four acoustic eigenvectors (short-time energy) Spectrum analysis Spectral entropy MFCC The training set consists of [various components]. The training samples of the acoustic branch also include hollow drum annotation information corresponding to the sampling positions, which is obtained by manual tapping detection, manual verification, or other non-destructive detection methods.

[0169] The acoustic branch employs a hybrid 1D-CNN+LSTM structure: 1D-CNN extracts local acoustic spectrum features; pooling layers reduce temporal redundancy; LSTM units capture the evolution of acoustic energy over time; and finally, fully connected layers output a 256-dimensional acoustic feature vector. , This branch captures the energy decay patterns and spectral structure differences of percussion or echo sounds over time.

[0170] S303, Visual Image Branch, training data consists of images and the corresponding texture feature vector High-level feature vectors The training set consists of [a set of data]. The training samples of the visual image branch also include hollow area annotation information corresponding to the image or image grid region. This annotation information is used to supervise the training of the convolutional neural network to extract high-level semantic features related to tile hollowness.

[0171] The visual image branch uses a ResNet-18 backbone network to extract surface texture and crack information; it outputs a 512-dimensional visual image feature vector. , This branch uses CNNs to abstractly represent surface geometry (tile boundaries, cracks, color differences, etc.).

[0172] S304, Transformer Fusion Module

[0173] Each modal feature is encoded with its corresponding modality code before being input into the Transformer fusion module. ,form:

[0174] ,

[0175] ,

[0176] in, This is a modal embedding encoding, where k1 is the feature dimension index and d is the total dimension of the feature vector. For modal position encoding function, Features of the Transformer fusion module after adding modality coding.

[0177] Features encoded by position / modality are concatenated along their vector dimensions:

[0178] .

[0179] concatenated feature vectors Input to the Transformer fusion module. Feature vector The labeled data structure is defined as follows:

[0180]

[0181] .

[0182] The core of the Transformer fusion module is a multi-head self-attention mechanism used to model inter-modal dependencies:

[0183] , , , ,

[0184] Through a self-attention mechanism, the model can automatically learn the correlation and complementarity between ultrasound, acoustic and visual features, thereby strengthening consistent features and suppressing noise features in the fused output.

[0185] The Transformer fusion module outputs a fused feature vector. The final distinction between hollow and non-hollow areas is made through a fully connected layer and a Softmax classifier.

[0186] ,

[0187] in, , , These are trainable weight parameters.

[0188] S304, The deep multimodal fusion model uses a joint loss function during the training phase:

[0189] ,

[0190] in, The cross-entropy loss is used for classification to learn how to distinguish between hollow and non-hollow areas; To constrain the reconstruction capability of time-frequency map autoencoders or voiceprint features for multimodal consistency reconstruction loss; To address the physical consistency loss based on the constraints of the sound wave propagation equation, the model output is forced to conform to the physical laws governing ultrasonic wave propagation. , These are hyperparameters used to adjust the relative impact of various losses.

[0191] Through joint optimization, the model can achieve both high classification accuracy and conformity to acoustic physics laws, ensuring that the results are interpretable and stable.

[0192] S4. Achieve collaboration with the cloud-based multimodal fusion model by deploying a lightweight ultrasound pre-screening model at the edge.

[0193] In practical applications, the computing power, power consumption, and storage of the edge computing unit carried by the wall-climbing robot are limited, making it unsuitable for directly deploying a deep multimodal fusion model. Therefore, a lightweight ultrasonic signal pre-screening model is deployed on the edge side to quickly screen the original ultrasonic time-domain waveform, reducing uplink bandwidth and cloud load.

[0194] The design process for a lightweight ultrasound pre-screening model is as follows:

[0195] S401, Input and Network Structure

[0196] Marginal lightweight ultrasound pre-screening model uses original ultrasound time-domain waveform sequences is the input, where T is the number of sampling points.

[0197] The lightweight ultrasound pre-screening model extracts ultrasound time-domain features through a three-layer lightweight 1D-CNN, compresses them using GAP, and then inputs them into a Sigmoid scoring layer to achieve real-time anomaly detection with low computing power.

[0198] First layer: 1D convolution + BN + ReLU:

[0199] ,

[0200] in, This is the first layer of convolution kernels. Represents convolution operation. Indicates bias. Indicates batch normalization. This represents the activation function.

[0201] Output characteristics:

[0202] .

[0203] Second layer 1D convolution + BN + ReLU:

[0204] ,

[0205] Output characteristics:

[0206] .

[0207] Third layer: 1D convolution + BN + ReLU:

[0208] ,

[0209] Output characteristics:

[0210] .

[0211] GAP (Global Average Pooling) performs time-series averaging on each channel:

[0212] ,

[0213] The final feature vector is obtained as follows:

[0214] .

[0215] GAP compresses dynamic ultrasound temporal features into a fixed-length vector, greatly reducing the complexity of edge computing.

[0216] The anomaly detection score is derived from the linear layer plus the sigmoid output layer.

[0217] ,

[0218] in, 'b' represents the classification weight, and 'b' represents the bias. This refers to the Sigmoid function.

[0219] Final score: 0 This indicates the probability or anomaly index of an anomaly occurring at that location.

[0220] S402, Definition of Abnormal Scores

[0221] The "abnormal score s" represents the probability or confidence that the time-domain waveform belongs to the "hollow echo mode". A higher value indicates a signal closer to the hollow type; a lower value indicates a signal closer to normal coupling. The distance between the output score s of the lightweight ultrasound pre-screening model and the classification boundary represents the uncertainty. The closer the score is to the midpoint between the two threshold intervals, the higher the uncertainty, requiring resampling to improve judgment quality. Figure 3 As shown.

[0222] S402, Relationship with Multimodal Fusion Model

[0223] The relationship between the two is shown in Table 3:

[0224] Table 3. Correlation between pre-screened models and multimodal models

[0225] Model Location Function Computing power requirements Lightweight ultrasound pre-screening model Edge side (robot end) Real-time rapid pre-screening of all acquired signals to determine whether uploading or resampling is necessary. Extremely low Multimodal fusion model Cloud (centralized server) Accurately identify and integrate key data for decision-making after screening. high

[0226] Intelligent detection and triggering mechanism based on anomaly score s:

[0227] Lightweight ultrasound pre-screening model outputs abnormality scores The system makes a determination based on a two-threshold strategy:

[0228] like If the sound is detected as hollow, the data is uploaded and recorded.

[0229] like : If the condition is deemed normal, no data will be uploaded;

[0230] like If a region is identified as suspicious, an intelligent resampling strategy will be triggered.

[0231] threshold and This is a system experience setting used to improve screening accuracy and recall.

[0232] Intelligent resampling strategy for suspicious areas: When a lightweight ultrasound pre-screening model identifies a certain area as "suspicious," the wall-climbing robot performs the following actions to improve data quality:

[0233] 1. Reduce stride distance: Improve spatial resolution by sampling more densely at measurement points.

[0234] 2. Change the probe contact force and angle: to compensate for ultrasonic coupling loss caused by poor contact.

[0235] 3. Enable more channels in the array in parallel: improve incident angle coverage and redundancy.

[0236] 4. Extend the sampling window to improve SNR: Improve the signal-to-noise ratio by sampling for a longer time window or by averaging cumulatively.

[0237] These steps enable the edge to acquire higher-quality ultrasound signals, which are then uploaded to the cloud-based multimodal model for accurate discrimination.

[0238] Deployment and acceleration of lightweight ultrasound pre-screening model: To adapt to low-power hardware, the lightweight ultrasound pre-screening model of this invention is designed as follows:

[0239] 1. ONNX Export: Export PyTorch / TensorFlow models to ONNX format for easy cross-platform deployment.

[0240] 2. TensorRT accelerates inference: TensorRT accelerates lightweight networks through graph optimization, operator fusion, and other techniques.

[0241] 3. FP16 / INT8 Quantization: FP16: Half-precision, reducing latency and power consumption; INT8: Integer quantization, which can significantly improve throughput while maintaining accuracy.

[0242] The final lightweight ultrasound pre-screening model achieves real-time (millisecond-level) ultrasound signal screening capability on edge devices.

Claims

1. A method for identifying hollow tiles based on multimodal data collected by a wall-climbing robot, characterized in that, include: Multimodal data synchronous acquisition was carried out using a wall-climbing robot, including ultrasonic modal, acoustic modal, and visual image modal data; Multimodal data is preprocessed to construct a multimodal dataset for training the constructed multimodal fusion model. The multimodal fusion model includes an ultrasound branch, an acoustic branch, a visual image branch, and a Transformer fusion module. The Transformer fusion module fuses the outputs of the ultrasound branch, the acoustic branch, and the visual image branch, and then uses a Softmax classifier to distinguish between hollow and non-hollow areas. The multimodal dataset includes an ultrasonic modal sample structure comprising "ultrasonic echo signal - time spectrum diagram", an acoustic modal sample structure comprising "discrete-time acoustic signal sequence - short-time energy / spectral amplitude / spectral entropy / Mel cepstral coefficient features", and a visual image modal sample structure comprising "RGB image - texture feature vector / high-level feature vector".

2. The method for identifying hollow tiles according to claim 1, characterized in that, The wall-climbing robot performs two-dimensional scanning along the wall with a set step size, sampling interval, and frequency, triggering ultrasonic and acoustic detection, as well as image acquisition, and aligning the acquired modal data in time.

3. The method for identifying hollow tiles according to claim 1, characterized in that, In the multimodal data preprocessing, the preprocessing of ultrasonic modal data includes: sequentially performing DC component removal, bandpass filtering, envelope detection, and STFT transformation on the acquired ultrasonic echo signal to generate a time-frequency diagram.

4. The method for identifying hollow tiles according to claim 1, characterized in that, In the multimodal data preprocessing, the preprocessing of acoustic ultrasonic modal data includes: performing analog-to-digital conversion on the acquired acoustic signals to obtain a discrete-time acoustic signal sequence; performing frame-by-frame processing on the discrete-time acoustic signal sequence; and using the framed signals to calculate the short-time energy, spectral amplitude, spectral entropy, and Mel-frequency cepstral coefficients of the sound wave.

5. The method for identifying hollow tiles according to claim 4, characterized in that, Based on the short-time energy of sound waves Spectral amplitude Spectral entropy Mel-frequency cepstral coefficients (MFCC) Methods for identifying hollow tiles include: when and When, it is marked as a high-confidence candidate point for hollowness; when ,and The higher-order components are greater than a preset multiple of the sum of their historical mean and corresponding standard deviations, or the spectral amplitude. If the energy percentage within the preset low-to-medium frequency band exceeds a threshold, the detection point is marked as a suspicious point. The suspicious point is then triggered by a wall-climbing robot for secondary sampling, recording the acoustic echo signal and combining the ultrasonic echo data for hollowness detection. , The short-time energy of all sampling points within the current detection area The arithmetic mean and standard deviation.

6. The method for identifying hollow tiles according to claim 1, characterized in that, In the multimodal data preprocessing, the preprocessing of visual image modal data includes: performing gridding on the acquired image; performing adaptive histogram equalization enhancement on the gridded image to enhance local contrast; performing local binary pattern texture description on the enhanced image and extracting texture feature vectors for each grid cell; extracting semantic features from the image using a convolutional neural network to obtain high-level feature vectors; and determining whether there is texture continuity disruption or semantic feature anomaly in the corresponding grid region based on the texture feature vectors and high-level feature vectors. When the anomaly meets a preset condition, the grid region is marked as a visually suspicious hollow region for subsequent multimodal fusion discrimination.

7. The method for identifying hollow tiles according to claim 1, characterized in that, The ultrasound branch uses a ResNet-18 network with the following hierarchical structure: 7×7 convolutional layers (stride 2) + batch normalization + ReLU; 4 residual blocks, each containing two 3×3 convolutional layers; and global average pooling to output the ultrasound feature vector. ; The acoustic branch employs a hybrid 1D-CNN+LSTM structure: 1D-CNN extracts local acoustic spectrum features; pooling layers reduce temporal redundancy; LSTM units capture the evolution trend of acoustic energy over time; and fully connected layers output acoustic feature vectors. ; The visual image branch uses a ResNet-18 backbone network to extract surface texture and crack information, outputting a visual image feature vector. .

8. The method for identifying hollow tiles according to claim 7, characterized in that, The Transformer fusion module outputs a fused feature vector as follows: , in, , , These are trainable weight parameters.

9. The method for identifying hollow tiles according to claim 1, characterized in that, The multimodal fusion model uses the following joint loss function during the training phase: , in, The cross-entropy loss is used for classification to learn how to distinguish between hollow and non-hollow areas; The multimodal consistency reconstruction loss constrains the reconstruction capability of spectrogram autoencoders or voiceprint features. For the physical consistency loss based on the constraints of the sound wave propagation equation, , It's a hyperparameter.

10. The method for identifying hollow tiles according to claim 1, characterized in that, It also includes a lightweight ultrasonic pre-screening model deployed on the edge side. The lightweight ultrasonic pre-screening model includes a three-layer lightweight one-dimensional convolutional neural network, which is used to extract time-domain features from ultrasonic time-domain echo signals. The extracted time-domain features are compressed in dimension by global average pooling and then input to the Sigmoid output layer to generate a confidence score for the corresponding sampling point as hollow, so as to realize real-time pre-screening of hollow tiles under low computing power edge computing conditions.