Methods, electronic equipment and storage media for detecting structural loosening
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本申请的主要目的在于提供一种建筑结构松动检测方法、电子设备及存储介质,旨在解决建筑外围护系统松动检测的现有技术存在准确性差、可靠性低的技术问题
本申请提供的建筑外围护系统松动音频敲击检测方法,在获取实验音频数据后,对其进行了多种形式的数据增强,并采用实验音频数据、物理增强音频数据和衍生增强音频数据以及对应的真实标签训练随机森林分类模型,有效扩展了训练样本的声学特征分布范围,弥补了真实采集样本数量不足、覆盖场景有限的缺陷;基于预设敲击参数对待测建筑结构进行敲击,保证了敲击激励的标准化与一致性,避免因敲击力度、角度、位置等人工操作差异导致采集的音频数据出现不必要的波动,提升了后续特征提取与模型判定的可靠性;通过提取音频数据的多维度音频特征,突破了传统单一时域或频域特征分析的局限性,能够全面表征结构敲击响应的声学特性,提高了检测精度;采用经过训练的随机森林分类模型对音频特征进行智能判定,兼具训练效率高、抗干扰能力强、输出结果稳定的优势,能够适配工程现场的复杂环境,实现快速、准确的结构松动检测。本申请解决了现有技术中建筑外围护系统松动检测方法准确性差、可靠性低的问题。
Smart Images

Figure CN122575416A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of building engineering testing technology, and in particular to a method for detecting structural loosening, electronic equipment, and storage medium. Background Technology
[0002] As an important component of building structure, the building envelope system requires regular loosening inspection of its connectors. This is a crucial technical basis for daily building maintenance, repair projects, renovation design, and insurance assessment.
[0003] Currently, the detection of loose connectors in building envelope systems mainly involves inspectors attaching sensors to the points to be inspected, then manually tapping the panels to generate vibration excitation. Data acquisition equipment collects the vibration response signals from the panels, and the looseness of the connectors is determined by analyzing changes in the natural frequency or other spectral characteristics of the vibration signals. However, the vibration excitation generated by manual tapping is not constant, making it difficult to establish a stable mapping between the loose state and signal characteristics. Analyzing only the natural frequency or other spectral characteristics of the vibration signal results in a single analytical dimension and low detection accuracy. Existing machine learning-based acoustic detection solutions typically use real tapping audio collected on-site as training samples to directly train a classification model. The trained model then identifies the state of the structure under test based on the tapping audio. However, in practical engineering applications, the number of effective samples available for model training is generally limited due to objective limitations such as complex on-site construction environments, large individual structural differences, and the difficulty in collecting loose anomaly samples. Furthermore, the range of acoustic features covered by the samples is relatively narrow.
[0004] This demonstrates that existing technologies for detecting loosening of building envelope systems suffer from poor accuracy and low reliability. Summary of the Invention
[0005] The main purpose of this application is to provide a method, electronic device and storage medium for detecting looseness in building structures, aiming to solve the technical problems of poor accuracy and low reliability in the existing technology for detecting looseness in building envelope systems.
[0006] To achieve the above objectives, this application proposes a method for detecting structural loosening, the method comprising: Acquire experimental audio data, and perform data enhancement on the experimental audio data based on a preset on-site background noise and / or a preset reverberation algorithm to obtain physically enhanced audio data; The experimental audio data is augmented based on a preset variational autoencoder, a pre-trained generative adversarial network, and / or a pre-trained diffusion model to obtain derived augmented audio data. The generative adversarial network is independently trained based on the experimental audio data according to the category of the tightness state of the building structure samples. The diffusion model is trained based on the experimental audio data through an iterative process of forward noise addition and backward noise reduction. The target random forest classification model is trained based on the experimental audio data, the physically enhanced audio data, the derived enhanced audio data, and the real labels corresponding to the experimental audio data, the physically enhanced audio data, and the derived enhanced audio data, respectively. The building structure under test is struck based on preset striking parameters; Audio data generated by the building structure under test is collected, and multi-dimensional audio features of the audio data are extracted, wherein the multi-dimensional audio features include at least time domain features and frequency domain features; The multi-dimensional audio features are input into the target random forest classification model, and the detection results of the building structure to be tested are predicted and output through the target random forest classification model.
[0007] For example, the step of predicting and outputting the detection result of the building structure to be tested using the target random forest classification model includes: The multi-dimensional audio features are input into the target random forest classification model, which is based on an ensemble of multiple decision trees. The multi-dimensional audio features are classified using each decision tree, and a classification result is output, which is either loose or not loose. Based on the classification results corresponding to all decision trees, the loosening judgment result is determined by majority voting, and the confidence level corresponding to the loosening judgment result is calculated at the same time. The loosening determination result and the corresponding confidence level are used as the detection result of the building structure under test, and the detection result is output.
[0008] For example, the step of outputting the detection result includes: Output the loosening determination result and confidence level of the tested building structure to the display screen; The confidence level is compared with a preset confidence threshold. When the confidence level is lower than the preset confidence threshold, a re-examination prompt message is output to the display screen.
[0009] For example, the step of acquiring experimental audio data includes: Various building structure samples with different fastening conditions were tapped, and experimental audio data corresponding to each building structure sample were collected.
[0010] For example, the step of performing data enhancement on the experimental audio data based on a preset on-site background noise and / or a preset reverberation algorithm to obtain physically enhanced audio data includes: Multiple preset ambient noises are superimposed onto the experimental audio data to obtain physically enhanced audio data; And / or, based on a preset reverb algorithm, apply reverb effects for various scenarios to the experimental audio data to obtain physically enhanced audio data.
[0011] For example, the derived enhanced audio data includes first derived enhanced audio data and / or second derived enhanced audio data, and the step of data augmentation of the experimental audio data based on a preset variational autoencoder, a pre-trained generative adversarial network, and / or a pre-trained diffusion model to obtain derived enhanced audio data includes: The experimental audio data is mapped to a low-dimensional latent space using a preset variational autoencoder, and the latent space mean vector corresponding to the experimental audio data is obtained. A Gaussian perturbation is introduced into the latent space mean vector to obtain the perturbation latent vector; The perturbation latent vector is input into the decoder of the variational autoencoder for reconstruction, generating the first derived enhanced audio data; Second derived enhanced audio data is generated based on multiple trained generative adversarial networks, wherein the generative adversarial networks include the WaveGAN model; Alternatively, a second derived enhanced audio data can be generated based on the trained diffusion model.
[0012] For example, the step of tapping the building structure under test based on preset tapping parameters includes: The preset striking parameters are obtained and transmitted to a preset vibrator. The preset striking parameters include at least striking force and striking frequency. The vibrator is controlled to strike the building structure under test based on the preset striking parameters.
[0013] For example, the step of extracting multi-dimensional audio features from the audio data includes: Read and preprocess the audio data, and output standardized audio data. The preprocessing method includes at least one of the following: format standardization, endpoint detection and effective tap signal extraction, DC component removal and baseline correction, noise suppression, and outlier sample detection and removal. Extract multi-dimensional audio features from the standardized audio data. The multi-dimensional audio features include time-domain features, frequency-domain features, and 13th-order Mel frequency cepstral coefficients. The time-domain features include at least one of the following: root mean square energy, zero-crossing rate, peak amplitude, and total energy. The frequency-domain features include at least one of the following: fundamental frequency, main frequency band energy distribution, and spectral centroid.
[0014] In addition, to achieve the above objectives, this application also proposes an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the building structure loosening detection method described above.
[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the building structure loosening detection method described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the building structure loosening detection method described above.
[0017] One or more technical solutions proposed in this application have at least the following technical effects: The method for detecting loosening of building envelope systems using audio tapping provided in this application, after acquiring experimental audio data, performs various forms of data augmentation. It then trains a random forest classification model using experimental audio data, physically augmented audio data, derived augmented audio data, and corresponding real labels. This effectively expands the acoustic feature distribution range of the training samples, compensating for the shortcomings of insufficient real-world sample quantity and limited scene coverage. The method taps the building structure under test based on preset tapping parameters, ensuring the standardization and consistency of the tapping excitation and avoiding unnecessary fluctuations in the acquired audio data due to differences in tapping force, angle, and position, thus improving the reliability of subsequent feature extraction and model judgment. By extracting multi-dimensional audio features from the audio data, it overcomes the limitations of traditional single time-domain or frequency-domain feature analysis, comprehensively characterizing the acoustic properties of the structural tapping response and improving detection accuracy. The method uses a trained random forest classification model for intelligent judgment of audio features, combining high training efficiency, strong anti-interference ability, and stable output results. It can adapt to the complex environment of engineering sites, achieving rapid and accurate structural loosening detection. This application solves the problems of poor accuracy and low reliability in existing methods for detecting loosening of building envelope systems. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating an embodiment of the building structure loosening detection method of this application. Figure 2 A flowchart illustrating steps S41 to S44 of the structural loosening detection method of this application; Figure 3 This is a flowchart illustrating Embodiment 2 of the building structure loosening detection method of this application; Figure 4 A flowchart illustrating steps A21 to A27 of the structural loosening detection method in this application; Figure 5 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the building structure loosening detection method in this application embodiment.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] Example 1 As a crucial component of building structure, the building envelope system requires regular loosening checks on its connectors. This serves as a vital technical basis for routine building maintenance, repair projects, renovation design, and insurance assessments. However, existing loosening detection methods suffer from poor consistency in audio excitation sources and limited analytical dimensions, resulting in low accuracy and reliability.
[0025] The main solution of this application embodiment is to replace manual striking with a vibrator whose striking parameters are controllable, thereby generating a stable audio excitation source from the source. By extracting multi-dimensional audio features from the audio data, the shortcomings of single-dimensional audio detection and analysis are avoided, thus improving the reliability of the detection method. A random forest classification model is trained using laboratory audio data under different experimental conditions, and the trained target random forest classification model is used to intelligently determine audio features, improving the anti-interference ability of the detection method and the accuracy of the detection results.
[0026] It should be noted that the executing entity in this embodiment can be a computer device with data processing, network communication, and program execution functions, such as a desktop computer, tablet computer, smartphone, smartwatch, etc., or an electronic device capable of performing the above functions. This embodiment does not impose specific limitations on this. The following uses a desktop computer as an example to describe this embodiment and the following embodiments.
[0027] Based on this, the embodiments of this application provide a method for detecting structural loosening, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the method of this application.
[0028] In this embodiment, the method for detecting loose building structures may include steps S10 to S40: Step S10: Obtain enhanced audio data, and train the preset initial random forest classification model based on the enhanced audio data to obtain the target random forest classification model. The enhanced audio data includes at least physical enhanced audio data and derived enhanced audio data. The enhanced audio data includes physically enhanced audio data and derived enhanced audio data. The physically enhanced audio data is obtained by enhancing the experimental audio data based on a preset on-site background noise and / or a preset reverberation algorithm. The derived enhanced audio data is obtained by enhancing the experimental audio data based on a preset variational autoencoder, a pre-trained generative adversarial network, and / or a pre-trained diffusion model. Furthermore, the generative adversarial network is independently trained based on the experimental audio data according to the category of the tightness state of the building structure samples. The diffusion model is trained based on the experimental audio data through an iterative process of forward noise addition and backward noise reduction.
[0029] The model training process includes: training a preset initial random forest classification model based on experimental audio data, physically enhanced audio data, derived enhanced audio data, and the real labels corresponding to the experimental audio data, physically enhanced audio data, and derived enhanced audio data, thereby obtaining the target random forest classification model.
[0030] It should be noted that augmented audio data refers to the training dataset obtained after expanding the original tapping audio samples. Its core function is to enrich the feature coverage of the samples and solve the problem of weak model generalization ability caused by insufficient real samples. Among them, physical augmented audio data refers to augmented audio data generated by simulating the acoustic propagation process in the real physical world, which can realistically reproduce the acoustic environment during on-site detection. Derived augmented audio data is a high-fidelity augmented sample synthesized by a deep learning generation model, which can simulate the acoustic feature fluctuations caused by differences in structural physical characteristics. The acoustic patterns of the samples are highly consistent with the real collected data. The initial random forest classification model is an ensemble learning model composed of multiple decision trees. It outputs the classification result through voting by multiple decision trees and has the characteristics of anti-overfitting, high training efficiency, and strong robustness. The target random forest classification model is an ensemble learning classification model for identifying the loose state of building structures, which is obtained after training and convergence with augmented audio data. It is the final usable form of the initial random forest classification model after parameter optimization.
[0031] Specifically, firstly, high signal-to-noise ratio (SNR) audio of building structures being struck is collected as the original seed sample, labeled as "normal" and "loose" according to the actual tightness of the structure. The original seed sample undergoes physical enhancement and derivative enhancement processing to generate physically enhanced audio data and derivative enhanced audio data that correspond one-to-one with the original sample labels. The physically enhanced audio data and derivative enhanced audio data are then mixed with the original seed sample to form a complete enhanced audio dataset. The enhanced audio dataset is divided into training and validation sets according to a preset ratio and input into the initial random forest classification model. Hyperparameters such as the number of decision trees, maximum depth, and feature sampling ratio are set. Iterative training is performed, combining the experimental audio data, physically enhanced audio data, and the real labels corresponding to the derivative enhanced audio data (set according to the actual looseness situation), with classification accuracy as the optimization objective. The convergence state of the model is verified through the validation set. After the model performance stabilizes, the final trained target random forest classification model is obtained.
[0032] Step S20: Tap the building structure to be tested based on preset tapping parameters; It should be noted that the preset impact parameters refer to a set of quantified parameters pre-stored in memory for standardized control of the vibrator's operating state, including but not limited to impact force, impact angle, impact speed, impact duration, and impact interval. Impact refers to the action of applying instantaneous impact force to the surface of a building structure, performed by the vibrator built into the audio impact detection device. The vibrator is electromagnetically or piezoelectrically driven, unlike manual impact, ensuring consistency in excitation each time. This step, through standardized vibration control, fundamentally solves the problem of poor excitation source consistency in existing technologies.
[0033] Specifically, the hammer head of the vibrator is automatically adjusted to be directly above the test point, and laser positioning is used to ensure that the striking angle is perpendicular to the surface of the building structure; the vibrator is started and a single striking action is performed according to the preset striking parameters; after the striking is completed, the hammer head of the vibrator automatically resets.
[0034] Step S30: Collect audio data generated by the building structure under test, and extract multi-dimensional audio features from the audio data. The multi-dimensional audio features include at least time-domain features and frequency-domain features. It should be noted that audio data refers to the sound wave signals generated by the vibration of a building structure after being struck, which are radiated into the air and collected by a microphone as digitized electrical signals. Time-domain characteristics refer to the features that describe the changing patterns of the audio signal over time, reflecting the amplitude and energy distribution of the signal over time. For example, time-domain characteristics include, but are not limited to, root-mean-square energy, zero-crossing rate, peak amplitude, and total energy. Frequency-domain characteristics refer to the features that describe the composition of the audio signal from a frequency perspective, reflecting information such as the frequency distribution, natural frequencies, and harmonic components. For example, frequency-domain characteristics include, but are not limited to, spectral centroid, spectral bandwidth, natural frequencies, and the amplitude ratio of the first harmonic to the second harmonic. Multi-dimensional audio features can comprehensively characterize the audio properties generated by the vibration of a building structure under impact, providing sufficient evidence for determining loosening. For example, multi-dimensional audio features may also include psychoacoustic features such as wavelet packet node energy spectra or cepstral coefficients of a Gamma Tone filter group.
[0035] Specifically, the built-in directional microphone is activated to collect audio signals within a preset time range after the tap at a preset sampling rate and sampling precision, while simultaneously collecting ambient background noise. An adaptive filtering algorithm is used to remove interference signals such as ambient noise and equipment operating noise from the collected audio data, and core time-domain features are extracted. A fast Fourier transform is performed on the audio data to obtain the signal spectrum, and core frequency-domain features are extracted from the spectrum. All extracted time-domain and frequency-domain features are combined in a preset order to generate a standardized feature vector, which is temporarily stored in the memory of the detection device. For example, the detection device collects audio data after a granite slab is struck, extracting the following time-domain features: root mean square value 0.08V, peak factor 4.2, kurtosis 3.5, short-time energy 0.0064 J; and frequency-domain features: spectral centroid 1200 Hz, spectral bandwidth 800 Hz, natural frequency 850 Hz, harmonic amplitude ratio 0.6. These features are combined to generate a standardized feature vector: [0.08, 4.2, 3.5, 0.0064, 1200, 800, 850, 0.6]. For example, the audio data can be collected using an audio acquisition device, which can be quickly fixed to the surface of the building structure under test using components such as rubber suction cups, magnetic suction cups, and backpack vacuum generators combined with sealing ring suction cups. After testing, it can be quickly disassembled and moved to the next measurement point.
[0036] Step S40: Input multi-dimensional audio features into the target random forest classification model, predict and output the detection results of the building structure to be tested through the target random forest classification model.
[0037] It should be noted that the target random forest classification model refers to an intelligent model that is pre-trained and optimized with a large amount of labeled data to establish a mapping relationship between audio features and the loosening state of building structures; the detection result refers to the classification result and the confidence level of the corresponding category of the connection of the building structure to be tested, which reflects the loosening state of the connection. For example, the loosening state includes three categories: normal, slightly loose, and severely loose.
[0038] Specifically: A pre-trained target random forest classification model is loaded from the memory of the detection device. This model has been trained, validated, and optimized using a large amount of building structure audio feature data. The standardized feature vector generated in step S20 is converted into the input format required by the model and input into the target random forest classification model. The model performs inference calculations, calculating the probability that the input feature vector belongs to one of three categories: normal, slightly loose, and severely loose. The category with the highest probability is selected as the final detection result, and the confidence level of that category is recorded. The detection result, confidence level, detection time, test point number, and corresponding feature vector are stored in a local database. The detection result is output, and the output method includes, but is not limited to, screen display, voice broadcast, and vibration feedback. For example, after inference, the model outputs the probabilities of the three categories: normal probability 5%, slightly loose probability 12%, and severely loose probability 83%. Therefore, the final detection result is "severely loose (confidence level 83%)".
[0039] This embodiment provides an audio-based method for detecting loosening of building envelope systems. It employs enhanced audio data, including derived enhanced audio data, to train a random forest classification model, effectively expanding the acoustic feature distribution range of the training samples and compensating for the shortcomings of insufficient real-world sample quantity and limited scene coverage. By striking the building structure under test based on preset striking parameters, the method ensures the standardization and consistency of striking excitation, avoiding unnecessary fluctuations in the collected audio data caused by differences in striking force, angle, and position, thus improving the reliability of subsequent feature extraction and model judgment. By extracting multi-dimensional audio features from the audio data, it overcomes the limitations of traditional single time-domain or frequency-domain feature analysis, comprehensively characterizing the acoustic properties of the structural striking response and improving detection accuracy. The trained random forest classification model intelligently judges the audio features, combining high training efficiency, strong anti-interference ability, and stable output results. It can adapt to the complex environment of engineering sites, achieving rapid and accurate structural loosening detection, and solving the problems of poor accuracy and low reliability in existing building envelope system loosening detection methods.
[0040] In one feasible implementation, please refer to Figure 2 The step S40, which involves predicting and outputting the detection results of the building structure under test using the target random forest classification model, may further include steps S41 to S44: Step S41: Input multi-dimensional audio features into the target random forest classification model. The random forest classification model is based on the ensemble of multiple decision trees. It should be noted that the pre-trained random forest classification model refers to an ensemble learning model that has been trained, validated, and optimized in advance using a large number of labeled audio feature datasets of building structures. Its core is a set of classifiers composed of multiple independent decision trees. Decision tree ensemble refers to a method that constructs multiple weak classifiers and fuses their outputs to obtain a strong classifier with better performance. It can effectively solve the problems of overfitting and poor generalization ability of a single decision tree.
[0041] Specifically, the pre-trained random forest classification model is loaded from the memory of the detection device; the model structure parameters are parsed, including the total number of decision trees, the maximum depth of each decision tree, the minimum number of sample splits, the split features of each node and the corresponding threshold; the standardized feature vector generated in step S20 is converted into the floating-point input tensor format required by the model; and the converted input tensor is distributed to each decision tree in the random forest model.
[0042] Step S42: Classify the multi-dimensional audio features using each decision tree and output the classification result, which is either loose or not loose. It should be noted that decision tree classification refers to the process by which a single decision tree starts from the root node, compares the feature values of the input feature vector at each node with a preset threshold, selects the corresponding branch to traverse downwards until it reaches a leaf node, and outputs the classification result corresponding to that leaf node; the splitting rules of each decision tree are automatically learned during the training process, and the splitting features and thresholds of different decision trees are different.
[0043] Specifically, each decision tree independently executes the following complete classification process: starting from the root node, obtain the splitting feature and splitting threshold corresponding to the current node; extract the feature value corresponding to the splitting feature from the input feature vector and compare it with the splitting threshold; if the feature value is less than or equal to the splitting threshold, proceed to the left child node; if the feature value is greater than the splitting threshold, proceed to the right child node; repeat the above steps until the leaf node is reached; output the classification result corresponding to the leaf node, and the classification result is a binary label: 0 represents not loosened, and 1 represents loosened.
[0044] Step S43: Based on the classification results corresponding to all decision trees, the majority voting method is used to determine the loosening judgment result, and the confidence level corresponding to the loosening judgment result is calculated at the same time. It should be noted that majority voting is a commonly used decision fusion method in ensemble learning. For binary classification problems, the classification result supported by more than half of the decision trees is taken as the final ensemble result. Confidence refers to the proportion of the number of decision trees that support the final decision result to the total number of decision trees. The value ranges from 0.5 to 1.0. The higher the confidence, the stronger the reliability of the decision result.
[0045] For example, the detection device collects the classification results of 100 decision trees and finds that 83 decision trees output 1 (loose) and 17 decision trees output 0 (not loose). Since 83 > 17, the final result is determined to be "loose", and the confidence level is calculated as 83 / 100 × 100% = 83%. The two results, "loose" and "83%", are temporarily stored in the memory.
[0046] Step S44: Use the loosening judgment result and the corresponding confidence level as the detection result of the building structure to be tested and output the detection result.
[0047] For example, the detection device displays "Loose (confidence 83%)" in red on the display screen, while simultaneously triggering a "beep" alarm and a flashing red LED (Light Emitting Diode).
[0048] This implementation uses a random forest classification model as the core detection method, balancing detection accuracy and inference speed. It is particularly suitable for field applications of embedded detection devices. By fusing the results of multiple decision trees through majority voting, the classification error of a single decision tree can be effectively reduced, significantly improving the reliability of the detection results. Compared to single models such as support vector machines and logistic regression, the random forest classification model has stronger anti-interference and generalization capabilities in complex field environments, effectively improving detection accuracy and robustness. Confidence scores quantify the reliability of the detection results, providing a scientific basis for rapid decision-making in engineering fields.
[0049] In one feasible implementation, the step of outputting the detection result in step S44 may further include steps S441 to S442: Step S441: Output the loosening judgment result and confidence level of the building structure under test to the display screen; For example, the loosening determination result and corresponding confidence level generated in step S33 are read from the memory; the display style is automatically matched according to the loosening determination result: if the loosening determination result is "not loose", the result is displayed in green font and the confidence level value is displayed in green; if the loosening determination result is "loose", the result is displayed in red font and the confidence level value is displayed in red; in the center of the main interface of the display screen, it is highlighted in large font and the format is "[determination result] (confidence level: XX%)"; in the lower area of the display screen, auxiliary information is displayed in small font, including but not limited to the test point number, test date and time.
[0050] Step S442: Compare the confidence level with the preset confidence threshold. When the confidence level is lower than the preset confidence threshold, output a re-examination prompt message to the display screen.
[0051] For example, a preset confidence threshold is read from the memory. The default threshold is 75%, which can be adjusted by the user in the system settings according to their needs, such as 85% for high-security buildings. The confidence level calculated in step S33 is compared with the preset confidence threshold. If the confidence level is greater than or equal to the preset confidence threshold, the result is deemed reliable, and the detection process ends normally. If the confidence level is lower than the preset confidence threshold, the result is deemed unreliable, and the following prompts are executed: a prompt box pops up in the center of the display screen, showing the text "The current result has low confidence; it is recommended to re-detect!" The corresponding building structure points to be tested are automatically marked as "to be re-inspected," and the point number and detection time are stored in the re-inspection task list.
[0052] This implementation outputs the loosening determination result and confidence level to the display screen for intuitive and clear visualization. Inspectors can quickly obtain the test conclusion at a glance without the need for complex data analysis. By comparing the confidence level with the preset confidence threshold, it automatically identifies low reliability results and issues a re-inspection prompt in a timely manner, which can effectively avoid misjudgment and missed detection. Based on the original technical solution, it improves the ease of use, accuracy and reliability of the test method.
[0053] In one feasible implementation, step S20, which involves striking the building structure under test based on preset striking parameters, may further include steps S21 to S22: Step S21: Obtain preset striking parameters and transmit the preset striking parameters to a preset vibrator. The preset striking parameters include at least striking force and striking frequency. It should be noted that the preset impact parameters refer to the set of quantitative control commands stored in the detection device for precise control of the vibrator's working state. In addition to the core impact force and impact frequency, they may also include auxiliary parameters such as impact angle, impact duration, interval between two impacts, and vibrator head rebound speed. The vibrator is an electric vibration device at the front end of the detection device, connected to a replaceable hammer head assembly. It can convert electrical signals into mechanical impact force and is the core component for generating standardized excitation sources.
[0054] Step S22: Control the vibrator to strike the building structure under test based on preset striking parameters.
[0055] It should be noted that controlling the vibrator to strike refers to the process by which the vibrator, according to the received control command, drives the electromagnetic coil to generate electromagnetic force, which drives the hammer head to move in a straight line, applying an instantaneous impact force to the surface of the building structure to be tested.
[0056] For example, testers can set the striking force and frequency via a remote controller, or select target striking parameters from a locally stored menu based on the material type, thickness, or installation method of the building structure under test. The remote controller packages these parameters into control commands according to a preset communication protocol format and transmits the control commands to the vibrator via the network. After receiving the control commands, the vibrator adjusts the angle of the hammer head so that the angle between the hammer head axis and the surface of the building structure under test is equal to the preset striking angle, drives the hammer head to move forward slowly until it makes slight contact with the surface of the building structure under test. Based on the preset striking parameters, the vibrator calculates the magnitude of the driving current and the energizing time of the electromagnetic coil. According to the preset striking frequency, the vibrator outputs the corresponding driving current to the electromagnetic coil, driving the hammer head to impact the surface of the building structure at a preset speed, completing a single strike. After the strike is completed, the hammer head automatically returns to its initial position.
[0057] This implementation method transmits preset impact parameters to the vibrator, achieving standardized transmission and execution of the impact parameters and ensuring the consistency of the excitation source from the outset. The adjustable preset impact parameters can accurately adapt to the testing needs of different types of building structures. Automated impact operations reduce human error, improve testing efficiency, and further enhance the accuracy and reliability of the testing method.
[0058] In one feasible implementation, step S30, which involves extracting multi-dimensional audio features from the audio data, may further include steps S31 to S32: Step S31: Read and preprocess audio data, and output standardized audio data. The preprocessing method includes at least one of the following: format standardization, endpoint detection and effective tap signal extraction, DC component removal and baseline correction, noise suppression, and outlier detection and removal. It should be noted that audio preprocessing refers to a series of standardization and purification processes performed on the raw audio data to eliminate various interferences and deviations introduced during the acquisition process. Specifically: format standardization refers to unifying the sampling rate, bit depth, and number of channels of the audio data, eliminating data format differences caused by different acquisition devices or parameter settings; endpoint detection and valid tap signal extraction refers to automatically identifying the start and end positions of valid tap segments in the audio signal and removing invalid silent segments before and after them; DC component removal and baseline correction refers to eliminating DC offset and slowly changing baseline drift in the audio signal, bringing the average signal value to zero; noise suppression refers to filtering out environmental background noise mixed in during the acquisition process, such as wind noise, human voices, and equipment operating noise, improving the signal-to-noise ratio; and abnormal sample detection and removal refers to automatically identifying and removing invalid or abnormal audio data generated due to operational errors, equipment malfunctions, etc.
[0059] For example, the detection device acquires a segment of raw audio data, formatted as 44.1 kHz, 16-bit stereo, lasting 3 seconds. After format normalization, it is converted to 48 kHz, 16-bit mono audio. Through endpoint detection, the valid tapping signal is identified from 0.12 seconds to 1.45 seconds, and this segment is extracted. The DC component of the signal is calculated to be 0.02 V, and 0.02 V is subtracted from each sampling point to complete DC removal. Using the first 0.1 seconds of silence, the ambient noise is estimated to be mainly low-frequency air conditioning noise, which is filtered out using spectral subtraction. Finally, the valid signal length is checked to be 1.33 seconds, and the peak amplitude is 0.82 V, both within a reasonable range, thus it is determined to be a valid sample, and normalized audio data is output.
[0060] Step S32: Extract multi-dimensional audio features from standardized audio data. The multi-dimensional audio features include time-domain features, frequency-domain features, and 13th-order Mel frequency cepstral coefficients. The time-domain features include at least one of the following: root mean square energy, zero-crossing rate, peak amplitude, and total energy. The frequency-domain features include at least one of the following: fundamental frequency, main frequency band energy distribution, and spectral centroid.
[0061] It should be noted that multi-dimensional audio features refer to a set of feature vectors that describe the characteristics of audio signals from different perspectives. By combining features in the time domain, frequency domain, and based on the characteristics of human hearing, they can comprehensively and accurately characterize the vibration and sound characteristics produced when a building structure is struck. Among them, the 13th-order Mel-Frequency Cepstral Coefficients (MFCCs) are features extracted based on the characteristics of human hearing. They can simulate the differences in human hearing perception of different frequencies of sound and have a strong ability to characterize subtle features of sound. They are one of the most widely used features in the field of sound recognition.
[0062] Specifically, the following core temporal features are calculated for standardized audio data: Root mean square energy: Calculated by the square root of the sum of squares of all sample points, reflecting the average energy level of the signal; Zero crossing rate: Calculates the number of times a signal crosses zero per unit time, reflecting the frequency and impulse characteristics of the signal; Peak amplitude: The maximum absolute value of all sampling points in the signal is extracted, reflecting the maximum impact intensity of the strike; Total Energy: Calculates the sum of squares of all sampling points, reflecting the total energy generated by the impact.
[0063] Perform a Fast Fourier Transform on the standardized audio data to obtain the signal's spectrum, and then calculate the following core frequency domain features: Fundamental frequency: The fundamental frequency of the signal is calculated using the autocorrelation method, which is the lowest natural frequency of the building structure vibration; Main frequency band energy distribution: The spectrum is divided into five equal-width sub-bands: 0~1 kHz, 1~2 kHz, 2~3 kHz, 3~4 kHz, and 4~5 kHz. The proportion of energy in each sub-band to the total energy is calculated. Spectral centroid: Calculates the weighted average frequency of the spectrum, with the weights being the amplitude of each frequency point, reflecting the location of the main frequency distribution of the signal.
[0064] The signal is passed through a first-order high-pass filter to boost the energy of the high-frequency components and compensate for high-frequency attenuation during sound propagation. The pre-emphasized signal is then framed, and each frame is multiplied by a Hamming window to reduce spectral leakage. A Fast Fourier Transform (FFT) is performed on each frame to obtain its power spectrum. The power spectrum is then passed through a 24-mel-scale triangular filter bank to obtain the Mel spectrum. The logarithm of the Mel spectrum is then taken to obtain the logarithmic Mel spectrum. A Discrete Cosine Transform (DCT) is performed on the logarithmic Mel spectrum, and the first 13 coefficients are used as the MFCC features for that frame. The average value of the MFCC features across all frames is calculated to obtain a 13-dimensional MFCC feature vector for the entire normalized audio data. The extracted time-domain features, frequency-domain features, and MFCC features are then concatenated in sequence to form a 20-dimensional normalized multi-dimensional audio feature vector.
[0065] This implementation improves the quality of audio data by preprocessing it. By extracting 20-dimensional multi-dimensional audio features from standardized audio data and introducing MFCC features, it enhances the ability to identify subtle features, comprehensively characterizes the characteristics of the audio, and improves the ability to capture and represent loose features. It solves the problems of high data noise, single features, and insufficient representation of loose features in the prior art, and significantly improves the accuracy and precision of building structure loosening detection.
[0066] Example 2 Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Step S10 may also include steps A10 to A30: Step A10: Tap the building structure samples with different fastening conditions and collect the experimental audio data corresponding to each building structure sample. It should be noted that the building structure sample refers to a standardized test component that is consistent with the actual building structure to be tested in terms of material, thickness, and installation method. It is used to collect audio data of different loose states in a controlled environment. Different tightness states refer to multiple levels of looseness of the building structure simulated by adjusting the torque value of the connectors, including but not limited to no looseness, slight looseness, moderate looseness, and severe looseness. The real label refers to the actual tightness state label corresponding to each experimental audio data, which is the core basis for training the machine learning model.
[0067] Specifically, the researchers prepared multiple test samples with the same material, thickness, and installation method as the actual building structure to be tested. By adjusting the tightening torque of the connectors with a torque wrench, they simulated building structure samples with different tightening states. Data was collected in a quiet laboratory environment using a vibrator, microphone, and data acquisition card that were exactly the same as those used in on-site testing. Professional testing personnel labeled all the collected audio data with the corresponding tightening state labels, and invalid data generated due to knocking errors, equipment failures, etc., were discarded.
[0068] Step A20: Perform sample augmentation on the experimental audio data to obtain augmented audio data; It should be noted that sample augmentation refers to the technique of generating a large number of new audio data samples by performing a series of transformations on the original audio data without changing the core semantic information of the audio data. Its purpose is to expand the size of the training dataset, increase the diversity of data, solve the problem of insufficient sample data in practical applications, and improve the robustness of the model to different environmental interferences.
[0069] Specifically, sample enhancement processing is performed on all the audio data collected in step A10. Sample enhancement processing can employ one or more of the following enhancement methods: time stretching, pitch transformation, volume adjustment, adding background noise, spectrum masking, frequency shifting, etc.
[0070] Step A30: Based on the experimental audio data, the enhanced audio data, and the real labels corresponding to the experimental audio data and the enhanced audio data, the preset initial random forest classification model is trained to obtain the target random forest classification model.
[0071] It should be noted that the initial random forest classification model refers to a random forest classification model with a fixed network structure or algorithm framework that has not been trained; model training refers to the process of inputting labeled feature data into the initial random forest classification model and continuously adjusting the model parameters through optimization algorithms, enabling the model to learn the mapping relationship between audio features and fixed states; the target random forest classification model refers to the final random forest classification model that has been trained and optimized to achieve preset performance indicators and can be used for field detection.
[0072] Specifically, the experimental audio data obtained in step A10 is merged with the enhanced audio data obtained in step A20 to construct a complete training dataset. A stratified sampling method is used to divide the training dataset into three sub-datasets: a training set, a validation set, and a test set, ensuring that the proportion of samples with different tight states in each sub-dataset is consistent with the training dataset. A random forest classification model is initialized, setting core parameters such as the number of decision trees, the maximum depth of each tree, the minimum number of sample splits, and the maximum number of features. The feature vectors of the training set and their corresponding ground truth labels are input into the initial random forest classification model. The Gini coefficient is used as the node splitting criterion to train each decision tree. A 5-fold cross-validation method is used to comprehensively evaluate the model, calculating the average accuracy, precision, and recall. A grid search method is used to optimize key hyperparameters of the model, such as the number of decision trees, the maximum depth, and the minimum number of samples for node splits. The hyperparameter combination that performs best on the validation set is selected as the final model parameters. The optimized model was subjected to final performance testing using an independent test set. The confusion matrix of the model on the test set was calculated, and the recognition accuracy for different tight states was analyzed. If the overall accuracy of the model reached above 90%, and the recall rate for each class was not less than 90%, the model was deemed to have met the performance standard; otherwise, supplementary sample data was added for retraining. Finally, the trained target model was serialized into a standard format file and saved to the memory of the detection device.
[0073] For example, sample augmentation was performed on the 588 original experimental audio data obtained in step A10. For each original data point, time stretching, volume adjustment, and background noise addition were randomly selected and combined to generate 8 augmented audio data points. A total of 588 × 8 = 4704 augmented audio data points were generated. After quality verification, 32 invalid data points were removed, resulting in 4672 valid augmented audio data points. All augmented audio data points inherited the true labels of the original data. The dataset of 5260 data points was divided into training, validation, and test sets in a 7:2:1 ratio, and 20-dimensional feature vectors were extracted from all data. A random forest model was initialized with 100 decision trees and a maximum depth of 12. After training and 5-fold cross-validation, the optimal parameters were obtained through grid search optimization: 120 decision trees and a maximum depth of 10. The model was finally tested using the test set. The overall accuracy was 95.8%, the recall rate was 97.2% for the fully secure class, 94.1% for the slightly loose class, 95.3% for the moderately loose class, and 96.7% for the severely loose class. All metrics met the preset requirements. The model was saved as a target random forest classification model and deployed to the field detection device.
[0074] This implementation method collects sample data in a controlled laboratory environment using equipment and parameters identical to those used in the field, ensuring consistency in the distribution of training data and field detection data. By employing sample augmentation techniques to simulate various environmental changes that may be encountered in the field, it not only effectively solves the problems of difficult and insufficient sample data collection in practical applications but also significantly increases the diversity of training data, substantially improving the model's generalization ability and environmental robustness. This addresses the issues of poor model generalization ability, environmental sensitivity, and low detection accuracy in existing technologies.
[0075] In one feasible implementation, the enhanced audio data includes physically enhanced audio data. The step of sample enhancement of the experimental audio data in step A20 to obtain enhanced audio data may further include steps A21-A22: Step A21: Superimpose various preset ambient background noises onto the experimental audio data to obtain physically enhanced audio data; It should be noted that physically enhanced audio data refers to enhanced audio data generated by simulating the acoustic propagation process in the real physical world, which can realistically reproduce the acoustic environment during on-site testing; on-site background noise refers to various environmental noises that actually exist at the building testing site, and is one of the main interference factors that cause a decrease in on-site testing accuracy.
[0076] For example, a background noise database is constructed: noise is collected at more than 10 typical building inspection sites, covering different scenarios such as construction sites, office buildings, shopping malls, residences, and bridges. The collection equipment and the on-site inspection device use the same microphone and data acquisition card, and the background noise is collected based on a preset duration. The collected noise is classified and labeled, which can include construction machinery such as tower crane operation, cutting machine, electric drill, and mixer; traffic such as car driving, horn, and subway operation; natural environment such as wind, rain, and thunder; human voice such as conversation, broadcast, and footsteps; equipment operation such as air conditioner, elevator, and water pump; building structure such as door and window opening and closing and pipe vibration; electromagnetic such as current noise and radio interference; and other types such as glass breaking and metal collision.
[0077] Specifically, for each original experimental audio data, different types of noise sub-segments are randomly selected from the on-site background noise library. According to the preset signal-to-noise ratio range, the noise gain coefficient to be superimposed is calculated. A segment of the same length as the experimental audio is randomly extracted from the noise sub-segments, multiplied by the gain coefficient, and then superimposed onto the original audio. Invalid enhanced audio data that cause severe signal distortion due to noise superposition is removed.
[0078] And / or, in step A22, based on a preset reverb algorithm, apply reverb effects for various scenarios to the experimental audio data to obtain physically enhanced audio data.
[0079] It should be noted that reverberation refers to the acoustic phenomenon that occurs when sound travels through a closed or semi-closed space and is reflected multiple times by objects such as walls, floors, and ceilings. Different sizes, shapes, and materials of spaces will produce different reverberation characteristics. The on-site environment of building testing is complex and diverse, ranging from open outdoor plazas to enclosed indoor halls, with huge differences in reverberation. This is another important reason why the performance of models trained in the laboratory deteriorates in the field.
[0080] Specifically, typical spatial scenarios most commonly encountered in building inspection, such as outdoor open spaces, corridors, offices, lobbies, stairwells, and elevator shafts, were selected. In each scenario, professional acoustic measurement equipment was used to play a logarithmic sweep signal while simultaneously recording the signal received by the microphone. The impulse response of the scenario was calculated through inverse filtering. Key acoustic parameters such as reverberation time, early reflection delay, intelligibility, and speech transmission index were measured and labeled to construct a reverberant impulse response library. For each original experimental audio data, an impulse response of a typical scenario was randomly selected from the reverberant impulse response library. Based on a preset parameter range, parameters such as reverberation time and reflection coefficient were randomly adjusted to generate a fine-tuned impulse response. The original audio and the fine-tuned impulse response were linearly convolved to obtain the audio with added reverberation. The convolved audio was normalized to adjust the amplitude to a reasonable range to avoid distortion. Invalid enhanced audio data with abnormal reverberation effects or severe signal distortion were discarded.
[0081] In this embodiment, the methods in steps A21 and A22 can be used alone for audio data or superimposed with physical enhancements for audio data. For example, multiple preset ambient background noises are superimposed onto the experimental audio data to obtain first physically enhanced audio data; based on a preset reverberation algorithm, reverberation effects from multiple scenes are applied to the experimental audio data to obtain second physically enhanced audio data; the first and second physically enhanced audio data are then merged to obtain the final physically enhanced audio data. As another example, multiple preset ambient background noises are superimposed onto the experimental audio data to obtain first physically enhanced audio data; based on a preset reverberation algorithm, reverberation effects from multiple scenes are applied to the first physically enhanced audio data to obtain second physically enhanced audio data; the second physically enhanced audio data is then used as the final physically enhanced audio data.
[0082] This implementation addresses the core issues of insufficient realism in traditional digital augmentation methods and domain shift between laboratory and field data by superimposing real-world noise and reverberation from multiple scenes onto the experimental audio data used to train the random forest classification model. This significantly improves the detection accuracy and robustness of the target random forest classification model in complex field environments, thereby enhancing the accuracy of building structure loosening detection methods.
[0083] In one feasible implementation, the derived enhanced audio data includes first derived data and / or second derived enhanced audio data. The step of sample enhancement of the experimental audio data in step A20 to obtain enhanced audio data may further include steps A23 to A27: Step A23: Map the experimental audio data to a low-dimensional latent space using the encoder of the preset variational autoencoder, and obtain the latent space mean vector corresponding to the experimental audio data. It should be noted that derived enhanced audio data refers to entirely new sample data automatically generated after a generative machine learning model learns the temporal distribution patterns of samples in real loose and non-loose samples. Derivative enhancement can generate new samples that do not appear in the original experimental data but conform to the distribution patterns of real samples of loose and non-loose building structures, which can greatly expand the boundaries of the dataset. Variational autoencoder (VAE) is a generative neural network based on a probabilistic graphical model. It can map high-dimensional original data to a low-dimensional continuous latent space and generate new data through sampling of the latent space. It has the advantages of stable training and controllable generation quality. The latent space mean vector corresponds to the standard representation of the original audio in the latent space, while Gaussian perturbation is used to simulate the feature shift caused by individual structural differences.
[0084] Step A24: Introduce a Gaussian perturbation into the latent space mean vector to obtain the perturbation latent vector; It should be noted that Gaussian perturbation refers to adding normally distributed random noise to the latent space mean vector. By utilizing the continuity and local similarity of the VAE latent space, a new sample similar to the original sample but with subtle differences is generated. Neighboring vectors in the latent space correspond to similar audio features. Therefore, by making a small perturbation to the mean vector of the original sample, a new sample that retains the original category features while having a certain degree of diversity can be generated without semantic bias. The perturbation latent vector is the latent space vector obtained by superimposing Gaussian perturbation, corresponding to the latent space representation of the original audio after a small feature shift.
[0085] Specifically, a VAE model consisting of a one-dimensional convolutional neural network encoder and a symmetric transposed convolutional decoder was constructed. A 128-dimensional latent space was set, and the VAE was trained using all experimental audio data. The loss function was the sum of the reconstruction loss and the KL divergence. After training, each original audio track was input into the encoder to extract the corresponding 128-dimensional latent space mean vector. Gaussian noise with a mean of 0 and a standard deviation of 0.1 was added to the latent space mean vector of each original audio track. The perturbed latent vector was then input into the VAE decoder to generate a new audio signal. Valid generated samples were selected through feature similarity verification and classification prediction verification, and the original labels were inherited. Invalid generated samples with low similarity or incorrect classification prediction were removed.
[0086] Step A25: The perturbation latent vector is input into the decoder of the variational autoencoder and reconstructed to generate the first derived enhanced audio data; It should be noted that waveform reconstruction is the process by which the decoder converts the latent vectors into a complete audio waveform. The optimization objective of model training is the loss function, which consists of two parts: reconstruction error and KL divergence. The reconstruction error measures the waveform difference between the reconstructed audio and the original audio, and is used to ensure the fidelity of the generated samples. The KL divergence measures the difference between the latent variable distribution and the standard Gaussian distribution, and is used to ensure the regularity and continuity of the latent space, and to ensure that the feature changes after perturbation are smooth and controllable.
[0087] Additionally, it should be noted that the encoder is a one-dimensional convolutional network at the front end of the model, whose function is to compress and map the high-dimensional audio waveform into probability distribution parameters in the low-dimensional latent space, namely the mean and logarithmic variance; the decoder is a one-dimensional transposed convolutional network at the back end of the model, whose function is to reconstruct the vectors in the low-dimensional latent space and restore them to the audio waveform with the same length and sampling rate as the original input.
[0088] For example, constructing by encoder and decoder A variational autoencoder is constructed. The input is a high signal-to-noise ratio seed audio waveform. ,in L The number of sampling points is used to obtain the mean of the posterior distribution of the latent variables through a one-dimensional convolutional encoder. Sum of logarithmic variance And through reparameterized sampling, the latent space mean vector is obtained. The data processing process can be represented by formula (1): (1) Where z is the latent space mean vector. It is a random vector. It follows a multivariate standard normal distribution.
[0089] decoder Will Reconstructed (Reconstructed audio waveform output by the decoder), training loss function The data processing can be expressed by formula (2) as the sum of the reconstruction error and the KL divergence: (2) in, The weight hyperparameters are used to balance reconstruction fidelity and latent space regularity. KL divergence measures the distance between the posterior distribution of the data obtained by the encoder and the pre-defined standard normal prior distribution. After training, for each seed sample, the encoder infers the mean of its latent variables. And by introducing a small Gaussian perturbation, the data processing can be represented by formula (3): (3) in, For perturbation latent vectors, The wave constant is The coefficient corresponding to the fluctuation constant is used to simulate subtle fluctuations in acoustic characteristics caused by manufacturing tolerances such as panel thickness and local stiffness. Input to the decoder to generate perturbation reconstruction samples This is the first derived enhanced audio data, with the same label as the original seed. Due to the continuity of the latent space, this small perturbation mainly causes a slight shift in the resonant frequency of the impact response, while the temporal envelope and attenuation characteristics are maintained, thus effectively simulating the spectral drift caused by individual manufacturing differences in the same batch of mounting plates.
[0090] And / or, in step A26, a second derived enhanced audio data is generated based on multiple trained generative adversarial networks, wherein the generative adversarial networks include a WaveGAN model, and the multiple trained generative adversarial networks are independently trained based on experimental audio data according to the category of the tightness state of the building structure samples. It should be noted that Generative Adversarial Network (GAN) is a generative model that learns data distribution through adversarial game between generator and discriminator, and can generate more realistic and diverse samples; Wave GAN is a GAN model specifically designed for audio data. It adopts one-dimensional convolution and transposed convolution structure, and can directly generate the original audio waveform, avoiding the information loss caused by spectrum conversion. Specifically, the experimental audio was divided into four subsets: completely tight, slightly loose, moderately loose, and severely loose. A Wave GAN or one-dimensional diffusion model was independently trained for each subset. New samples were generated using the trained category-specific models. More data was generated for categories with fewer samples to balance the distribution. High-quality generated samples were selected based on Fraser audio distance and subjective evaluation, and labeled with the category to which the corresponding model belonged. Experienced testers then subjectively evaluated the generated samples, assessing authenticity, distinguishability, and absence of background noise. Invalid generated samples with excessively high Fraser audio distance values or unsatisfactory subjective evaluations were discarded.
[0091] For example, taking Conditional WaveGAN as an example, the generator With random noise and category labels As input, the audio is progressively upsampled through a one-dimensional transposed convolution to output fake audio. Discriminator This is a one-dimensional convolutional network that receives audio and labels, and outputs true / false scores. WGAN-GP loss is introduced to stabilize the training process, and the training loss function is... The data processing procedure can be represented by formulas (4) and (5): (4) (5) in, For the distribution of real audio datasets, This indicates that the average of a batch of real samples is calculated. The interpolated samples are composed of the real sample x and the generated sample x. Linear interpolation is obtained. The gradient penalty coefficient is... The loss of generator G, The discriminant score is the score that the discriminator outputs for the generated samples. This represents the average discrimination score of the discriminator for the generated fake audio. After training convergence, new samples with high-fidelity time-domain waveforms and frequency-domain features can be generated in batches by randomly sampling noise and specifying categories.
[0092] Alternatively, in step A27, a second derived enhanced audio data is generated based on the trained diffusion model, wherein the trained diffusion model is obtained by iterative processes of forward noise addition and reverse noise reduction based on experimental audio data.
[0093] It should be noted that one-dimensional diffusion models are a new generation of generative models that have emerged in recent years. They generate data by progressively removing noise, and have the advantages of high generation quality, stable training, and fewer model collapses. Training independently by category means training a separate generative model for each fixed state. This allows each model to focus on learning the data distribution of a single category, avoiding feature confusion between different categories, and significantly improving the label accuracy of generated samples.
[0094] Specifically, using experimental audio data as the training basis, a forward noise addition process is first defined, where Gaussian noise is progressively added to clean audio according to a preset total number of diffusion steps. Each step corresponds to a fixed noise intensity, and after all steps, the audio is completely transformed into random Gaussian noise. A noise prediction network with a one-dimensional convolutional neural network as its core is constructed, and the noisy frequencies at different time steps and the corresponding step information are used as network inputs to train the network to predict the real noise added in the current step. The mean square error between the predicted noise and the real noise is used as the denoising loss, and the network parameters are iteratively optimized through backpropagation until the loss converges and the network denoising accuracy reaches the target. After training, starting from pure Gaussian noise, noise is gradually removed through the noise prediction network in the reverse order of forward noise addition, and finally, high-fidelity second-derived enhanced audio data is generated.
[0095] For example, a one-dimensional diffusion model can be used to generate the second derived enhanced audio data. First, the forward noise addition process is defined for the one-dimensional diffusion model, and the data processing process can be represented by formula (6): (6) in, Given a forward diffusion conditional probability distribution, the noisy frequency at time t-1 is known. The frequency with noise at time t is obtained. Gaussian distribution, Let be the noise scheduling coefficient at time t. The mean term of a Gaussian distribution. The covariance matrix is a Gaussian distribution. This represents the total diffusion time step. A noise prediction network based on U-Net is then trained. Minimize the denoising loss, loss function The data processing procedure can be represented by formulas (7) and (8): (7) , (8) in, This indicates all clean audio files within the batch. Random noise Calculate the mean value at random sampling time step t. Used to measure the difference between the actual noise and the network's predicted noise. The cumulative signal retention coefficient, This indicates traversing all time steps from s=1 to s=t.
[0096] In actual reasoning, from Gaussian noise Initially, high-quality audio is generated through gradual noise reduction via back diffusion. Starting from pure noise, it gradually recovers into a clear percussion audio, and its energy distribution and formant structure closely approximate the real percussion sound.
[0097] In this embodiment, the methods in steps A23-A25, A26, and A27 can be used individually for audio data or superimposed on derived enhancements for audio data. For example, based on a preset variational autoencoder, experimental audio data is mapped to a low-dimensional latent space to obtain the latent space mean vector corresponding to the experimental audio data; Gaussian perturbation is introduced into the latent space mean vector, and the perturbated latent vector is input into the decoder of the variational autoencoder for reconstruction to generate first derived enhanced audio data; second derived enhanced audio data is generated based on multiple trained generative adversarial networks; and the first and second derived enhanced audio data are merged as derived enhanced audio data.
[0098] Preferably, please refer to Figure 4 The methods in steps A23 to A27 are combined with steps A21 to A22 in the previous embodiment, so that the enhanced audio data in step A20 includes physical enhanced audio data and derived enhanced audio data. For specific implementation methods, please refer to the detailed description of the above steps, which will not be repeated here.
[0099] This implementation employs a multi-path enhancement design that generates first derived enhanced audio data through latent space perturbation using a variational autoencoder and second derived enhanced audio data through a generative adversarial network or diffusion model. This design can accurately simulate the subtle acoustic feature drift caused by structural manufacturing tolerances, enriching the fine-grained feature distribution of samples. It can also synthesize high-fidelity new samples in batches through generative models, significantly expanding the scale of training samples. At the same time, the method of training the generative model independently by category ensures the accuracy of the generated sample categories. Multiple enhancement paths can be flexibly combined or selected individually, effectively solving the problems of insufficient real-world samples and limited feature coverage. This significantly improves the generalization ability and detection accuracy of the target random forest classification model, thereby enhancing the accuracy of the building structure loosening detection method.
[0100] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the building structure loosening detection method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0101] This application also provides an audio impact detection device, which includes: The vibration module, including a vibrator and a hammer, is used to strike the building structure under test based on preset striking parameters. The audio acquisition module, including a microphone, preamplifier, acquisition card, and data storage device, is used to acquire and store audio data generated by the building structure under test. The data analysis and control module, including a microprocessor, battery, remote controller, and display screen, is used to control the excitation module and audio acquisition module and output the test results of the building structure under test. The fixing module, including the housing and fixing unit, is used to adsorb and fix the entire device to the surface of the building structure to be tested.
[0102] For example, to aid in understanding this audio tapping detection device, this application provides a specific implementation of the audio tapping detection device: Vibration module: Consists of an electromagnetically or piezoelectrically driven vibrator and replaceable hammerheads. The vibrator outputs a constant frequency and force of impact, with adjustable force ranging from 5 to 20 N and adjustable frequency from 1 to 5 Hz. The hammerheads are made of rubber or nylon with a Shore hardness of 60 to 70, approximately 15 to 20 mm in diameter, and have a slightly convex spherical end face. This design effectively transmits excitation energy while avoiding damage to the surface of the tested panel and ensuring consistent contact points.
[0103] The audio acquisition module includes a retractable, flexible high-precision microphone, a preamplifier, a 16-bit AD acquisition card, and data storage. The microphone has a frequency response range of 20 Hz to 10 kHz, with a sampling point 5 to 10 cm from the impact point, used to acquire the audio signal radiated after the structure is struck. The preamplifier gain is adjustable from 10 to 100 times and integrates an anti-aliasing filter (cutoff frequency 10 kHz). The AD acquisition card supports multi-channel synchronous acquisition with a sampling frequency of at least 50 kHz. All data is stored in the memory, supporting local memory card and cloud upload. The high-precision microphone can be replaced with a MEMS digital microphone array, using beamforming algorithms to directionally amplify the sound signal in the direction of the impact, further suppressing environmental noise.
[0104] The data analysis and control module consists of a microprocessor, a lithium battery pack, a remote controller, and a 2.4-inch LCD display, all integrated within the housing. The microprocessor runs an algorithm program for detecting component loosening. The remote controller connects to the vibrator and audio acquisition module via 4G or Bluetooth to control the start / stop of tapping and audio acquisition. The LCD display shows the tapping status, signal quality, loosening determination result, and corresponding confidence level in real time. The battery pack has a battery life of ≥8 hours and supports fast charging. Shielded cables are used for all components to reduce electromagnetic interference.
[0105] Mounting Module: Includes the housing and mounting components. The housing has an IP 65 protection rating, suitable for dusty and humid outdoor environments. The mounting components allow the entire device to be quickly and easily attached to the surface of the external protective system. After testing, it can be easily disassembled and moved to the next testing point. These components can be rubber suction cups, magnetic suction cups, or backpack vacuum generator combined with sealing ring suction cups, etc. The modules are detachably connected via bolts or snap-fit structures, facilitating disassembly and maintenance.
[0106] Compared with existing technologies, the audio impact detection device provided in this application has a high degree of integration, allowing for immediate testing and operation. The adjustable impact parameters of the vibrator ensure consistent excitation energy for each impact. The audio acquisition module uses a high-precision microphone to acquire high-fidelity audio signals and ensures complete recording of transient responses. Its IP65 protection rating and long-lasting battery meet the all-weather operation requirements of dusty and humid outdoor curtain walls. The display screen provides visual output of the test results, increasing the ease of use of the audio impact detection device. The audio impact detection device provided in this application, employing the building structure loosening detection method described in the above embodiments, can solve the technical problems of poor accuracy and low reliability in building envelope system loosening detection methods.
[0107] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the building structure loosening detection method in Embodiment 1 above.
[0108] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0109] like Figure 5As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. While electronic devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0110] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0111] The electronic device provided in this application, employing the building structure loosening detection method described in the above embodiments, can solve the technical problems of poor accuracy and low reliability in building envelope system loosening detection methods. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the building structure loosening detection method provided in the above embodiments, and other technical features of the electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0112] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0113] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0114] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the building structure loosening detection method described in the above embodiments.
[0115] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0116] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0117] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to: Acquire enhanced audio data, and train a preset initial random forest classification model based on the enhanced audio data to obtain a target random forest classification model, wherein the enhanced audio data includes at least derived enhanced audio data; The building structure under test is struck based on preset striking parameters; Audio data generated by the building structure under test is collected, and multi-dimensional audio features of the audio data are extracted. The multi-dimensional audio features include at least time domain features and frequency domain features. Input multi-dimensional audio features into the target random forest classification model, and use the target random forest classification model to predict and output the detection results of the building structure to be tested.
[0118] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0120] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0121] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described method for detecting loose building structures. This solves the technical problems of poor accuracy and low reliability in methods for detecting loose building envelope systems. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the building structure loosening detection method provided in the above embodiments, and will not be repeated here.
[0122] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the building structure loosening detection method described above.
[0123] The computer program product provided in this application can solve the technical problems of poor accuracy and low reliability in the detection method of loose building envelope systems. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the building structure loosening detection method provided in the above embodiments, and will not be repeated here.
[0124] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for detecting structural loosening in buildings, characterized in that, The method includes: Acquire experimental audio data, and perform data enhancement on the experimental audio data based on a preset on-site background noise and / or a preset reverberation algorithm to obtain physically enhanced audio data; The experimental audio data is augmented based on a preset variational autoencoder, a pre-trained generative adversarial network, and / or a pre-trained diffusion model to obtain derived augmented audio data. The generative adversarial network is independently trained based on the experimental audio data according to the category of the tightness state of the building structure samples. The diffusion model is trained based on the experimental audio data through an iterative process of forward noise addition and backward noise reduction. The target random forest classification model is trained based on the experimental audio data, the physically enhanced audio data, the derived enhanced audio data, and the real labels corresponding to the experimental audio data, the physically enhanced audio data, and the derived enhanced audio data, respectively. The building structure under test is struck based on preset striking parameters; Audio data generated by the building structure under test is collected, and multi-dimensional audio features of the audio data are extracted, wherein the multi-dimensional audio features include at least time domain features and frequency domain features; The multi-dimensional audio features are input into the target random forest classification model, and the detection results of the building structure to be tested are predicted and output through the target random forest classification model.
2. The method as described in claim 1, characterized in that, The step of predicting and outputting the detection results of the building structure to be tested using the target random forest classification model includes: The multi-dimensional audio features are input into the target random forest classification model, which is based on an ensemble of multiple decision trees. The multi-dimensional audio features are classified using each decision tree, and a classification result is output, which is either loose or not loose. Based on the classification results corresponding to all decision trees, the loosening judgment result is determined by majority voting, and the confidence level corresponding to the loosening judgment result is calculated at the same time. The loosening determination result and the corresponding confidence level are used as the detection result of the building structure under test, and the detection result is output.
3. The method as described in claim 2, characterized in that, The step of outputting the detection result includes: Output the loosening determination result and confidence level of the tested building structure to the display screen; The confidence level is compared with a preset confidence threshold. When the confidence level is lower than the preset confidence threshold, a re-examination prompt message is output to the display screen.
4. The method as described in claim 1, characterized in that, The steps for obtaining experimental audio data include: Various building structure samples with different fastening conditions were tapped, and experimental audio data corresponding to each building structure sample were collected.
5. The method as described in claim 1, characterized in that, The step of enhancing the experimental audio data based on a preset background noise and / or preset reverberation algorithm to obtain physically enhanced audio data includes: Multiple preset ambient noises are superimposed onto the experimental audio data to obtain physically enhanced audio data; And / or, based on a preset reverb algorithm, apply reverb effects for various scenarios to the experimental audio data to obtain physically enhanced audio data.
6. The method as described in claim 1, characterized in that, The derived enhanced audio data includes first derived enhanced audio data and / or second derived enhanced audio data. The step of data augmentation of the experimental audio data based on a preset variational autoencoder, a pre-trained generative adversarial network, and / or a pre-trained diffusion model to obtain derived enhanced audio data includes: The experimental audio data is mapped to a low-dimensional latent space using a preset variational autoencoder, and the latent space mean vector corresponding to the experimental audio data is obtained. A Gaussian perturbation is introduced into the latent space mean vector to obtain the perturbation latent vector; The perturbation latent vector is input into the decoder of the variational autoencoder for reconstruction, generating the first derived enhanced audio data; Second derived enhanced audio data is generated based on multiple trained generative adversarial networks, wherein the generative adversarial networks include the WaveGAN model; Alternatively, a second derived enhanced audio data can be generated based on the trained diffusion model.
7. The method as described in claim 1, characterized in that, The step of striking the building structure under test based on preset striking parameters includes: The preset striking parameters are obtained and transmitted to a preset vibrator. The preset striking parameters include at least striking force and striking frequency. The vibrator is controlled to strike the building structure under test based on the preset striking parameters.
8. The method as described in claim 1, characterized in that, The step of extracting multi-dimensional audio features from the audio data includes: Read and preprocess the audio data, and output standardized audio data. The preprocessing method includes at least one of the following: format standardization, endpoint detection and effective tap signal extraction, DC component removal and baseline correction, noise suppression, and outlier sample detection and removal. Extract multi-dimensional audio features from the standardized audio data. The multi-dimensional audio features include time-domain features, frequency-domain features, and 13th-order Mel frequency cepstral coefficients. The time-domain features include at least one of the following: root mean square energy, zero-crossing rate, peak amplitude, and total energy. The frequency-domain features include at least one of the following: fundamental frequency, main frequency band energy distribution, and spectral centroid.
9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the building structure loosening detection method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the building structure loosening detection method as described in any one of claims 1 to 8.