A cement pavement intelligent acoustic void detection method, system, device and medium

By using unmanned driving equipment and acoustic detection technology, combined with noise suppression and multi-model classification, a high-precision distribution map of void defects in cement pavement is generated, which solves the problems of accuracy and efficiency in the detection of void defects in existing technologies and realizes efficient and accurate maintenance decision support.

CN122042821BActive Publication Date: 2026-07-21XIAN CHANGDA HIGHWAY MAINTENANCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAN CHANGDA HIGHWAY MAINTENANCE TECH CO LTD
Filing Date
2026-04-17
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies are insufficient for achieving centimeter-level precise location of voids in cement pavements and rapid inspection of large-scale road networks. Furthermore, the lack of effective fusion and in-depth mining of multi-source heterogeneous data results in limited reliability of detection results and maintenance decision support capabilities.

Method used

Unmanned equipment is used for acoustic wave detection. The road surface acoustic wave signal is obtained by applying mechanical impact through the acoustic vibration wheel. Combined with noise suppression, frequency domain transformation and multi-model classification, a standardized acoustic feature vector is generated to achieve spatiotemporal alignment and disease marking. Finally, a disease distribution map is generated on the cloud platform.

Benefits of technology

It has achieved efficient and accurate clearance detection, improved the automation level of detection and centimeter-level positioning accuracy, optimized the scientific nature of maintenance decisions, and provided key technical support for preventive maintenance of road infrastructure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122042821B_ABST
    Figure CN122042821B_ABST
Patent Text Reader

Abstract

The application relates to a cement pavement intelligent acoustic void detection method, system, device and medium, which comprises the following steps: after a user configuration instruction is acquired, an unmanned device is controlled to travel along a preset path, and a periodic mechanical impact is applied to a pavement by an acoustic vibration wheel. A pavement acoustic signal generated by the impact is acquired in real time, and spatial coordinate data and a timestamp are synchronously collected. Noise suppression and frequency domain conversion are performed on the acoustic signal to generate a standardized acoustic feature vector. The pavement acoustic signal, the feature vector, the coordinate data and the timestamp are input into a multi-source fusion module to generate an acoustic feature-position data packet. The feature vector is input into a pre-trained classification model group for real-time inference, and a void state identifier and a confidence degree are output. A disease marking instruction is triggered based on the coordinate data of the void state. A disease coordinate set is constructed, the acoustic signal is converted into a frequency domain feature map, and a spatialized disease distribution map is generated. The application realizes real-time detection and accurate positioning of a pavement void state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of road infrastructure testing technology, and in particular to a method, system, equipment and medium for intelligent acoustic void detection of cement pavement. Background Technology

[0002] With the continuous growth of my country's highway network, especially the operational mileage and service life of high-grade cement concrete pavements, the pavement structure is prone to hidden defects such as voids and loosening in its base layer under the long-term coupling effect of vehicle loads and environmental factors. These defects are initially hidden within the pavement structure and are difficult to detect effectively through traditional manual visual inspections, but they significantly weaken the road's load-bearing capacity, posing potential traffic safety hazards and imposing huge cost pressures on subsequent maintenance. To address this challenge, the industry has gradually developed non-destructive testing technologies such as ground-penetrating radar and falling-weighted deflectometers to achieve non-destructive detection of hidden defects. These technologies, based on the principles of electromagnetic wave reflection or mechanical impact, can penetrate the pavement structure to a certain extent to obtain internal information, contributing to the development of road inspection technology.

[0003] However, in practical applications, detection methods based on stress wave propagation characteristics often suffer from limitations in accuracy and depth due to the inhomogeneity of road materials and environmental noise interference, making it difficult to achieve centimeter-level precise location of road defects. Furthermore, some technical solutions rely on heavy machinery or complex on-site setups, resulting in low detection efficiency and failing to meet the needs of rapid inspection of large-scale road networks. Moreover, existing methods lack effective fusion and in-depth analysis of multi-source heterogeneous data during the detection process, limiting the reliability of defect identification results and their ability to support maintenance decisions. In particular, how to transform discrete point-like detection results into an intuitive, continuous spatial distribution pattern of road defects, and how to achieve visualized integration and intelligent analysis of detection information on a cloud platform, remain pressing technical problems that need to be solved. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a method, system, equipment, and medium for intelligent acoustic void detection of cement pavement.

[0005] Firstly, this application provides an intelligent acoustic void detection method for cement pavement, employing the following technical solution: A smart acoustic void detection method for cement pavement, the detection method comprising: Retrieve user configuration commands containing target area and driving speed parameters; In response to the user configuration command, the driverless device is controlled to travel along a preset path, while periodic mechanical impacts are applied to the road surface through the acoustic vibration wheel on the driverless device. The sound wave signal of the road surface generated by the mechanical impact is acquired in real time by the soundprint acquisition module, and the spatial coordinate data and timestamp of the unmanned driving equipment are acquired by the positioning system. The road surface acoustic wave signal is subjected to noise suppression and frequency domain transformation to generate a standardized acoustic signature feature vector. The road surface acoustic wave signal, acoustic feature vector, spatial coordinate data and timestamp are input into the multi-source fusion module to generate a spatiotemporally aligned acoustic-location data packet; The voiceprint feature vector is input into a pre-trained classification model group for real-time inference, generating a classification result that includes the empty state identifier and confidence level. When the void status indicator indicates a void status, a disease marking command is triggered based on the spatial coordinate data; The voiceprint-location data packet and classification results are associated with a structured data packet. When a disease marking instruction exists, it is associated with the data packet and uploaded to the cloud service platform. In the cloud service platform, a set of disease coordinates is constructed based on the structured data packet, and the road surface acoustic wave signal is converted into a frequency domain feature map; By fusing the disease coordinate set and frequency domain feature map, a spatialized disease distribution map is generated.

[0006] By adopting the above technical solution, a highly efficient and accurate method for detecting voids in cement pavement has been achieved. First, through acoustic vibration wheel excitation and synchronous data acquisition, physical impact is transformed into a quantifiable acoustic-spatial information flow, overcoming the subjectivity of traditional detection methods. Second, noise suppression and frequency domain transformation processing improve feature robustness, while a dual-model classifier ensures inference reliability. Finally, cloud-based data fusion and visualization generate an intuitive disease distribution map, achieving an upgrade from point-based detection to area-based management. This technical solution combines edge computing, machine learning, and geographic information technology, not only improving the automation level and centimeter-level positioning accuracy of detection but also optimizing the scientific nature of maintenance decisions through data closed-loop management, providing key technical support for preventative maintenance of road infrastructure.

[0007] Optionally, the step of performing noise suppression and frequency domain transformation on the road surface acoustic wave signal to generate a standardized acoustic signature feature vector includes: Acquire road surface acoustic wave signals in real time collected by the acoustic signature acquisition module; The acoustic signal is subjected to noise suppression processing, and the pulse interference noise is suppressed by spectral subtraction to generate a denoised wave signal; The denoised signal is divided into equal-length time segments, and a windowing function is applied to each frame. Perform frequency domain transformation on the windowed time segment of equal length to generate a frequency domain energy spectrum; Multi-scale acoustic signature features are extracted from the frequency domain energy spectrum to generate the original feature vector; The original feature vector is normalized to generate a standardized voiceprint feature vector.

[0008] By employing the above technical solutions, efficient feature extraction of road surface acoustic signals is achieved. Noise suppression and frame-by-frame windowing improve the time-frequency analysis quality of the signal, laying the foundation for feature extraction. Multi-scale feature fusion (such as MFCC and energy entropy) captures the acoustic fingerprint of delamination defects from both global and local dimensions, enhancing pattern discriminability. Normalization ensures the standardization and comparability of feature vectors, providing stable input for model inference. This technical solution closely integrates digital signal processing with machine learning feature engineering, not only overcoming noise interference in complex environments but also optimizing the accuracy and robustness of defect detection through multi-scale representation.

[0009] Optionally, the step of inputting the voiceprint feature vector into a pre-trained classification model group for real-time inference to generate a classification result containing the empty state identifier and confidence level includes: The voiceprint feature vectors are input in parallel into a pre-trained classification model group consisting of a support vector machine model and a deep residual network model; The wavelet packet energy entropy features in the voiceprint feature vector are classified using the support vector machine model, and the first classification probability is output. The Mel-Cepstral Coefficient features in the voiceprint feature vector are processed by the deep residual network model to output the second classification probability; The first classification probability and the second classification probability are weighted and fused to generate a fused classification probability. Based on the comparison result between the fusion classification probability and the preset confidence threshold, a void status identifier is generated; The fusion classification probability is used as the confidence level and associated with the empty state identifier to form the classification result.

[0010] By adopting the above technical solution, firstly, the heterogeneous combination of Support Vector Machine (SVM) and Deep Residual Network (DRN) compensates for the limitations of a single model; SVM ensures stability under small sample sizes, while ResNet enhances sensitivity under complex patterns. Secondly, the weighted fusion strategy optimizes classification confidence and reduces the risk of misjudgment caused by environmental interference. Finally, the correlation between confidence and state label provides actionable decision outputs, supporting closed-loop automation from data to label. This technical solution combines the advantages of traditional machine learning and deep learning, not only improving the accuracy of the detection system but also enhancing the interpretability of the results through probabilistic output, achieving high-precision and robust identification of road surface void states.

[0011] Optionally, when the void status indicator indicates a void status, the step of triggering a disease marking command based on the spatial coordinate data includes: Receive the classification results, which include the empty status identifier and the corresponding confidence level; When the empty status indicator shows an empty status, extract the spatial coordinate data synchronized with the classification results; Perform position optimization processing on the spatial coordinate data to generate marked position coordinates; A disease marking instruction is generated based on the marked location coordinates and associated with the current timestamp.

[0012] By adopting the above technical solution, the confidence threshold comparison mechanism introduces a risk-controlled decision-making threshold, avoiding mislabeling of low-reliability detection results; the extraction and optimization of spatial coordinate data ensures centimeter-level alignment between the marked position and the disease point, improving maintenance accuracy; and the association between instructions and timestamps enhances the traceability of operations, supporting subsequent data analysis and optimization. This technical solution combines machine learning output with high-precision positioning technology, not only achieving a high degree of automation in marking triggering but also overcoming errors caused by environmental dynamics through optimized processing, providing efficient and reliable technical support for preventive road maintenance.

[0013] Optionally, the steps of constructing a set of disease coordinates based on the structured data packet and converting the road surface acoustic signal into a frequency domain feature map in the cloud service platform include: Parse the disease marking instructions and associated spatial coordinate data in the structured data packet; Filter the spatial coordinate data associated with the disease markers to generate an initial set of disease coordinates; Perform spatial clustering analysis on the initial set of disease coordinates, and merge adjacent spatial coordinates to form a new set of disease coordinates; Extract the road surface acoustic wave signal from the acoustic signature-location data packet; The road surface acoustic wave signal is subjected to time-domain segmentation processing to generate a standardized time-domain signal frame sequence; Apply a windowed Fourier transform to each time-domain signal frame in the standardized time-domain signal frame sequence to generate the corresponding frequency-domain spectrum sequence; The frequency domain spectrum sequence is spliced ​​into a two-dimensional frequency domain feature map.

[0014] By adopting the above technical solution, the entire process of detecting data, from raw reception to feature visualization, has been optimized. Parsing and filtering structured data packets ensures the input of high-reliability disease data and reduces noise interference; spatial clustering analysis improves the spatial continuity of coordinate data, making disease distribution more consistent with actual physical laws; time-frequency transformation and feature map stitching generate intuitive acoustic representations, supporting in-depth analysis and visualization of diseases. This technical solution tightly integrates edge computing and cloud analytics, not only improving the efficiency and reliability of data processing but also enhancing the interpretability of disease identification through feature map generation, providing key technical support for intelligent decision-making in road maintenance.

[0015] Optionally, the step of fusing the disease coordinate set and frequency domain feature map to generate a spatialized disease distribution map includes: Spatial interpolation is performed on the set of disease coordinates to generate a continuous disease density distribution surface; Extract the spectral feature vector associated with the disease coordinates from the frequency domain feature map; Map the spectral feature vector to the spatial grid corresponding to the disease density distribution surface; Based on the correlation analysis between the spectral feature vectors and disease density data within the spatial grid, a fused feature matrix is ​​generated. A two-dimensional spatialized disease distribution map is rendered based on the fused feature matrix.

[0016] By adopting the above technical solution, an efficient conversion from multi-source data to a comprehensive visual atlas was achieved. First, spatial interpolation transforms discrete disease points into continuous density surfaces, enhancing the continuity of the spatial representation. Second, spectral feature extraction and correlation analysis enable adaptive fusion of acoustic and spatial data, improving the information quality of the atlas. Finally, two-dimensional rendering technology generates an intuitive and easy-to-understand distribution map, supporting precise decision-making in road maintenance. This technical solution closely integrates the spatiality and acoustic characteristics of the detection data, not only optimizing the macroscopic perspective of disease identification but also reducing the risk of false alarms through fusion analysis, providing reliable technical support for intelligent maintenance.

[0017] Optionally, after the step of generating the spatialized disease distribution map, the method further includes: Obtain a maintenance decision rule base, including the mapping relationship between disease severity classification standards and treatment measures; The confidence data and disease density data in the spatialized disease distribution map are weighted and fused together. Based on the weighted results, the treatment priority parameters in the maintenance decision rule base are matched to generate a maintenance decision map that includes the boundary of the treatment area and construction plan suggestions.

[0018] By adopting the above technical solutions, the scientific and automated nature of road maintenance decision-making has been achieved. The maintenance decision-making rule base, optimized based on damage tolerance theory and historical data, ensures the engineering rationality of damage level classification and treatment recommendations. The weighted fusion of confidence level and damage density enhances the comprehensiveness of risk assessment, overcoming the limitations of single indicators. The priority matching mechanism, through multi-dimensional matrices and efficient retrieval, achieves precise decision-making and dynamic adjustment. The spatial output of the maintenance decision map intuitively presents treatment plans, significantly improving the efficiency and accuracy of maintenance work. This technical solution transforms digital decisions into visualized engineering guidance. The interactive features of the map allow maintenance personnel to quickly locate key areas and obtain construction details, improving the executability of decisions and achieving a closed loop from data analysis to construction implementation.

[0019] Secondly, this application provides an intelligent acoustic void detection system for cement pavement, which adopts the following technical solution: A smart acoustic void detection system for cement pavement, the detection system comprising: The detection task configuration module is used to obtain user configuration instructions containing target area and driving speed parameters; The execution control module is used to control the unmanned driving device to travel along a preset path in response to the user configuration command, while applying periodic mechanical impacts to the road surface through the acoustic vibration wheel on the unmanned driving device; The multi-source data acquisition module is used to acquire the road surface sound wave signal generated by the mechanical impact in real time through the acoustic fingerprint acquisition module, and to acquire the spatial coordinate data and timestamp of the unmanned driving device through the positioning system. The voiceprint feature extraction module is used to perform noise suppression and frequency domain transformation processing on the road surface sound wave signal to generate a standardized voiceprint feature vector. The spatiotemporal data fusion module is used to perform spatiotemporal alignment processing on the road surface acoustic wave signal, voiceprint feature vector, spatial coordinate data and timestamp to generate a voiceprint-location data packet. The intelligent classification module is used to input the voiceprint feature vector into a pre-trained classification model group for real-time inference and generate a classification result that includes the empty state identifier and confidence level. The disease marking triggering module is used to trigger a disease marking command based on the spatial coordinate data when the void status indicator indicates a void status. The data encapsulation and transmission module is used to associate the voiceprint-location data packet and the classification result into a structured data packet. When there is a disease marking instruction, it is associated with the data packet and uploaded to the cloud service platform. The frequency domain feature map generation module is used to construct a set of disease coordinates based on the structured data packet in the cloud service platform, and convert the road surface acoustic wave signal into a frequency domain feature map; The spatial map generation module is used to fuse the disease coordinate set and frequency domain feature map to generate a spatial disease distribution map.

[0020] Thirdly, this application provides a computer device, which adopts the following technical solution: A computer device includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to perform the steps of the method as described in the first aspect.

[0021] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the first process of a method for detecting voids in cement pavement using intelligent acoustic wave technology, which is one embodiment of this application.

[0023] Figure 2 This is a schematic diagram of the second process of a smart acoustic delamination detection method for cement pavement according to one embodiment of this application.

[0024] Figure 3 This is a schematic diagram of the third process of the intelligent acoustic void detection method for cement pavement according to one embodiment of this application.

[0025] Figure 4 This is a schematic diagram of the fourth process of the intelligent acoustic void detection method for cement pavement according to one embodiment of this application.

[0026] Figure 5 This is a schematic diagram of the fifth process of the intelligent acoustic void detection method for cement pavement according to one embodiment of this application.

[0027] Figure 6 This is a schematic diagram of the sixth process of the intelligent acoustic void detection method for cement pavement according to one embodiment of this application.

[0028] Figure 7 This is a schematic diagram of the seventh process of the intelligent acoustic void detection method for cement pavement according to one embodiment of this application. Detailed Implementation

[0029] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-7 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0030] This application discloses an intelligent acoustic void detection method for cement pavement.

[0031] Reference Figure 1 A method for detecting voids in cement pavement using intelligent acoustic wave technology, specifically including: Step S101: Obtain user configuration instructions containing target area and driving speed parameters; The target area parameter defines the spatial range of the detection, typically based on Geographic Information System (GIS) coordinates or actual road segment descriptions, ensuring that detection activities are focused on specific road segments and avoiding resource waste. The driving speed parameter directly relates to the spatiotemporal resolution of data acquisition: excessive speed may lead to insufficient acoustic signal sampling, affecting feature extraction accuracy; excessive speed reduces detection efficiency. In some embodiments, the default detection speed is 0.4 m / s, a setting that balances data quality and operational efficiency.

[0032] From a system control perspective, the user configuration command can be input through a human-machine interface (such as a touch screen) to trigger the subsequent path planning module. Its essence is to transform the user's intention into a parameterized task that the machine can execute, laying the foundation for the autonomous navigation of unmanned driving equipment.

[0033] Step S102: In response to the user configuration command, control the unmanned driving equipment to drive along the preset path, and at the same time apply periodic mechanical impact to the road surface through the acoustic vibration wheel on the unmanned driving equipment; Among them, the unmanned driving equipment travels based on a preset path. Its path planning relies on laser SLAM+GNSS+vision fusion positioning technology to ensure that the vehicle moves accurately along a linear or curved trajectory, avoiding human error. The acoustic vibration wheel, as a mechanical impact source, uses the principle of generating stress waves through physical impact. When the wheel contacts the road surface, the impact energy excites a vibration response in the road structure. If there are defects such as voids in the roadbed, its vibration modes will differ from the normal area, thus showing differences in the acoustic signal.

[0034] In this embodiment, the intelligent pulsed acoustic vibration wheel weighs 30kg and is connected to the vehicle via a universal joint. This design ensures the stability and directionality of the impact. From a signal generation perspective, the periodic mechanical impact creates a controllable acoustic excitation environment, avoiding interference from random noise and providing a clean signal source for subsequent acoustic signature acquisition.

[0035] Step S103: The road surface sound wave signal generated by mechanical impact is acquired in real time through the soundprint acquisition module, and the spatial coordinate data and timestamp of the unmanned driving equipment are acquired through the positioning system. In this embodiment, the voiceprint acquisition module (e.g., using a directional microphone) converts sound pressure fluctuations generated by mechanical impact into electrical signals based on the piezoelectric effect or electromagnetic induction principle. Its operating frequency band of 20Hz-20kHz covers the main energy range of road vibration, and its high directivity design suppresses environmental noise. The positioning system (e.g., GNSS) calculates latitude and longitude coordinates by receiving satellite signals, achieving centimeter-level accuracy using RTK technology, and compensates for errors caused by signal obstruction using an inertial navigation unit. The timestamp, provided by the system clock with sub-millisecond accuracy, establishes a strict temporal relationship between the sound wave signal and spatial coordinates.

[0036] It should be noted that the core of this step is to solve the problem of synchronizing heterogeneous data: acoustic signals are time series, while spatial coordinates are geographic information. Only through high-precision timestamps can we ensure that each acoustic sample corresponds to a unique location in subsequent analysis.

[0037] Step S104: Perform noise suppression and frequency domain transformation on the road surface sound wave signal to generate a standardized soundprint feature vector; Noise suppression can employ adaptive filters (such as the LMS algorithm) to dynamically estimate the statistical characteristics of environmental noise (such as wind noise or vehicle noise) and subtract noise components from the original signal. Spectral subtraction can also be used to enhance the effective frequency band energy of the impact sound wave by analyzing the static characteristics of the signal spectrum. Frequency domain transformation converts the time-domain signal into a frequency-domain representation. This is necessary because cavitation defects often manifest acoustically as enhancement or attenuation of specific frequency components, which are difficult to directly identify in the time domain. During processing, the signal is first framed and windowed (e.g., using a Hamming window) to smooth the spectrum, then subjected to a Fast Fourier Transform (FFT) to obtain the amplitude spectrum, and finally mapped to a Mel scale to simulate human hearing characteristics, generating feature vectors such as Mel frequency cepstral coefficients (MFCC). Finally, the features are normalized to eliminate the influence of amplitude fluctuations and ensure the comparability of data from different test batches.

[0038] Step S105: Input the road surface acoustic wave signal, acoustic feature vector, spatial coordinate data and timestamp into the multi-source fusion module to generate a spatiotemporally aligned acoustic-location data packet; The core algorithm of the multi-source fusion module (such as Kalman filtering or simple time alignment) uses timestamps as key indices to bundle the original road surface acoustic wave signal, acoustic feature vector (representing acoustic properties), and spatial coordinate data (representing geographical location) into a unified data unit. The generated data packet adopts a standardized format (such as JSON or binary stream) and includes a timestamp primary key, an acoustic feature array, and latitude and longitude coordinates. The necessity of this fusion lies in the fact that de-emergence diagnosis is only meaningful when combined with acoustic anomalies and spatial location; isolated sound features cannot locate defects.

[0039] Step S106: Input the voiceprint feature vector into the pre-trained classification model group for real-time inference and generate a classification result containing the empty state identifier and confidence level. In this group, classification models typically employ ensemble learning strategies, such as a parallel combination of Support Vector Machines (SVM) and Deep Residual Networks (ResNet). The SVM model operates based on wavelet packet energy entropy features. Its principle is to decompose the signal frequency band through wavelet transform and calculate the energy entropy of each sub-band to capture the changes in spectral energy distribution caused by the disease. The ResNet model, on the other hand, directly processes Mel-frequency cepstral coefficients and uses deep neural networks to learn an abstract representation of frequency domain features.

[0040] Specifically, during real-time inference, the two models output probability results separately, and then the final classification result is generated through weighted fusion (e.g., setting weights based on validation set accuracy). The empty state label is a binary classification output (normal / empty), while the confidence score quantifies the reliability of the judgment (a probability value between 0 and 1). The technical advantage of this step lies in the parallel operation of the two engines: SVM is suitable for small sample scenarios, while ResNet excels at complex patterns, and their combination improves generalization ability.

[0041] Step S107: When the void status indicator indicates a void status, trigger the disease marking command based on the spatial coordinate data; The preset confidence threshold (e.g., 0.8) is an adjustable parameter that balances detection reliability with false positive rate: a threshold that is too high may miss minor diseases, while a threshold that is too low may easily produce false positives. When the confidence level of the classification result exceeds the preset confidence threshold, the system determines it to be in a "blank state" and then triggers a disease marking instruction.

[0042] In this embodiment, after generating the defect marking instruction, an actuator, such as a vehicle-mounted spraying device, is driven based on spatial coordinate data to mark the corresponding road surface location (e.g., spraying red paint). This step achieves closed-loop feedback, directly converting acoustic detection results into physical actions to ensure that no defect points are missed.

[0043] Step S108: Associate the voiceprint-location data packet and the classification result into a structured data packet. When there are disease markers, associate them together and upload them to the cloud service platform. The structured data packets employ an index-based association method, using timestamps as the primary key to bind voiceprint-location data packets (raw data), classification results (inference output), and potential disease markers into a complete log. This structure ensures data traceability: for example, it allows tracing back the voiceprint features and model confidence corresponding to a specific marker point. The upload process can be executed via a 4G / 5G wireless industrial router, employing compression and encryption protocols to guarantee transmission efficiency and security.

[0044] Step S109: Construct a set of disease coordinates based on structured data packets in the cloud service platform, and convert the road surface acoustic wave signal into a frequency domain feature map; The disease coordinate set is generated by extracting the spatial coordinates (latitude and longitude) of all data points marked as anomalies from the classification results, providing a foundation for spatial analysis. Simultaneously, the road surface acoustic wave signal is converted into a frequency domain feature map, such as a Mel-Spectrogram time-frequency spectrum. Its generation process includes short-time Fourier transform and Mel-scale mapping, ultimately presenting the time-frequency-energy relationship in the form of a heatmap. The horizontal axis of the feature map represents the timestamp (associated with the detection time sequence), the vertical axis represents the frequency components, and the color depth represents the energy intensity, making the frequency domain features related to the de-energization (such as energy anomalies in specific frequency bands) intuitively visible.

[0045] Step S110: Merge the disease coordinate set and frequency domain feature map to generate a spatialized disease distribution map.

[0046] Specifically, the logical principle of this step is to achieve spatial modeling of diseases based on Geographic Information Systems (GIS) and data visualization technology. The fusion process uses spatial interpolation algorithms (such as Kriging) to transform the discrete set of disease coordinates into a continuous spatial surface, while simultaneously overlaying thermal information from frequency domain feature maps to generate two-dimensional or three-dimensional distribution maps. In the maps, color depth is associated with confidence levels, with high-confidence areas highlighted with warm colors to visually display disease clusters. Such maps can reveal the distribution patterns of latent diseases, such as voids that are prone to occur at road joints or slope changes.

[0047] The above implementation presents an efficient and accurate method for detecting voids in cement pavement. First, by using an acoustic vibration wheel for excitation and synchronous data acquisition, physical impact is transformed into a quantifiable acoustic-spatial information flow, overcoming the subjectivity of traditional detection methods. Second, noise suppression and frequency domain transformation improve feature robustness, while a dual-model classifier ensures reliable inference. Finally, cloud-based data fusion and visualization generate an intuitive disease distribution map, upgrading from point-based detection to area-based management. This technical solution combines edge computing, machine learning, and geographic information technology, not only improving the automation level and centimeter-level positioning accuracy of detection but also optimizing the scientific nature of maintenance decisions through data closed-loop management, providing key technical support for preventative maintenance of road infrastructure.

[0048] Reference Figure 2 As one implementation of step S104, the step of performing noise suppression and frequency domain transformation processing on the road surface sound wave signal to generate a standardized soundprint feature vector includes: Step S201: Obtain the road surface sound wave signal collected in real time by the soundprint acquisition module; The road surface acoustic wave signal originates from the periodic mechanical impact of the intelligent pulsed acoustic vibration wheel on the road surface. When the wheel strikes the road surface, stress waves propagate within the road structure. If defects such as voids exist in the base layer, its vibration mode will differ from the normal area, resulting in characteristic changes in the acoustic response. The acoustic signature acquisition module, based on the piezoelectric effect or capacitive principle, converts sound pressure fluctuations into analog electrical signals. Its operating frequency band of 20Hz-20kHz covers the main energy range of road surface vibration. The highly directive design effectively suppresses environmental noise (such as wind noise or traffic interference), ensuring the purity of the acquired signal. During the acquisition process, the module performs analog-to-digital conversion at a high sampling rate (e.g., 44.1kHz) to generate a time-domain discrete signal sequence.

[0049] Step S202: Apply noise suppression processing to the acoustic signal and use spectral subtraction to suppress impulse interference noise to generate a denoised signal; The noise suppression process first employs an adaptive filter (such as the LMS algorithm), which dynamically estimates the power spectral density of the ambient background noise and adjusts the filter coefficients through feedback to subtract noise components from the original signal. Secondly, spectral subtraction is applied to suppress impulse interference noise: by analyzing the amplitude characteristics of the signal's short-time spectrum, sudden noise bands (such as transient interference from passing vehicles) are identified and attenuated. The combination of these two methods forms a layered noise reduction strategy that addresses both continuous background noise and optimizes for transient impulses.

[0050] Step S203: Divide the denoised signal into equal-length time segments and apply a windowing function frame by frame; The framing operation cuts the continuous denoised wave signal into equal-length time segments (e.g., 23ms per frame with 50% overlap). The basis for this is that the road impact response can be regarded as a stationary process in the short term, while the long-term signal may be non-stationary due to changes in vehicle speed. Processing each frame signal independently can improve the temporal resolution.

[0051] In the embodiments of this application, the application of windowing functions (such as Hamming windows) aims to reduce spectral leakage: the window function smoothly decays to zero at the frame edges, avoiding frequency domain distortion caused by truncation effects. For example, the Hamming window weakens the discontinuity of frame boundaries through cosine weighting, making the Fourier transform results more focused on the frequency band characteristics of the current frame.

[0052] Step S204: Perform frequency domain transformation on the windowed time segment of equal length to generate a frequency domain energy spectrum; The frequency domain transformation is achieved through Fast Fourier Transform (FFT), which mathematically transforms time-domain convolution operations into frequency-domain multiplication, significantly improving computational efficiency. After performing an FFT on each windowed frame, a linear spectrum is obtained, with frequency on the horizontal axis and amplitude on the vertical axis, reflecting the energy intensity of the signal in each frequency band. Subsequently, the linear spectrum is mapped to a Mel-scale filter bank: the Mel scale is designed based on the characteristics of human hearing, offering high resolution in the low-frequency range and low resolution in the high-frequency range, better aligning with acoustic perception. The filter bank output is the Mel spectral coefficient, which compresses high-frequency redundant information and highlights key frequency bands related to the defect (such as the 1-5kHz abnormal resonance region).

[0053] It should be noted that void defects often manifest as an increase or decrease in energy at a specific frequency, which is difficult to identify directly in the time domain waveform, but the frequency domain energy spectrum can intuitively expose these characteristics.

[0054] Step S205: Extract multi-scale acoustic signature features from the frequency domain energy spectrum to generate the original feature vector; The multi-scale acoustic signature features include Mel-frequency cepstral coefficients (MFCC) and wavelet packet energy entropy, which are complementary: MFCC generates cepstral coefficients by decorrelation of the Mel-frequency spectral coefficients through discrete cosine transform, which has the advantage of characterizing the vocal tract characteristics of the signal (similar to speech recognition) and can capture spectral envelope changes caused by gaps; wavelet packet energy entropy decomposes the signal into multiple resolutions based on wavelet transform (such as the Daubechies wavelet basis) and calculates the information entropy value of each sub-band energy, which is sensitive to transient impact components (such as the mechanical impact response in this embodiment). During feature extraction, MFCC focuses on the global spectral shape, while energy entropy focuses on the local energy distribution, and the two are concatenated to form the original feature vector. By fusing acoustic representations at different scales, the limitation of a single feature in representing the response of complex road surfaces is overcome.

[0055] Step S206: Normalize the original feature vector to generate a standardized voiceprint feature vector.

[0056] The normalization process includes zero-mean normalization and standard deviation scaling: zero-mean normalization adjusts the mean of each feature dimension to zero, eliminating DC bias or systematic bias; standard deviation scaling unifies the variance of each dimension to 1, preventing certain large numerical features (such as energy entropy) from dominating model training. Mathematically, it is a linear transformation that ensures the comparability of features from different detection batches and road sections. For example, without normalization, the MFCC coefficients may fluctuate due to differences in microphone gain, misleading the classification model. After normalization, the feature vectors fall within a uniform numerical range, accelerating model convergence and improving robustness.

[0057] In the above embodiments, efficient feature extraction of road surface acoustic signals is achieved. Noise suppression and frame-by-frame windowing improve the time-frequency analysis quality of the signal, laying the foundation for feature extraction. Multi-scale feature fusion (such as MFCC and energy entropy) captures the acoustic fingerprint of delamination defects from both global and local dimensions, enhancing pattern discriminability. Normalization ensures the standardization and comparability of feature vectors, providing stable input for model inference. This technical solution closely integrates digital signal processing with machine learning feature engineering, not only overcoming noise interference in complex environments but also optimizing the accuracy and robustness of defect detection through multi-scale representation.

[0058] Reference Figure 3 As one implementation of step S106, the step of inputting the voiceprint feature vector into a pre-trained classification model group for real-time inference and generating a classification result containing the empty state identifier and confidence level includes: Step S301: Input the voiceprint feature vectors in parallel into a pre-trained classification model group consisting of a support vector machine model and a deep residual network model; The parallel input architecture feeds voiceprint feature vectors into two pre-trained models simultaneously: Support Vector Machine (SVM) is a traditional model based on statistical learning theory, which is good at handling high-dimensional features and small sample scenarios; Deep Residual Network (ResNet) is a deep learning model that solves the gradient vanishing problem through residual connections and is suitable for learning complex nonlinear patterns.

[0059] In this embodiment, pre-training means that the model has already optimized its parameters using historical declassified data and can be directly deployed for real-time inference. The principle behind this parallel design lies in the complementarity of the models—SVM, based on the principle of minimizing structural risk, has strong generalization ability but limited ability to capture complex patterns; ResNet automatically learns abstract features through deep neural networks but requires a large amount of data. This step reduces the risk of misjudgment by a single model through dual-engine redundancy, laying the foundation for subsequent probabilistic fusion.

[0060] Step S302: Classify the wavelet packet energy entropy features in the voiceprint feature vector using a support vector machine model, and output the first classification probability. Among them, the wavelet packet energy entropy feature in the voiceprint feature vector is calculated based on the wavelet packet decomposition technique: the wavelet packet decomposes the signal into multiple sub-bands, and the energy entropy value quantifies the uniformity of energy distribution in each sub-band. Voiding disease often causes energy to concentrate or disperse in a specific frequency band, resulting in abnormal entropy values.

[0061] Specifically, the SVM model uses a radial basis function (RBF) kernel to map features to a high-dimensional space, in which the optimal classification hyperplane is solved to maximize the margin between the two classes of samples (empty and normal). The first classification probability is derived from the hyperplane distance using Platt scaling or a similar method, representing the probability estimate that the sample belongs to the empty class; the closer the value is to 1, the higher the probability. The advantage of this step is that SVM is robust to small sample data, especially when the training data is limited, providing a stable probability output.

[0062] Step S303: Process the Mel-Cepstral Coefficient features in the voiceprint feature vector using a deep residual network model and output the second classification probability; Among them, the deep residual network (ResNet) model focuses on processing Mel-frequency cepstral coefficients (MFCC) features, which simulate the characteristics of human hearing. It extracts spectral envelope information through cepstral analysis and effectively captures the resonant frequency changes caused by voiding.

[0063] In this embodiment, the ResNet model consists of multiple stacked residual blocks, each containing a convolutional layer, a batch normalization layer, and an activation function. Residual connections allow for direct backpropagation of gradients, mitigating the degradation problem of deep networks. The model automatically learns temporal dependencies and local patterns in MFCC features, ultimately outputting a second classification probability through fully connected layers and a softmax function. This probability reflects the pattern confidence gained by the network based on a large amount of data training, offering advantages in adaptability to complex acoustic environments (such as noise or changes in road surface material).

[0064] Step S304: Weighted fusion of the first classification probability and the second classification probability to generate a fused classification probability; Specifically, weighted fusion will convert the first classification probability P output by the SVM into a weighted fusion. svm and the second classification probability P output by ResNet resnet A linear combination is performed, and the weight coefficients α and β can be dynamically set based on the model's performance on the validation set (e.g., α=0.4, β=0.6), satisfying α+β=1 to maintain the consistency of the probability space.

[0065] In this embodiment of the application, the formula for calculating the fusion classification probability is P. f = α·P svm + β·P resnet The principle behind this fusion lies in probability calibration: SVM may be conservative on some boundary samples, while ResNet is more sensitive to complex patterns. Weighted averaging can smooth out the bias of individual models and improve the overall confidence. Ultimately, the fused classification probability still falls within the [0,1] interval.

[0066] Step S305: Generate a blanking status identifier based on the comparison result between the fusion classification probability and the preset confidence threshold; The preset confidence threshold (e.g., 0.8) is an adjustable parameter, set based on a trade-off between false positives and false negatives according to business needs: a higher threshold makes the system more cautious, marking only high-confidence samples as "empty," reducing false positives but potentially causing false negatives; a lower threshold makes the system more sensitive, covering more potential defects but increasing false positives. During comparison, if the fused classification probability is greater than or equal to the confidence threshold, a "empty" status is generated; otherwise, it is marked as "normal."

[0067] Step S306: The fusion classification probability is used as the confidence level and associated with the empty state identifier to form the classification result.

[0068] The fusion classification probability is assigned a confidence level, quantifying the reliability of the classification result. For example, a probability of 0.95 indicates high confidence, while 0.7 indicates uncertainty. The confidence level is then associated with a null status identifier (such as a binary tag) to form a complete classification result, which can be stored as a data log or transmitted to a tagging device in real time. Through this association, users not only know the classification conclusion but can also assess the risk level using the confidence level, facilitating subsequent prioritization or manual review.

[0069] In the above implementation, firstly, the heterogeneous combination of Support Vector Machine (SVM) and Deep Residual Network (DRN) compensates for the limitations of a single model; SVM ensures stability under small sample sizes, while ResNet enhances sensitivity under complex patterns. Secondly, the weighted fusion strategy optimizes classification confidence and reduces the risk of misjudgment caused by environmental interference. Finally, the correlation between confidence and state label provides actionable decision outputs, supporting closed-loop automation from data to label. This technical solution combines the advantages of traditional machine learning and deep learning, not only improving the accuracy of the detection system but also enhancing the interpretability of the results through probabilistic output, achieving high-precision and robust identification of road surface void states.

[0070] Reference Figure 4 As one implementation of step S107, when the void status indicator indicates a void status, the step of triggering a disease marking command based on spatial coordinate data includes: Step S401: Receive the classification result, which includes the empty state identifier and the corresponding confidence level; Among them, the void status identifier is a binary label (such as "normal" or "void"), which represents the machine learning model's judgment on the disease at the current detection point; the confidence level is a continuous probability value (range 0-1), which quantifies the reliability of the model's judgment. A high confidence level (such as close to 1) indicates that the model is very certain about the void status judgment, while a low confidence level (such as 0.6) indicates that there is uncertainty.

[0071] Step S402: When the empty state indicator indicates an empty state, extract the spatial coordinate data synchronized with the classification results; The spatial coordinate data originates from GNSS satellite positioning equipment, employing centimeter-level RTK technology to record the latitude and longitude coordinates of the unmanned vehicle in real time, and is associated with the voiceprint acquisition timestamp through a sub-millisecond synchronization mechanism. Extraction is achieved through time indexing: when the "empty state" indicator shows an empty state, the system retrieves spatial coordinate data from the cache at the same time based on the timestamp of the classification result. For example, the system might detect a point as empty (confidence level 0.9), and then extract its corresponding coordinates (e.g., longitude 108.95°E, latitude 34.27°N) to provide a spatial anchor point for the marker.

[0072] Step S403: Perform position optimization processing on the spatial coordinate data to generate marked position coordinates; The location optimization process is based on statistical principles. A typical method involves acquiring a continuously collected spatial coordinate dataset within a preset time window (such as 10 coordinate points within the last second) and calculating the geometric center of these points as the final marked location. Mathematically, this is centroid calculation, which reduces random errors caused by vehicle vibrations and signal fluctuations through averaging. For example, if consecutive coordinate points drift by ±5 cm due to road bumps, geometric center calculation can converge them to near their true location.

[0073] It should be noted that the necessity of optimization stems from the dynamic nature of the actual detection environment: the unmanned electric four-wheeled vehicle may experience slight trajectory deviations due to slopes or obstacles while driving, and directly using single-point coordinates may lead to misalignment of the markers.

[0074] Step S404: Generate a disease marking instruction based on the marking location coordinates and associate it with the current timestamp.

[0075] The disease marking instruction is a structured command containing 3D geographic information (latitude, longitude, and elevation) of the marked location coordinates, a marking action type code (such as the instruction code for spraying red paint), and a trigger timestamp. The generation process is based on protocol encapsulation: the system encodes the marked location coordinates into a standard format (such as the WGS84 coordinate system) and combines this with preset action codes to generate an instruction frame. The associated timestamp provides audit trails, facilitating subsequent analysis of the temporal relationship between marking actions and detection events.

[0076] In this embodiment, the disease marking command is accurately parsed and executed by the disease marking device (such as a vehicle-mounted spraying device), and the timestamp ensures that the cloud platform can synchronously record the operation log. When an anomaly is detected and the threshold is reached, an alarm is triggered and the physical marking device is driven to make a disease mark at the designated location, which reflects the seamless connection between detection and maintenance.

[0077] In the above implementation, the confidence threshold comparison mechanism introduces a risk-controlled decision-making threshold, avoiding mislabeling of low-reliability detection results; the extraction and optimization of spatial coordinate data ensures centimeter-level alignment between the marked position and the disease point, improving maintenance accuracy; and the association between instructions and timestamps enhances the traceability of operations, supporting subsequent data analysis and optimization. This technical solution combines machine learning output with high-precision positioning technology, not only achieving a high degree of automation in marker triggering but also overcoming errors caused by environmental dynamics through optimized processing, providing efficient and reliable technical support for preventive road maintenance.

[0078] In addition, to improve the robustness of detection and judgment, an evaluation unit mechanism based on spatial distance can be introduced. Step S107 also includes: Based on the continuous driving path of the autonomous driving equipment, the acoustic signals and corresponding classification results within each preset spatial distance range are divided into an independent confidence evaluation unit; wherein, the preset spatial distance is determined based on the preset driving speed and sampling rate to ensure that each evaluation unit contains a preset number of acoustic samples. Sliding scan judgment is performed on a continuous confidence evaluation unit sequence; If the fusion classification probability of two consecutive adjacent confidence evaluation units exceeds the preset confidence threshold, the current detection area is determined to be a deterministic void disease, and a high-priority disease labeling instruction is generated. If the fusion classification probability of a single, non-contiguous adjacent confidence evaluation unit exceeds the preset confidence threshold, the current detection area is determined to be in a suspected empty state, and a low-priority record or prompt instruction is generated without triggering the entity marking instruction.

[0079] Specifically, the system can treat the acoustic signals and their corresponding classification results within a 0.6-meter range along a continuous vehicle travel path as an independent confidence evaluation unit. The spatial scale of this evaluation unit is determined based on a preset travel speed (e.g., default 0.4 m / s) and sampling rate, ensuring that each unit contains a sufficient number of acoustic samples for reliable evaluation. When the system performs a judgment, it will slide and scan the continuous sequence of evaluation units. If the combined confidence (i.e., fusion classification probability) of two consecutive adjacent evaluation units both exceed a preset confidence threshold (e.g., 0.8), the current detection area is determined to be a definitive "gap" defect, and a high-priority defect marking instruction is immediately triggered. If only a single isolated evaluation unit has a confidence exceeding the threshold, it is determined to be in a "suspected gap" state, and the system can take a lower-priority recording or prompt action instead of immediately triggering entity marking.

[0080] Understandably, this dual verification logic effectively reduces the risk of misjudgment caused by transient environmental interference or single-point measurement errors. While ensuring that high-confidence defects are not missed, it distinguishes and controls the false alarm rate of suspected signals, achieving a more refined balance between detection reliability and false alarm rate. This mechanism is closely integrated with the original spatial coordinate data extraction, optimization, and marking instruction generation process, making the final defect marking more prudent and reliable.

[0081] Reference Figure 5 As one implementation of step S109, the steps of constructing a set of disease coordinates based on structured data packets in the cloud service platform and converting the road surface acoustic wave signal into a frequency domain feature map include: Step S501: Parse the disease marking instructions and associated spatial coordinate data in the structured data packet; The parsing operation is based on a predefined data format (such as JSON or binary encoding). The disease marking instruction usually includes the marking location coordinates, action type code and timestamp. The spatial coordinate data comes from the centimeter-level RTK measurement results of the GNSS positioning system and adopts a latitude and longitude coordinate system (such as WGS84).

[0082] Specifically, the parsing algorithm identifies index fields (such as timestamps) in the data packet, establishes a mapping relationship between instructions and coordinates, and ensures that each marked instruction corresponds to an accurate geographical location. Its technical principle is similar to database query optimization, quickly retrieving related entries through key-value matching. For example, when a marked instruction is parsed, the system extracts the spatial coordinates of the same moment from the coordinate dataset based on its timestamp.

[0083] Step S502: Filter the spatial coordinate data associated with disease markers to generate an initial set of disease coordinates; The filtering process iterates through all spatial coordinate data, retaining only coordinate points marked with disease indicators while discarding coordinates in a normal state (below a preset confidence threshold), ensuring that each point in the set represents a high probability of disease. For example, if the preset confidence threshold is 0.8, a point with a confidence level of 0.6 (below the threshold) is filtered out, while points with a confidence level of 0.9 are retained. When generating the initial set of disease coordinates, the system also associates corresponding timestamps for subsequent time-series analysis.

[0084] Step S503: Perform spatial clustering analysis on the initial disease coordinate set and merge adjacent spatial coordinates to form a disease coordinate set; Specifically, spatial clustering analysis employs distance-based algorithms (such as DBSCAN or custom neighborhood merging). The principle is to set a distance threshold ε (e.g., 1 meter), construct a neighborhood with each coordinate point as its center, merge all overlapping neighborhoods into a single defect area, and calculate the geometric center as the representative point. This process solves the problem of discrete point redundancy: during detection, due to vehicle vibration or positioning drift, the same defect may be repeatedly marked as multiple adjacent points; after clustering, these are merged into a single region. For example, if the distance between three points is less than ε, they are merged into one defect point, with its coordinates taken as the centroid of the three points. This step, through the clustering algorithm, improves the analysis from point-like noise to area-like regularity, enhancing the continuity of spatial analysis.

[0085] Step S504: Extract the road surface acoustic wave signal from the voiceprint-location data packet; The acoustic signal field consists of the original time-domain waveform recorded by the voiceprint acquisition device, typically with a sampling rate of 44.1 kHz and a quantization depth of 16 bits, preserving complete acoustic details. During extraction, the system decodes the acoustic signal from the data packet based on the timestamp index and verifies its synchronization with the spatial coordinates.

[0086] Step S505: Perform time-domain segmentation processing on the road surface acoustic wave signal to generate a standardized time-domain signal frame sequence; In this embodiment, time-domain segmentation employs a sliding window technique: a fixed time window length T (e.g., 23 milliseconds) and step size S (e.g., 11.5 milliseconds, 50% overlap) are set to divide the continuous acoustic signal into a sequence of equal-length frames. A Hamming window function is then applied to each frame to reduce spectral leakage—the window function smoothly decays to zero at the frame edges, avoiding frequency distortion caused by truncation. Normalization processing includes amplitude normalization to eliminate the influence of gain differences in the acquisition devices. For example, a 1-second acoustic signal may be divided into approximately 87 frames (50% overlap), each processed independently. This step is necessary because road surface acoustic waves are non-stationary signals; after segmentation, each frame can approximate a stationary process, making the Fourier transform more efficient.

[0087] Step S506: Apply a windowed Fourier transform to each time-domain signal frame in the standardized time-domain signal frame sequence to generate the corresponding frequency-domain spectrum sequence. The windowed Fourier transform first applies a window function (such as a Hamming window) to weight each frame of the time-domain signal to reduce boundary effects. Then, a Fast Fourier Transform (FFT) is performed to map the signal from the time domain to the frequency domain, generating a linear spectrum with frequency on the horizontal axis and amplitude on the vertical axis. The spectrum reflects the energy intensity of the signal in each frequency band; signal gaps often manifest as enhanced or attenuated resonant peaks at specific frequencies (e.g., 1-5kHz). For example, a 23ms frame of signal, after FFT, yields a spectrum of 0-22kHz.

[0088] Step S507: The frequency domain spectrum sequence is spliced ​​into a two-dimensional frequency domain feature map.

[0089] The splicing operation arranges the frequency spectrum of each signal frame in a time sequence, forming a two-dimensional matrix: the horizontal axis represents time (correlated detection timing), the vertical axis represents frequency components, and the matrix element values ​​represent energy intensity (such as logarithmic amplitude). Subsequently, the matrix may be converted to a Mel-scale to simulate auditory characteristics and normalized to enhance contrast. The generated two-dimensional frequency domain feature map is essentially a time-spectrum map of the sound wave, such as a Mel-Spectrogram, whose color depth visualizes energy distribution. This step elevates the discrete spectrum to a continuous feature map, facilitating machine learning models or manual analysis of the spatiotemporal evolution of diseases.

[0090] The above implementation optimizes the entire process of detection data processing, from raw reception to feature visualization. By parsing and filtering structured data packets, high-reliability disease data input is ensured, reducing noise interference. Spatial clustering analysis improves the spatial continuity of coordinate data, making disease distribution more consistent with actual physical laws. Time-frequency transformation and feature map stitching generate intuitive acoustic representations, supporting in-depth disease analysis and visualization. This technical solution tightly integrates edge computing and cloud analytics, not only improving data processing efficiency and reliability but also enhancing the interpretability of disease identification through feature map generation, providing key technical support for intelligent decision-making in road maintenance.

[0091] Reference Figure 6 As one implementation of step S110, the step of fusing the disease coordinate set and frequency domain feature map to generate a spatialized disease distribution map includes: Step S601: Perform spatial interpolation on the disease coordinate set to generate a continuous disease density distribution surface; Specifically, spatial interpolation can employ mathematical interpolation methods (such as radial basis function interpolation). The principle is based on the spatial density values ​​of each point in the disease coordinate set (e.g., the number of disease points per unit area). A regular grid coordinate system is constructed in the target area, and then the density value of each grid node is calculated using radial basis functions (such as Gaussian functions or multiple quadratic functions), forming a smooth, continuous surface. During interpolation, grid nodes closer to the disease points are assigned higher density values, and vice versa, simulating the diffusion effect of disease influence.

[0092] It is understandable that actual road defects often exhibit regional characteristics, and discrete point coordinates cannot fully express the degree of defect clustering and spatial trends. This embodiment generates continuous surfaces, which can more intuitively identify high-risk areas and support macro-level decision-making.

[0093] Step S602: Extract the spectral feature vector associated with the disease coordinates from the frequency domain feature map; Among them, the frequency domain feature map contains rich acoustic information (such as specific frequency resonance caused by voiding), but it needs to be combined with the disease coordinates to give it spatial meaning; after extracting the spectral feature vector, each disease point not only has location attributes, but also carries an acoustic fingerprint, which enhances the information density of the map.

[0094] Specifically, the extraction process first identifies the timestamp of each point in the disease coordinate set, and then locates the spectral data of the same timestamp interval in the frequency domain feature map (e.g., by extracting the corresponding frame through time axis indexing). This process ensures a strict correspondence between acoustic features and spatial location. The spectral feature vector is usually composed of multi-dimensional statistical features, such as calculating the mean, variance, and spectral entropy of the extracted spectrum, forming a fixed-length vector representation.

[0095] Step S603: Map the spectral feature vector to the spatial grid corresponding to the disease density distribution surface; The mapping operation is based on a regular grid generated from the disease density distribution surface, where each grid node has unique spatial coordinates (such as latitude and longitude). The system assigns the spectral feature vector to the nearest grid node according to its spatial coordinates corresponding to its timestamp, so that each grid node contains both the disease density value and the spectral feature vector. In some embodiments, the mapping algorithm may use nearest neighbor interpolation or bilinear interpolation to ensure the spatial continuity of the feature vector.

[0096] The necessity of this step lies in the fact that it solves the problem of scale inconsistency between acoustic features and spatial grids: the disease density surface is a continuous spatial representation, while the spectral feature vector is point data. After mapping, the two can be operated on point by point under a unified grid.

[0097] Step S604: Based on the correlation analysis between the spectral feature vectors and disease density data within the spatial grid, a fused feature matrix is ​​generated; In this embodiment, the correlation analysis can be performed using the Pearson correlation coefficient, which essentially measures the linear relationship between the spectral feature vector (such as multi-dimensional energy statistics) and the disease density value within a grid. The result is normalized to the [0,1] interval as the fusion weight. For example, if the spectral features of a certain grid show high-frequency energy anomalies and are positively correlated with high disease density, then a higher weight is assigned. When generating the fusion feature matrix, the system combines the weights with the original feature vectors to form an enhanced feature representation for each grid.

[0098] It is understandable that simply overlaying spatial and acoustic data may lead to information redundancy or noise amplification, while correlation analysis can screen out acoustic patterns closely related to the disease and improve the signal-to-noise ratio of the spectrum.

[0099] Step S605: Render a two-dimensional spatialized disease distribution map based on the fused feature matrix.

[0100] The rendering process can be based on a fused feature matrix, using a heatmap to represent the density distribution of diseases, with color levels (such as red-yellow-green) indicating high or low density. At the same time, the principal component analysis (PCA) dimensionality reduction results of the spectral feature vector are superimposed as a texture layer, displaying acoustic feature differences through pattern or color changes. Finally, the coordinates of core disease points with confidence levels exceeding a threshold are marked to provide accurate location references.

[0101] In this embodiment, the rendering algorithm can use graphics libraries (such as OpenGL or WebGL) to generate raster or vector graphics, ensuring no distortion during scaling. This step realizes the conversion from digital data to visual information. The two-dimensional spatial form of the map facilitates maintenance personnel to quickly identify disease clusters and optimize resource allocation. Finally, multi-layer rendering balances macroscopic trends and microscopic details, making the map both holistic and accurate.

[0102] The above implementation achieves efficient conversion from multi-source data to a comprehensive visual atlas. First, spatial interpolation transforms discrete disease points into continuous density surfaces, enhancing the continuity of the spatial representation. Second, spectral feature extraction and correlation analysis enable adaptive fusion of acoustic and spatial data, improving the information quality of the atlas. Finally, two-dimensional rendering technology generates an intuitive and easy-to-understand distribution map, supporting precise decision-making in road maintenance. This technical solution closely integrates the spatiality and acoustic characteristics of the detection data, not only optimizing the macroscopic perspective of disease identification but also reducing the risk of false alarms through fusion analysis, providing reliable technical support for intelligent maintenance.

[0103] Reference Figure 7 As a further implementation of the intelligent acoustic void detection method for cement pavement, after step S110 of generating a spatialized defect distribution map, the method further includes: Step S701: Obtain the maintenance decision rule base, including the mapping relationship between disease level classification standards and treatment measures; In this embodiment of the application, the maintenance decision rule base is essentially a pre-built expert system database, which can be constructed based on the Damage Tolerance Theory, which originates from the mechanics of materials and is used to evaluate the remaining life and safety of a structure in the presence of initial defects or damage.

[0104] Specifically, the disease classification standard uses quantitative indicators to divide void diseases into multiple risk ranges based on spatial dimensions (such as void area and depth) and temporal dimensions (such as disease development rate). For example, when a continuous area of ​​void is detected to be greater than 1 square meter and the depth is greater than 5 centimeters, the system automatically identifies it as "high risk". This classification ensures the objectivity and repeatability of disease assessment.

[0105] The mapping relationship between treatment measures is constructed using a decision tree algorithm. A decision tree is a machine learning model based on conditional branching that learns the correspondence between "disease level - construction plan" through a historical maintenance case database. For example, a high-risk level is mapped to the "grooving and grouting + base replacement" combined construction method, while medium and low levels may correspond to "local grouting" or "monitoring and observation". The dynamic optimization mechanism of the rule base allows it to be iteratively updated by continuously integrating new detection data and maintenance effect feedback, thereby improving the adaptability of decision-making.

[0106] Step S702: Weighted fusion of confidence data and disease density data in the spatialized disease distribution map; Among them, the confidence level data in the disease distribution map is a probability value between 0 and 1, reflecting the credibility of the classification results; the disease density data is obtained through spatial statistical calculation, representing the number of outliers per unit area, and quantifying the clustering intensity of the disease.

[0107] In this embodiment, the weighted fusion process first performs data normalization, using the Min-Max standardization method to scale the confidence level and disease density to a uniform numerical range (e.g., 0-1), eliminating the influence of dimensional differences on the fusion results. Weight allocation applies information entropy theory; the system dynamically calculates weight coefficients based on historical data. For example, a high confidence level (e.g., greater than 0.9) indicates reliable model judgment, and a higher weight is assigned (e.g., 0.7); while a high disease density (e.g., exceeding 5 points / square meter) reflects concentrated spatial risk, thus increasing the density weight (e.g., 0.3).

[0108] Subsequently, the fusion calculation generates a comprehensive risk index using a weighted summation formula: Comprehensive Risk Index = α × Confidence Level + β × Disease Density, where α and β are preset weights, and α + β = 1. This comprehensive risk index integrates the accuracy of the model's judgment with the actual distribution characteristics of the disease, avoiding misjudgments that may be caused by a single indicator. For example, isolated points with high confidence but low density will not be overemphasized.

[0109] Step S703: Based on the weighted results, match the treatment priority parameters in the maintenance decision rule base to generate a maintenance decision map that includes the boundary of the treatment area and construction plan suggestions.

[0110] Specifically, the matching process is based on a pre-set three-dimensional decision matrix. The dimensions of this matrix include factors such as the comprehensive risk index range, traffic load level (e.g., heavy-load lanes or ordinary road sections), and material service life. For example, when the risk index is greater than 0.8 and the area is in a high traffic load zone, the system will automatically match the highest priority parameter P1.

[0111] In this embodiment, the matching engine employs a hash table retrieval mechanism. A hash table is an efficient data structure that uses a hash function to directly map key-value pairs (such as risk indices) to storage locations, enabling fast queries. In this application, the system uses the comprehensive risk index as the key-value pair and indexes the corresponding urgency parameters (such as P1-P5 levels) in the rule base. Simultaneously, the engine introduces a timeliness correction factor to automatically increase the priority level of developing diseases (such as those with a detected weekly disease density growth rate exceeding 10%), ensuring dynamic adaptability in decision-making.

[0112] Understandably, transforming abstract risk quantification results into actionable priority parameters takes into account the multi-factor influence in actual engineering and avoids simplistic "one-size-fits-all" threshold decisions. Efficient hash table retrieval supports real-time decision-making needs, while the multi-dimensional matrix design enhances the system's responsiveness to complex scenarios.

[0113] Next, the process of generating the boundary of the treatment area first applies a boundary generation algorithm, such as Delaunay triangulation. This algorithm is a computational geometry method that connects discrete high-risk point clouds into a triangular mesh, ensuring that the circumcircle of any triangle in the mesh does not contain other points, thus forming a smooth topological structure. Subsequently, the α-shape algorithm is used to remove outliers, generating accurate closed polygonal boundaries for the treatment area, avoiding the arbitrariness of manual delineation. Construction plan suggestions are embedded through GIS layer association technology. The system calls the corresponding construction method database from the rule base according to priority parameters. For example, the boundary of the high-risk area is automatically associated with the "high-pressure grouting" process parameter library, including equipment configuration and material usage.

[0114] Ultimately, the output maintenance decision map is a spatial decision map (SpaceDM), whose layer structure includes a base layer (such as a road BIM model), a risk layer (rendering a comprehensive risk index using a heat map), and a decision layer (displaying treatment boundaries and construction method labels).

[0115] The above implementation achieves scientific and automated road maintenance decision-making. The maintenance decision-making rule base, optimized based on damage tolerance theory and historical data, ensures the engineering rationality of damage level classification and treatment recommendations. The weighted fusion of confidence level and damage density enhances the comprehensiveness of risk assessment, overcoming the limitations of single indicators. The priority matching mechanism, through multi-dimensional matrices and efficient retrieval, achieves precise decision-making and dynamic adjustment. The spatial output of the maintenance decision map intuitively presents treatment plans, significantly improving the efficiency and accuracy of maintenance work. This technical solution transforms digital decisions into visualized engineering guidance. The interactive nature of the map allows maintenance personnel to quickly locate key areas and obtain construction details, improving the executability of decisions and achieving a closed loop from data analysis to construction implementation.

[0116] This application also discloses an intelligent acoustic void detection system for cement pavement.

[0117] A smart acoustic void detection system for cement pavement, the system comprising: The detection task configuration module is used to obtain user configuration instructions containing target area and driving speed parameters; The execution control module is used to control the unmanned driving equipment to travel along a preset path in response to user configuration commands, while applying periodic mechanical impacts to the road surface through the acoustic vibration wheel on the unmanned driving equipment. The multi-source data acquisition module is used to acquire road surface sound wave signals generated by mechanical impact in real time through the acoustic fingerprint acquisition module, and to acquire spatial coordinate data and timestamps of the unmanned driving equipment through the positioning system. The voiceprint feature extraction module is used to perform noise suppression and frequency domain transformation on the road surface sound wave signal to generate a standardized voiceprint feature vector. The spatiotemporal data fusion module is used to perform spatiotemporal alignment processing on road surface sound wave signals, voiceprint feature vectors, spatial coordinate data and timestamps to generate voiceprint-location data packets; The intelligent classification module is used to input the voiceprint feature vector into a pre-trained classification model group for real-time inference and generate classification results that include the empty state identifier and confidence level. The disease marking trigger module is used to trigger a disease marking command based on spatial coordinate data when the void status indicator indicates a void status. The data encapsulation and transmission module is used to associate the voiceprint-location data packet and the classification result into a structured data packet. When there are disease markers, they are associated together and uploaded to the cloud service platform. The frequency domain feature map generation module is used to construct a set of disease coordinates based on structured data packets in the cloud service platform and convert the road surface acoustic wave signal into a frequency domain feature map. The spatial map generation module is used to fuse disease coordinate sets and frequency domain feature maps to generate spatial disease distribution maps.

[0118] The intelligent acoustic void detection system for cement pavement according to an embodiment of this application can implement any of the above methods, and the specific working process of each module in the system can be referred to the corresponding process in the above method embodiments.

[0119] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0120] This application also discloses a computer device.

[0121] A computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described intelligent acoustic void detection method for cement pavement.

[0122] This application also discloses a computer-readable storage medium.

[0123] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the intelligent acoustic void detection methods for cement pavement.

[0124] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0125] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A method for intelligent acoustic void detection in cement pavement, characterized in that, The detection method includes: Retrieve user configuration commands containing target area and driving speed parameters; In response to the user configuration command, the driverless device is controlled to travel along a preset path, while periodic mechanical impacts are applied to the road surface through the acoustic vibration wheel on the driverless device. The sound wave signal of the road surface generated by the mechanical impact is acquired in real time by the soundprint acquisition module, and the spatial coordinate data and timestamp of the unmanned driving equipment are acquired by the positioning system. The road surface acoustic wave signal is subjected to noise suppression and frequency domain transformation to generate a standardized acoustic signature feature vector. The road surface acoustic wave signal, acoustic feature vector, spatial coordinate data and timestamp are input into the multi-source fusion module to generate a spatiotemporally aligned acoustic-location data packet; The voiceprint feature vector is input into a pre-trained classification model group for real-time inference, generating a classification result that includes the empty state identifier and confidence level. When the void status indicator indicates a void status, a disease marking command is triggered based on the spatial coordinate data; The voiceprint-location data packet and classification results are associated with a structured data packet. When disease markers are present, they are also associated together and uploaded to the cloud service platform. In the cloud service platform, a set of disease coordinates is constructed based on the structured data packet, and the road surface acoustic wave signal is converted into a frequency domain feature map; By fusing the disease coordinate set and frequency domain feature map, a spatial disease distribution map is generated; The steps for generating a spatialized disease distribution map by fusing the disease coordinate set and frequency domain feature map include: Spatial interpolation is performed on the set of disease coordinates to generate a continuous disease density distribution surface; Extract the spectral feature vector associated with the disease coordinates from the frequency domain feature map; Map the spectral feature vector to the spatial grid corresponding to the disease density distribution surface; Based on the correlation analysis between the spectral feature vectors and disease density data within the spatial grid, a fused feature matrix is ​​generated. A two-dimensional spatialized disease distribution map is rendered based on the fused feature matrix.

2. The intelligent acoustic void detection method for cement pavement according to claim 1, characterized in that, The steps of performing noise suppression and frequency domain transformation on the road surface acoustic wave signal to generate a standardized acoustic signature vector include: Acquire road surface acoustic wave signals in real time collected by the acoustic signature acquisition module; The acoustic signal is subjected to noise suppression processing, and the pulse interference noise is suppressed by spectral subtraction to generate a denoised wave signal; The denoised signal is divided into equal-length time segments, and a windowing function is applied to each frame. Perform frequency domain transformation on the windowed time segment of equal length to generate a frequency domain energy spectrum; Multi-scale acoustic signature features are extracted from the frequency domain energy spectrum to generate the original feature vector; The original feature vector is normalized to generate a standardized voiceprint feature vector.

3. The intelligent acoustic void detection method for cement pavement according to claim 1, characterized in that, The steps of inputting the voiceprint feature vector into a pre-trained classification model group for real-time inference to generate a classification result containing a void state identifier and confidence level include: The voiceprint feature vectors are input in parallel into a pre-trained classification model group consisting of a support vector machine model and a deep residual network model; The wavelet packet energy entropy features in the voiceprint feature vector are classified using the support vector machine model, and the first classification probability is output. The Mel-Cepstral Coefficient features in the voiceprint feature vector are processed by the deep residual network model to output the second classification probability; The first classification probability and the second classification probability are weighted and fused to generate a fused classification probability. Based on the comparison result between the fusion classification probability and the preset confidence threshold, a void status identifier is generated; The fusion classification probability is used as the confidence level and associated with the empty state identifier to form the classification result.

4. The intelligent acoustic void detection method for cement pavement according to claim 1, characterized in that, When the void status indicator indicates a void status, the steps for triggering a disease marking command based on the spatial coordinate data include: Receive the classification results, which include the empty status identifier and the corresponding confidence level; When the empty status indicator shows an empty status, extract the spatial coordinate data synchronized with the classification results; Perform position optimization processing on the spatial coordinate data to generate marked position coordinates; A disease marking instruction is generated based on the marked location coordinates and associated with the current timestamp.

5. The intelligent acoustic void detection method for cement pavement according to claim 1, characterized in that, The steps of constructing a set of disease coordinates based on the structured data packet and converting the road surface acoustic wave signal into a frequency domain feature map in the cloud service platform include: Parse the disease marking instructions and associated spatial coordinate data in the structured data packet; Filter the spatial coordinate data associated with the disease markers to generate an initial set of disease coordinates; Perform spatial clustering analysis on the initial set of disease coordinates, and merge adjacent spatial coordinates to form a new set of disease coordinates; Extract the road surface acoustic wave signal from the acoustic signature-location data packet; The road surface acoustic wave signal is subjected to time-domain segmentation processing to generate a standardized time-domain signal frame sequence; Apply a windowed Fourier transform to each time-domain signal frame in the standardized time-domain signal frame sequence to generate the corresponding frequency-domain spectrum sequence; The frequency domain spectrum sequence is spliced ​​into a two-dimensional frequency domain feature map.

6. The intelligent acoustic void detection method for cement pavement according to claim 5, characterized in that, Following the step of generating a spatialized disease distribution map, the method further includes: Obtain a maintenance decision rule base, including the mapping relationship between disease severity classification standards and treatment measures; The confidence data and disease density data in the spatialized disease distribution map are weighted and fused together. Based on the weighted results, the treatment priority parameters in the maintenance decision rule base are matched to generate a maintenance decision map that includes the boundary of the treatment area and construction plan suggestions.

7. A smart acoustic void detection system for cement pavement, characterized in that, The detection system is used to perform the intelligent acoustic void detection method for cement pavement according to any one of claims 1 to 6, the detection system comprising: The detection task configuration module is used to obtain user configuration instructions containing target area and driving speed parameters; The execution control module is used to control the unmanned driving device to travel along a preset path in response to the user configuration command, while applying periodic mechanical impacts to the road surface through the acoustic vibration wheel on the unmanned driving device; The multi-source data acquisition module is used to acquire the road surface sound wave signal generated by the mechanical impact in real time through the acoustic fingerprint acquisition module, and to acquire the spatial coordinate data and timestamp of the unmanned driving device through the positioning system. The voiceprint feature extraction module is used to perform noise suppression and frequency domain transformation processing on the road surface sound wave signal to generate a standardized voiceprint feature vector. The spatiotemporal data fusion module is used to perform spatiotemporal alignment processing on the road surface acoustic wave signal, voiceprint feature vector, spatial coordinate data and timestamp to generate a voiceprint-location data packet. The intelligent classification module is used to input the voiceprint feature vector into a pre-trained classification model group for real-time inference and generate a classification result that includes the empty state identifier and confidence level. The disease marking triggering module is used to trigger a disease marking command based on the spatial coordinate data when the void status indicator indicates a void status. The data encapsulation and transmission module is used to associate the voiceprint-location data packet and the classification result into a structured data packet. When there is a disease marking instruction, it is associated with the data packet and uploaded to the cloud service platform. The frequency domain feature map generation module is used to construct a set of disease coordinates based on the structured data packet in the cloud service platform and convert the road surface acoustic wave signal into a frequency domain feature map. The spatial map generation module is used to fuse the disease coordinate set and frequency domain feature map to generate a spatial disease distribution map.

8. A computer device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and execute the method as described in any one of claims 1 to 6.