Domain adaptive disease identification method and system based on multi-frequency ground penetrating radar

By using a domain-adaptive defect identification method based on multi-band ground-penetrating radar, the problem of model performance degradation caused by frequency band differences in ground-penetrating radar equipment is solved, achieving high adaptability and accuracy in cross-band defect identification and outputting a comprehensive report to support road maintenance.

CN122110099AActive Publication Date: 2026-05-29EAST CHINA JIAOTONG UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EAST CHINA JIAOTONG UNIVERSITY
Filing Date
2026-04-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing disease identification models are difficult to reuse directly under different frequency bands of ground penetrating radar equipment, resulting in model performance degradation. Furthermore, disease identification is prone to misjudgment under different frequency bands, and they cannot effectively cope with the problems of differences in road structure type, material dielectric properties and environment.

Method used

By acquiring a multi-band ground-penetrating radar detection dataset, cross-band response difference deconstruction, disease physical constraint extraction, and multi-view representation construction are performed to generate a cross-domain disease representation information set. Furthermore, source-target domain collaborative-driven shared feature alignment and disease semantic discrimination are performed to generate a disease identification result information set, and finally, a multi-band domain adaptive disease identification report is output.

Benefits of technology

It enables one-time training applicable to multiple frequency bands, reduces the dependence on new frequency band data annotation, improves engineering efficiency, and outputs an integrated report of disease type, accurate three-dimensional location and severity level, supporting precise road maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122110099A_ABST
    Figure CN122110099A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent identification, in particular to a domain self-adaptive disease identification method and system based on a multi-frequency ground penetrating radar. The method comprises the following steps: acquiring a multi-frequency ground penetrating radar detection data set, performing cross-frequency response difference deconstruction, disease physical constraint extraction and multi-view representation construction based on the multi-frequency ground penetrating radar detection data set, and generating a cross-domain disease representation information set; performing source domain-target domain collaborative driving shared feature alignment and disease semantic discrimination based on the cross-domain disease representation information set, and generating a disease identification result information set; and performing positioning output and severity representation for a road disease target based on the disease identification result information set, and generating and outputting a multi-frequency domain self-adaptive disease identification report. In the nondestructive detection process of road engineering, the application greatly reduces the dependence on new frequency band data labeling, improves engineering practical efficiency, and provides direct and comprehensive decision support for road precise maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent identification technology, and in particular to a domain-adaptive disease identification method and system based on multi-band ground-penetrating radar. Background Technology

[0002] In the field of non-destructive testing and intelligent identification technology for road engineering, ground-penetrating radar (GPR) can image the pavement structure layers based on the differences in the electromagnetic properties of underground media, making it an important means of non-destructive testing for roads. Existing research has enabled the use of 3D GPR to locate and classify defects from multi-view radar images, including longitudinal and horizontal planes, and to identify various typical internal defects such as poor interlayer bonding, voids, cracks, and mixture segregation.

[0003] However, existing disease identification models typically rely on training samples under fixed equipment, fixed center frequencies, fixed road structures, or fixed operating conditions. When the acquisition equipment switches from low-frequency bands to high-frequency bands, or from a single frequency band to a full-bandwidth mode, the radar waveform will change significantly in terms of resolution, penetration depth, amplitude distribution, boundary sharpness, and texture scale, making it difficult to directly reuse the original model.

[0004] Furthermore, for defects such as poor interlayer contact and thin-layer voids in asphalt pavements, electromagnetic waves at the interlayer interfaces are affected by the interlayer dielectric difference, propagation path, and phase difference, resulting in significant thin-layer interference. These defects may exhibit different appearances at different frequency bands, such as single-peak enhancement, single-peak weakening, or waveform separation. If identification relies solely on traditional amplitude thresholds or single image features, it is highly likely that "apparent changes caused by frequency bands" will be misjudged as "changes in defect type."

[0005] Furthermore, in actual engineering projects, there are also differences in road surface structure types, material dielectric properties, temperature and humidity environments, and structural interference. Even if the defects themselves are the same, the radar image distribution will show significant drift, resulting in false detections, missed detections, and model instability when deploying across projects.

[0006] Therefore, it is necessary to propose a domain-adaptive disease identification method that can achieve feature transfer and distribution alignment across multiple frequency bands, structures, and operating conditions, so as to improve the model's cross-scenario reusability and engineering practicality. Summary of the Invention

[0007] This application provides a domain-adaptive disease identification method and system based on multi-band ground-penetrating radar to solve the above-mentioned technical problems.

[0008] In a first aspect, this application provides a domain-adaptive disease identification method based on multi-band ground-penetrating radar, the method comprising: A multi-band ground-penetrating radar (GPR) detection dataset is acquired. Based on this dataset, cross-band response difference deconstruction, physical constraint extraction of road surface defects, and multi-view representation construction are performed to generate a cross-domain defect representation information set. Based on this cross-domain defect representation information set, source-target domain collaborative-driven shared feature alignment and defect semantic discrimination are performed to generate a defect identification result information set. Based on this defect identification result information set, location output and severity representation of road surface defects are performed to generate and output a multi-band domain adaptive defect identification report.

[0009] The above technical solutions effectively overcome the problem of model performance degradation caused by frequency band differences in ground-penetrating radar equipment, achieving highly adaptive recognition that is applicable to multiple frequency bands with "one-time training"; significantly reducing the dependence on new frequency band data annotation and improving engineering efficiency; and finally outputting an integrated report on the type of damage, accurate three-dimensional location and severity level, providing direct and comprehensive decision support for precise road maintenance.

[0010] Optionally, the generation process of the cross-domain disease characterization information set includes: the multi-band ground-penetrating radar detection dataset includes source domain detection data and target domain detection data collected under different center frequencies and different bandwidth modes; based on the multi-band ground-penetrating radar detection dataset, radar image characterization information is generated through echo preprocessing and multi-view reconstruction; based on the multi-band ground-penetrating radar detection dataset, disease physical constraint analysis information is generated through inter-layer interface echo analysis and time-frequency response deconstruction; combining the radar image characterization information and the disease physical constraint analysis information, and combining domain attribute difference analysis, cross-band shared feature extraction and domain offset characterization decoupling are performed to generate the cross-domain disease characterization information set.

[0011] Optionally, the process of generating the radar image characterization information includes: based on the source domain detection data and the target domain detection data, performing time-zero correction, DC removal processing, background removal, gain compensation, and bandpass filtering to obtain initial purified echo data; based on the initial purified echo data, constructing an A-scan echo sequence, a B-scan longitudinal section image, and a horizontal slice image as the basic structural characterization for disease identification; performing amplitude normalization, layer tracking, and anomalous echo enhancement processing on the basic structural characterization to generate a comparable multi-view radar image feature set; summarizing the A-scan echo sequence, the B-scan longitudinal section image, the horizontal slice image, and the multi-view radar image feature set to generate the radar image characterization information used to characterize the significant differences in disease severity at different frequency bands.

[0012] Optionally, the process of generating the physical constraint analysis information of the disease includes: locating the interlayer interface reflection waveforms based on the source domain detection data and the target domain detection data, extracting the interface peak amplitude, peak polarity, and adjacent peak spacing as basic waveform clues for the interface response; simultaneously, performing two-way travel time measurement, dominant frequency offset analysis, and frequency band energy distribution analysis on the interlayer interface reflection waveforms to identify their corresponding propagation path differences and frequency response differences; calculating the interlayer approximate reflection coefficient and constructing a physical constraint parameter set characterizing the interlayer contact state based on the interface peak amplitude, peak polarity, two-way travel time, and frequency band energy distribution; and performing correlation mapping and consistency analysis between the basic waveform clues and the physical constraint parameter set to generate the physical constraint analysis information of the disease.

[0013] Optionally, the decoupling of cross-band shared feature extraction and domain offset representation includes: inputting multi-view features from the radar image representation information into a shared coding branch and a domain difference coding branch, and extracting intrinsic features and frequency band difference features respectively through parallel mapping; introducing the interlayer approximate reflection coefficient, the two-way travel time, and the frequency band energy distribution from the physical constraint analysis information of the disease into the shared coding branch to impose physical consistency constraints on the intrinsic features of the disease; analyzing the distribution offset degree of the frequency band difference features under different frequency bands, different equipment responses, and different pavement structures to generate a domain difference semantic description; and constructing a cross-domain disease representation information set containing shared disease semantic information and domain offset semantic information based on the intrinsic features of the disease and the domain difference semantic description.

[0014] Optionally, the process of generating the disease identification result information set includes: based on the cross-domain disease representation information set, determining whether the current target domain sample is a labeled adapted sample or an unlabeled migration sample; when the target domain sample is the labeled adapted sample, performing shared feature refinement and disease category boundary correction based on category supervision to generate a supervised identification result; when the target domain sample is the unlabeled migration sample, performing target domain feature alignment and category semantic compensation based on pseudo-label iteration to generate a migration identification result; associating and encapsulating the generated supervised identification result or the migration identification result with the corresponding disease location marker and confidence marker to generate the disease identification result information set.

[0015] Optionally, the target domain feature alignment and category semantic compensation based on pseudo-label iteration includes: performing initial disease category prediction on target domain samples based on the cross-domain disease representation information set, and selecting target domain samples that meet the confidence threshold condition as pseudo-label samples; aligning the pseudo-label samples with the source domain labeled and adapted samples in the input domain network, and using the maximum mean difference algorithm to achieve feature distribution convergence between the target domain and the source domain; during the feature distribution convergence process, introducing the inter-layer approximate reflection coefficient and the physical consistency loss of the two-way travel time to correct and compensate for the category semantic relationship of the pseudo-label samples; and jointly solidifying the iteratively updated target domain features and the category semantic relationship to complete the target domain feature alignment and category semantic compensation.

[0016] Optionally, the process of refining shared features and correcting disease category boundaries based on category supervision includes: based on the cross-domain disease characterization information set, performing category center aggregation on source domain samples and labeled adapted samples to construct shared feature prototypes corresponding to different disease categories; combining the disease physical constraint analysis information, analyzing the distribution differences of each disease category in interlayer approximate reflection coefficient, two-way travel time, and frequency band energy ratio to form physical prior boundaries; based on the shared feature prototypes and the physical prior boundaries, correcting the discrimination boundaries between adjacent disease categories by shrinking or expanding them to reduce category aliasing under cross-frequency band conditions; and outputting the category discrimination results after boundary correction as the supervised identification results.

[0017] Optionally, the generation process of the multi-band domain adaptive disease identification report includes: adaptively weighting and fusing the multi-view results in the disease identification result information set according to the identification confidence of longitudinal and horizontal viewpoints under different frequency bands to generate a joint identification result of the target disease; based on the joint identification result, jointly locating the longitudinal range, lateral range, and interlayer depth of the disease to generate a disease spatial distribution image; simultaneously, calculating the disease severity level according to the interlayer approximate reflection coefficient, the two-way travel time, and the category confidence, and forming a corresponding disease warning marker; integrating the joint identification result, the disease spatial distribution image, and the disease warning marker to generate the multi-band domain adaptive disease identification report.

[0018] Secondly, this application provides a domain-adaptive disease identification system based on multi-band ground-penetrating radar, the system comprising: The cross-domain disease characterization module is used to acquire a multi-band ground-penetrating radar detection dataset. Based on the multi-band ground-penetrating radar detection dataset, it performs cross-band response difference deconstruction, disease physical constraint extraction, and multi-view characterization construction to generate a cross-domain disease characterization information set. The disease identification module is used to perform source-target domain collaborative-driven shared feature alignment and disease semantic discrimination based on the cross-domain disease characterization information set to generate a disease identification result information set. The result output module is used to perform location output and severity characterization of road surface disease targets based on the disease identification result information set, and generate and output a multi-band domain adaptive disease identification report. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram illustrating an application scenario provided in one embodiment of this application; Figure 2 A flowchart of a domain adaptive disease identification method based on multi-band ground-penetrating radar provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a domain adaptive disease identification system based on multi-band ground penetrating radar provided in an embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0022] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0023] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.

[0024] Existing disease identification models typically rely on training samples under fixed equipment, fixed center frequencies, fixed road structures, or fixed operating conditions. When the acquisition equipment switches from low-frequency bands to high-frequency bands, or from a single frequency band to a full-bandwidth mode, the radar waveform will change significantly in terms of resolution, penetration depth, amplitude distribution, boundary sharpness, and texture scale, making it difficult to directly reuse the original model.

[0025] Based on this, this application provides a domain-adaptive disease identification method and system based on multi-band ground-penetrating radar.

[0026] Figure 1 This is a schematic diagram illustrating an application scenario provided by this application. In the process of non-destructive testing in road engineering, the method provided in this application significantly reduces the reliance on new frequency band data annotation, improves engineering efficiency, and provides direct and comprehensive decision support for precise road maintenance.

[0027] Specifically, the method of this application is applied to any server that communicates with a ground-penetrating radar data platform and obtains multi-band ground-penetrating radar detection datasets provided by the ground-penetrating radar data platform through the server.

[0028] For specific implementation details, please refer to the following examples.

[0029] Figure 2 This is a flowchart illustrating a domain-adaptive disease identification method based on multi-band ground-penetrating radar according to an embodiment of this application. The method of this embodiment can be applied to servers in the above scenarios. Figure 2 As shown, the method includes: S201. Obtain a multi-band ground-penetrating radar detection dataset. Based on the multi-band ground-penetrating radar detection dataset, perform cross-band response difference deconstruction, disease physical constraint extraction, and multi-view characterization construction to generate a cross-domain disease characterization information set.

[0030] A multi-band ground-penetrating radar (GPR) detection dataset can be a collection of raw radar echo signals collected by GPR devices with different center frequencies (such as 200MHz, 900MHz, 2GHz) and different bandwidths at different times, on different road sections, or under different detection modes. The data comes from a GPR data platform.

[0031] Cross-band response difference deconstruction can be a process of separating information that characterizes the inherent properties of the disease (such as the "hyperbolic" shape of cavities and the "phase axis" misalignment of cracks) from the original radar echo signal through signal processing and feature engineering, and the interference or variation information caused by radar equipment frequency band characteristics, environmental noise, etc.

[0032] Physical constraint extraction of road defects can be achieved by mining and quantifying parameters or features that are directly related to the dielectric properties of road materials and the physical laws of electromagnetic wave propagation from radar echo signals and have clear physical significance.

[0033] Multi-view characterization can be constructed from raw radar data to create feature representations from various perspectives or dimensions to comprehensively describe the disease. The most basic ones include: A-scan (single-channel waveform, reflecting the reflection intensity of a point as it changes with depth) and B-scan (longitudinal profile image, composed of continuous A-scans, reflecting the underground profile along the survey line).

[0034] A cross-domain disease characterization information set can be a structured feature set formed after the above deconstruction, extraction and construction. It is not a single feature vector, but a composite containing multiple components, such as: {shared disease features, domain difference features, physical constraint parameters, multi-view image features}.

[0035] Specifically, multi-band ground-penetrating radar (GPR) detection is a common practice in engineering projects. Radars with different center frequencies (e.g., 400MHz and 2GHz) have inherent differences in resolution, penetration depth, and waveform details. This "domain offset" causes a high-precision identification model trained on data from a single frequency band to experience a sharp performance drop when directly applied to data from another frequency band. Essentially, the model misinterprets the differences in data distribution caused by the device's frequency band characteristics as changes in the disease characteristics themselves. Existing cross-device disease identification methods mostly rely on targeted re-labeling and training, which is costly and inefficient. This solution addresses this core pain point by systematically deconstructing cross-band response differences. It aims to separate the original mixed signals into "shared features" that characterize the essence of the disease and "domain-differential features" that reflect the device's frequency band characteristics. Simultaneously, it introduces physical constraints on the disease derived from the principles of electromagnetic wave propagation (e.g., reflection coefficient, two-way travel time) to provide interpretable, cross-domain, and universally applicable prior knowledge for the features. Furthermore, it constructs multi-view representations such as B-scan and horizontal slicing to comprehensively describe the disease. The cross-domain disease characterization information set generated by this series of processes lays a robust and semantically rich data foundation for subsequent adaptive identification that can be applied to multiple frequency bands with a single model.

[0036] S202. Based on the cross-domain disease representation information set, perform source domain-target domain collaborative-driven shared feature alignment and disease semantic discrimination to generate a disease identification result information set.

[0037] Source-target domain collaborative driving considers both source domain (labeled and knowledge-rich) and target domain (few or no labels and needing adaptation) data during model training and inference, enabling both to jointly guide the updating of model parameters.

[0038] Shared feature alignment can be achieved by optimizing algorithms to make the distribution of shared disease features extracted from source domain data and target domain data as close as possible or overlap in the feature space. Ideally, the shared feature vectors of the same type of disease should be clustered in the same area in the feature space, regardless of which frequency band they come from.

[0039] Semantic disease discrimination can be based on the alignment of shared features, and then use a classification network to classify the features and identify specific disease types (such as cracks, voids, poor interlayer bonding, loosening, etc.).

[0040] The disease identification result information set can be a data set containing detailed information about each detected disease. Each record should include at least: the disease's unique identifier ID, the predicted disease category, the confidence score of the prediction, and the preliminary location information of the disease in the original radar data (such as the B-scan image index and approximate depth range).

[0041] Specifically, after obtaining structured cross-domain disease representations, the core challenge lies in how to reliably identify target domain data (such as field data from another low-frequency radar) using knowledge from well-labeled source domain data (e.g., a fine-grained dataset from a high-frequency radar) with sparsely labeled or completely unlabeled target domain data. Traditional transfer learning methods, with their simple fine-tuning, are prone to causing the model to forget source domain knowledge or overfit to the target domain, making it difficult to achieve a balance. This solution addresses this key challenge of adaptive learning by employing a source-target domain collaborative driving mechanism. During model training, the data from both domains work together, forcing the network to extract features that can accurately classify in the source domain while remaining indistinguishable between the two domains. This achieves shared feature alignment, essentially unifying the model's "cognitive perspective" on data from different frequency bands. Based on this, disease semantic discrimination is performed, mapping the aligned, physically constrained features to specific disease categories. The disease identification result information set generated in this step not only includes category prediction but also associates preliminary location and confidence levels, providing a core basis for the final decision.

[0042] S203. Based on the information set of road surface defects identification results, perform location output and severity characterization for road surface defects, and generate and output a multi-frequency domain adaptive defect identification report.

[0043] A multi-band domain adaptive disease identification report can be a comprehensive electronic or paper report that integrates all analysis results. The report should include at least the following: an overview of the detection (time, road segment, radar band used), a summary table of disease statistics, detailed information for each disease (type, precise three-dimensional coordinates, size, severity level, confidence level, corresponding radar image screenshot), and a summary of maintenance recommendations based on severity and spatial distribution.

[0044] Specifically, after obtaining preliminary disease identification results, it is still necessary to address the ultimate needs of engineering maintenance: Where exactly is the disease located? What is its severity? This is to guide maintenance operations and resource allocation. Many existing intelligent identification methods stop at category output, lacking spatial and status quantification information directly linked to maintenance decisions. This solution addresses this last-mile problem from "identification" to "application." Based on the initial judgment results, it outputs the location, integrates multi-view information to refine the extent and boundaries of the disease in three-dimensional space; simultaneously, it combines the physical constraint parameters extracted in step one (such as the abnormal amplitude of the reflection coefficient) and the geometric dimensions of the disease to characterize its severity, achieving a leap from "presence or absence of disease" to "disease risk level." Finally, an automated multi-band domain adaptive disease identification report is generated by integrating all information, presenting the detection overview, detailed disease list, and maintenance recommendations in a visualized and structured format.

[0045] The method provided in this embodiment firstly standardizes and preprocesses the raw multi-band radar data and performs time-frequency analysis to construct multi-view representations such as A-scan, B-scan, and horizontal slices. By introducing physical parameters such as electromagnetic wave reflection coefficient and two-way travel time as cross-domain invariant constraints, and combining them with a deep learning dual-branch network, the essential characteristics of the defects and the frequency band characteristics of the equipment are decoupled to form a structured cross-domain defect representation. Then, methods such as domain adversarial training or maximum mean difference alignment are used to drive the defect features of the source domain (such as high-frequency labeled data) and the target domain (such as low-frequency unlabeled data) to align in the latent space. On this basis, the aligned shared features are used to perform semantic classification and preliminary localization of the defects. Finally, based on multi-view confidence fusion, the defects are accurately located, and a comprehensive intelligent report including severity assessment is generated by integrating physical parameters, geometric dimensions, and model confidence, and output to the R&D personnel. This solution effectively overcomes the performance degradation problem caused by the frequency band difference of ground penetrating radar equipment, and achieves highly adaptive identification with "one-time training and application to multiple frequency bands"; it greatly reduces the dependence on new frequency band data annotation and improves the efficiency of engineering application; the final output of the integrated report of disease type, accurate three-dimensional location and severity level provides direct and comprehensive decision support for precise road maintenance.

[0046] In some embodiments, the multi-band ground-penetrating radar (GPR) detection dataset includes source domain detection data and target domain detection data acquired under different center frequencies and bandwidth modes. Based on the multi-band GPR detection dataset, radar image characterization information is generated through echo preprocessing and multi-view reconstruction. Based on the multi-band GPR detection dataset, physical constraint analysis information of the disease is generated through inter-layer interface echo analysis and time-frequency response deconstruction. Combining the radar image characterization information and the physical constraint analysis information of the disease, and combining domain attribute difference analysis, cross-band shared feature extraction and domain offset characterization decoupling are performed to generate a cross-domain disease characterization information set.

[0047] Source domain detection data can be radar datasets with detailed disease annotations used for model training. It usually comes from historical projects, laboratory simulations, or heterogeneous forward simulations and contains known disease types, locations, and severity labels, serving as a knowledge base for model learning; for example, data of a highway section collected and annotated using a 1.5 GHz center frequency radar.

[0048] Target domain detection data can be new radar datasets that are to be identified, with scarce or no labels. It comes from the actual engineering scenarios that need to be detected. Its data distribution (due to differences in frequency band, equipment, and environment) differs from the source domain and is the object that the model needs to generalize to. For example, field data collected by a new radar with a center frequency of 900MHz on another urban road.

[0049] Radar image representation information can be a set of image features generated after a series of standardized preprocessing and multi-view reconstruction of the original radar echo signal.

[0050] Physical constraint analysis information for diseases can be a set of quantitative parameters that are analyzed and calculated from radar echoes based on the principle of electromagnetic wave propagation, and are used to describe the state of interlayer interfaces and the physical properties of diseases.

[0051] Domain attribute difference analysis can be a systematic process of identifying, quantifying, and labeling the fundamental attribute factors that cause differences in radar data distribution.

[0052] Shared feature extraction can be a process of automatically learning and extracting essential features that are common and stable to the disease itself from multi-frequency and multi-domain data through a specific model structure (such as a shared weight encoder).

[0053] Domain offset representation decoupling can be a mechanism that actively separates "shared features" from "unique features" caused by differences in domain attributes during the feature extraction process through design (such as using adversarial learning, dedicated network branches, etc.).

[0054] Specifically, traditional methods for processing multi-band radar data typically involve mixing all data and directly inputting it into the network for training. The model may misinterpret changes in image texture and resolution caused by frequency band differences as features of the disease category; for example, it might misclassify a blurry crack under 900MHz radar as "poor interlayer adhesion." The necessity of this step lies in its structured process: first, standardizing and expanding the perspective of the original signal; then, injecting physical laws as supervision; and finally, using the network architecture to forcibly separate the essence of the disease from its frequency band appearance. This constructs a high-quality input that is both rich in information and feature-decoupled for subsequent domain-adaptive recognition.

[0055] In the specific analysis process, firstly, a consistent preprocessing pipeline is applied to the raw A-scan data of the source and target domains: zero-correction is used to align the start time, a bandpass filter (such as a Chebyshev Type I filter) is applied to retain the effective frequency band and suppress high and low frequency noise, the direct wave background is removed by the moving average method, and an exponential gain function is used to compensate for depth attenuation to obtain the initial purified echo data. Subsequently, a multi-view representation is reconstructed: consecutive A-scans are arranged to generate B-scan grayscale images; the three-dimensional volume data is projected to maximum amplitude within a specific depth range (such as within the pavement structure layer) to generate horizontal slice images; and the B-scan images are subjected to layer calibration and amplitude normalization based on phase axis tracking to generate enhanced longitudinal profile feature maps. These three elements together constitute the radar image representation information. Simultaneously, physical constraint analysis of the defects is performed: In the preprocessed A-scan sequence, the interlayer interface reflected wave is located by finding zero-crossing points and extreme points, and its peak amplitude A_peak and polarity P are extracted; the two-way travel time Δt of the reflected wave is calculated; a short-time Fourier transform is performed on this waveform segment to calculate its dominant frequency f_center and high-low frequency energy ratio E_ratio. Next, the interlayer reflection coefficient r is approximately calculated using the formula r≈(A_peak*c) / (2*Z0*sqrt(ε)) (where c is the speed of light, Z0 is the free-space wave impedance, and ε is the estimated dielectric constant of the upper layer). {r, Δt, f_center, E_ratio} are packaged into a physical constraint parameter set. Finally, feature extraction and decoupling are performed: a dual-branch coding network is constructed. Radar image representation information (such as image patch sequences) is input into the network. Features F_shared are extracted through a shared coding branch (e.g., composed of multiple TransformerEncoder layers). Simultaneously, the domain attribute labels of the representation information (such as frequency encoding and structure type encoding) are input along with the image into another domain difference coding branch (the structure of which may be the same as or shallower than the shared branch) to extract features F_domain. During the training of the shared coding branch, a set of physical constraint parameters is used as auxiliary supervision signals. For example, it ensures that the norm of a subspace vector of F_shared extracted from an image region containing a highly reflective interface is strongly correlated with the calculated r value, thereby imposing a physical consistency constraint. By optimizing a domain discrimination loss (e.g., making it difficult for a classifier to distinguish whether F_shared comes from the source domain or the target domain), F_shared is encouraged to extract as much domain information as possible. Ultimately, F_shared, F_domain, and the original set of physical constraint parameters together constitute a structured cross-domain disease representation information set.

[0056] In alternative or modified implementations, multi-view reconstruction may exclude horizontal slicing and use only B-scan and its enhanced features. In physical constraint analysis, the reflection coefficient r can be calculated using a simplified amplitude ratio method or by directly regressing from the waveform using a neural network. Decoupling shared features from domain-discrete features can be achieved without a two-branch network, using a gradient inversion layer: after a single feature extraction network, features are simultaneously fed into both the disease classifier (main task) and the domain classifier (auxiliary task), but during backpropagation to the domain classifier, the gradient is multiplied by a negative coefficient, forcing the network to learn features that are useless to the domain classifier, i.e., domain-invariant. The domain discrimination loss can also be replaced with a maximum mean difference loss, directly minimizing the distance between the source and target domain F_shared features in the regenerating kernel Hilbert space.

[0057] In some embodiments, based on source domain detection data and target domain detection data, time-zero correction, DC removal processing, background removal, gain compensation, and bandpass filtering are performed to obtain initial purified echo data. Based on the initial purified echo data, A-scan echo sequences, B-scan longitudinal section images, and horizontal slice images are constructed as the basic structural representation for disease identification. Amplitude normalization, layer tracking, and anomalous echo enhancement processing are performed on the basic structural representation to generate a comparable multi-view radar image feature set. The A-scan echo sequences, B-scan longitudinal section images, horizontal slice images, and multi-view radar image feature sets are summarized to generate radar image representation information for characterizing the significant differences in diseases at different frequency bands.

[0058] Zero-time correction refers to correcting the "zero-time" drift in the radar acquisition system caused by factors such as circuit delay and cable length, ensuring that the starting position of all A-scan data on the time axis accurately corresponds to the instant when the antenna transmits electromagnetic waves.

[0059] DC removal processing can refer to eliminating the DC (zero-frequency) component in radar signals introduced by hardware bias or environmental interference.

[0060] Background removal can refer to subtracting systematic, large-scale, and slowly varying background responses (such as direct waves, ground reflections, and reflections from fixed interfaces) from radar profile data to highlight local anomalies or relative changes in target objects.

[0061] Gain compensation can refer to compensating for the increase of radar signal amplitude with depth (or two-way travel time) to correct the energy attenuation caused by geometric diffusion and medium absorption when electromagnetic waves propagate in underground media, so that the weak reflection information in deep areas can be highlighted.

[0062] Bandpass filtering can be achieved by using digital filters to allow radar signal components within a specific frequency range (passband) to pass through, while suppressing noise and interference outside the passband range (stopband).

[0063] Initial purified echo data refers to intermediate data obtained after the original radar echo signal has undergone a series of standardized preprocessing steps (including time zero correction, DC removal, background removal, gain compensation, and bandpass filtering). It has eliminated major acquisition system errors, environmental background interference, and irrelevant frequency band noise, and retained the effective signal that reflects the differences in the true electromagnetic properties of the underground medium after energy correction.

[0064] A-scan echo sequence can be a single-channel voltage signal sequence received by radar at a single antenna location point, which varies with time (or equivalent depth). It is the most basic radar data unit, which directly reflects the intensity and time relationship of reflection and scattering events generated when electromagnetic waves encounter different medium interfaces on the vertical downward propagation path at that point. It is the basis for all images and advanced features.

[0065] A B-scan longitudinal profile image is a two-dimensional image formed by arranging multiple A-scan data collected continuously along a survey line in spatial order and displaying them in grayscale. Its horizontal axis represents the survey line distance, the vertical axis represents the two-way travel time (equivalent depth), and the pixel grayscale value represents the echo amplitude.

[0066] A horizontal slice image can be a two-dimensional planar image formed by extracting the amplitude of all A-scans at a specific depth point within a certain two-way travel time (depth) or a narrow depth range from a three-dimensional radar data volume and arranging them according to the antenna plane position.

[0067] Specifically, traditional methods for processing multi-band radar data often employ simple or fixed preprocessing parameters, neglecting the inherent differences in noise characteristics, attenuation patterns, and background responses among different frequency bands. This results in preprocessed data still retaining a large number of artifacts and intensity differences related to equipment and the environment, making subsequent cross-domain feature extraction and comparison difficult, and allowing models to easily learn these "non-pathological" patterns left over from preprocessing. The necessity of this step lies in using a systematic and parameterizable preprocessing pipeline tailored to the physical characteristics of radar signals to "purify" and "normalize" the raw echo data to a standardized and comparable state to the greatest extent possible, laying a reliable foundation for constructing image representations centered on real underground structures that are not dominated by acquisition conditions.

[0068] In the specific analysis process, the system sequentially executes the preset standardized preprocessing procedures: 1) Time-zero correction: First, for each A-scan data channel, the time zero offset is determined by finding the moment when the signal amplitude first exceeds the noise threshold (e.g., 3 times the standard deviation) or by calculating the peak value through cross-correlation with the reference template channel, and all channels are cyclically shifted and aligned. 2) DC removal: For each corrected A-scan channel, the arithmetic mean of all its sampling points is calculated, and then this mean is subtracted from each sampling point of that channel. 3) Background removal: A sliding window averaging method is used. For each row of the B-scan data matrix (at the same time depth), a moving window of length L (e.g., 11 channels) is set along the survey line direction. The average amplitude of all A-scans within the window at that time depth is calculated as the background estimate of the center channel of the window at that depth point. All channels and all depth points are traversed to form a background profile, which is then subtracted from the original profile. 4) Gain Compensation: A time-varying exponential gain function G(t) = exp(α*t) is applied, where t is the two-way travel time and α is a pre-set gain coefficient based on the medium attenuation characteristics. The amplitude of each A-scan channel at each time point is multiplied by the corresponding G(t). 5) Bandpass Filtering: Based on the acquisition frequency band label of the current data, the pre-set Butterworth bandpass filter coefficients are called. For example, for nominal 1.5GHz data, a 6th-order digital filter with a passband of 0.8GHz to 2.2GHz is used for zero-phase filtering. After preprocessing, multi-view construction and feature enhancement are performed: The processed A-scan channels are arranged in the acquisition order to directly form the A-scan echo sequence; they are stacked into a two-dimensional matrix and displayed in grayscale image form to form the B-scan longitudinal profile image; for the three-dimensional data volume, within a certain structural layer depth range determined by layer tracking, the maximum amplitude value is extracted along the vertical direction to generate a horizontal slice image. Next, feature-level processing is performed: amplitude normalization is applied to the B-scan image, linearly scaling the pixel values ​​of the entire image to the [0,1] interval. Layer tracking is performed using the energy aggregation method, searching for the continuous path with the highest energy along the time axis at each location point to identify the main layer interfaces. Finally, anomalous echo enhancement is performed: within the neighborhood of the identified layer (e.g., 5 sampling points above and below), the deviation of each pixel from the mean of its local window (e.g., 5 channels × 5 points) is calculated. The deviation value is multiplied by an enhancement coefficient and then superimposed back onto the original image, thus highlighting local anomalies. In the field of ground-penetrating radar, "channel" specifically refers to "Trace" or "A-scan," which is a complete time-domain waveform signal acquired along the survey line direction, with the antenna moving at a fixed interval (e.g., 1 cm). It represents a one-dimensional data sequence in the longitudinal (depth) direction. "Point" refers to a sampling point in the time (corresponding to depth) direction within a "channel" (A-scan) signal. For example, a channel may have 512 time sampling points.The “5 lines × 5 points” above describes a local three-dimensional data block extracted from a three-dimensional radar data volume.

[0069] In alternative or modified implementations, time-zero correction can be achieved using precise synchronization based on hardware trigger signals, provided the device supports it. Background removal can use median filtering instead of moving average to better suppress isolated strong interference points. Gain compensation can employ a combination model of spherical diffusion compensation (proportional to time) and absorption compensation (combined with an exponential function). The bandpass filter design can utilize Chebyshev filters to obtain steeper roll-off characteristics, or dynamically adjust the passband range based on spectral analysis of measured data. Layer tracking can use a U-Net neural network to perform semantic segmentation on B-scan images, directly outputting the layer category for each pixel. Anomaly echo enhancement can use the Retinex image enhancement algorithm or directional filters to enhance linear anomalies (such as cracks) with specific orientations.

[0070] In some embodiments, the interlayer interface reflection waveforms in the source domain detection data and target domain detection data are used for localization, and the interface peak amplitude, peak polarity, and adjacent peak spacing are extracted as basic waveform clues for the interface response. Simultaneously, the interlayer interface reflection waveforms are subjected to two-way travel time measurement, dominant frequency offset analysis, and frequency band energy distribution analysis to identify the corresponding propagation path differences and frequency response differences. Based on the interface peak amplitude, peak polarity, two-way travel time, and frequency band energy distribution, the interlayer approximate reflection coefficient is calculated, and a set of physical constraint parameters characterizing the interlayer contact state is constructed. The basic waveform clues and the physical constraint parameter set are correlated and mapped and analyzed for consistency to generate physical constraint analysis information for the disease.

[0071] The interlayer interface reflection waveform can be a strong reflection in-phase axis with a certain temporal and spatial continuity, caused by the difference in dielectric properties between different structural layers of the road surface (such as surface layer and base layer) in a radar B-scan image.

[0072] The peak amplitude of the interface can be the amplitude of the peak or trough with the largest absolute value in the extracted single-channel interlayer interface reflection waveform.

[0073] The polarity of the wave crest can be the phase direction (positive or negative) of the first major wave crest of the reflection waveform at the interlayer interface.

[0074] The spacing between adjacent peaks can be the time difference between the first adjacent secondary peak (or trough) and the main peak on both sides of the main peak (or trough) of the reflected waveform at the interlayer interface.

[0075] Two-way travel time can be the total propagation time of an electromagnetic wave from the radar antenna to the target interlayer interface and back to the antenna.

[0076] The dominant frequency offset can be the amount of drift of the instantaneous dominant frequency relative to the center frequency of the radar antenna after performing time-frequency analysis (such as wavelet transform) on the reflection waveform of the interlayer interface.

[0077] Frequency band energy distribution can be the proportion of signal energy in different frequency ranges (such as low frequency band, mid frequency band, and high frequency band) after performing spectral analysis on the reflection waveform of the interlayer interface. Different diseases (such as voids and water content) will cause characteristic changes in the spectral morphology of the reflection signal.

[0078] The interlayer approximate reflection coefficient can be a physical quantity that characterizes the intensity of electromagnetic wave reflection at the interlayer interface. The calculation formula is r=(Ap-Ab) / (Ap+Ab), where Ap is the peak amplitude of the interface and Ab is the reference background amplitude estimated based on the properties of the upper and lower layers or extracted from the data. The larger the absolute value of r, the greater the difference in dielectric between the layers, and the worse the contact condition may be.

[0079] Specifically, traditional disease identification methods heavily rely on the apparent features of radar images, such as texture and shape. However, for diseases like poor interlayer contact and thin-layer voids, their image features are greatly affected by the radar center frequency (potentially showing single-peak at low frequencies and double-peak at high frequencies). Relying solely on image features can easily lead to cross-frequency band misjudgments. The necessity of this step lies in directly starting from the physical essence of radar echoes (wave equation) to extract quantitative parameters (such as reflection coefficient and travel time) that are decoupled from the image appearance and have clear physical meaning. This provides a stable basis for the core discrimination logic, unaffected by frequency band appearance changes, fundamentally improving the interpretability and cross-domain robustness of interlayer disease identification.

[0080] In the specific analysis process, the system first locates the interlayer interface reflection waveform: based on the pre-processed B-scan image and layer tracking results, along each measurement point position (channel number), the A-scan signal segment of the channel is extracted as the interlayer interface reflection waveform within the time window corresponding to the target layer (such as the surface layer-base layer interface) (e.g., 10 sampling points above and below the layer line). Subsequently, feature extraction is performed in parallel: 1) Extracting basic waveform clues: in the extracted waveform segments, the peak point with the largest absolute value is found, and its amplitude is recorded as Ap; the sign of the peak point relative to the waveform baseline (zero value) is determined to determine the peak polarity (positive or negative); with the main peak as the center, the secondary peak point after the first zero crossing point is searched forward and backward, and the time difference between it and the main peak is calculated to obtain the adjacent peak spacing ΔT. 2) Analytical Propagation and Frequency Response: Determine the sampling time t_peak corresponding to the peak point of the energy envelope of the waveform segment, subtract the antenna direct wave time t0, and obtain the two-way travel time Δt; perform continuous wavelet transform on the waveform segment, calculate the center frequency f_instant corresponding to the scale with the highest energy in its scalogram, and subtract it from the radar system center frequency f_center to obtain the main frequency offset Δf; simultaneously, divide the power spectrum of the waveform segment into three frequency bands: low frequency (e.g., 0-fc / 3), mid frequency (fc / 3-2fc / 3), and high frequency (2fc / 3-fc), calculate the proportion of energy in each frequency band to the total energy, and obtain the frequency band energy distribution vector [E_low, E_mid, E_high], where E_low (low frequency band energy) refers to the radar signal in the low frequency subband (e.g., 0.5 - 1.5). The energy integral within the GHz range mainly reflects the scattering and attenuation characteristics of deep or large-scale structures (such as base layer voids and thick, loose layers); the E_mid (mid-frequency band energy) refers to the energy integral of the radar signal within the mid-frequency sub-band (e.g., 1.5 - 2.5 GHz), which is most sensitive to dielectric changes and interface reflections of common mid-layer defects (such as interlayer voids and moderate honeycombing); the E_high (high-frequency band energy) refers to the energy integral of the radar signal within the high-frequency sub-band (e.g., 2.5 - 3.5 GHz), which mainly captures the subtle scattering and diffraction characteristics caused by shallow, fine structures (such as surface cracks and small-scale spalling). 3) Calculate the physical constraint parameters: The interlayer approximate reflection coefficient r is calculated using the formula r=(Ap-Ab) / (Ap+Ab). The reference background amplitude Ab is preset in the following way: On the B-scan image, a region that has been manually confirmed or determined by the model to be "structurally intact" is selected, and the interlayer interface reflection waveforms of all measuring points within this region are extracted. The statistical median of its peak amplitude is calculated as Ab. Finally, the system correlates Ap, polarity, ΔT, Δt, Δf, [E_low,E_mid,E_high] and the calculated r to construct a complete physical constraint parameter set record for each analysis point, forming physical constraint analysis information for the disease.

[0081] In alternative or modified implementations, the location of the interlayer interface reflection waveform can be precisely determined by matching it with a standard healthy interface waveform template using a waveform matching algorithm based on Dynamic Time Warping (DTW). The dominant frequency offset analysis can use Hilbert-Huang Transform (HHT) instead of wavelet transform to adaptively obtain the instantaneous frequency. The calculation of the frequency band energy distribution can be directly based on the power spectrum after Fast Fourier Transform (FFT) partitioning and integration. The preset reference background amplitude Ab can be more dynamic; for example, based on the theoretical value of the material dielectric constant in the current road design drawings, the theoretical reflection amplitude can be generated as Ab through forward modeling; or, when a healthy area cannot be obtained, a sliding window can be used to statistically analyze the median of the interface peak amplitudes of several adjacent roads (e.g., five roads before and after) as a local adaptive Ab.

[0082] In some embodiments, multi-view features from radar image characterization information are input into a shared coding branch and a domain difference coding branch, and intrinsic features and frequency band difference features are extracted respectively through parallel mapping. The interlayer approximate reflection coefficient, two-way travel time, and frequency band energy distribution from the physical constraint analysis information of the disease are introduced into the shared coding branch to impose physical consistency constraints on the intrinsic features of the disease. The distribution offset of frequency band difference features under different frequency bands, different equipment responses, and different pavement structures is analyzed to generate a domain difference semantic description. Based on the intrinsic features of the disease and the domain difference semantic description, a cross-domain disease characterization information set containing shared disease semantics and domain offset semantics is constructed.

[0083] The shared coding branch and the domain difference coding branch can refer to two feature extraction sub-networks set up in parallel in a domain adaptive recognition network. The shared coding branch is responsible for extracting features (intrinsic features of the disease) that are essentially related to the disease and remain stable across different frequency bands and devices from the input; the domain difference coding branch is responsible for extracting features (frequency band difference features) that are related to a specific frequency band, device response, or acquisition conditions. The two are functionally separated through different network parameters and training objectives.

[0084] The intrinsic characteristics of a disease can be feature vectors output by a shared coding branch that can characterize the disease's intrinsic physical properties (such as abnormal morphology, interface discontinuity, and media variation).

[0085] Frequency band difference features can be feature vectors output by the domain difference coding branch, which are mainly used to distinguish which specific frequency band or domain the data comes from.

[0086] Physical consistency constraints can be introduced by adding an additional loss function term when training shared coding branches, requiring that the disease-related features extracted by the network be consistent with the pre-calculated physical parameters (such as the interlayer approximate reflection coefficient r and the two-way travel time Δt) in terms of statistical trends or semantics.

[0087] Domain difference semantic description can be structured information obtained by statistical analysis or further encoding of frequency band difference features, used to quantitatively characterize the distribution differences between different domains (such as different frequency bands).

[0088] Specifically, traditional deep learning models, when processing mixed multi-frequency band data, learn all features indiscriminately, resulting in network weights encoding both disease information and frequency band difference information. This causes the model to perform well on the training set, but once applied to new frequency band data (the target domain), the discrimination boundary of the entire model fails because the frequency-related features shift. The necessity of this step lies in designing a parallel dual-branch network structure to force the model to decouple the two types of information, "what is the disease" and "which frequency band the data comes from," from the root of the problem by isolating the interference of domain shift, so that the core disease discrimination relies only on stable and highly generalizable intrinsic disease features.

[0089] In the specific analysis process, the system first constructs and initializes a dual-branch coding network. The shared coding branch and the domain difference coding branch can be constructed based on the same backbone network (such as ResNet-18), but have independent network parameters. The processing flow is as follows: 1) Feature extraction: The generated multi-view radar image feature set (such as normalized B-scan image blocks) is simultaneously input into both branches. The shared coding branch outputs a defect intrinsic feature vector F_c; the domain difference coding branch outputs a frequency band difference feature vector F_d. 2) Physical consistency constraint: During the training phase, the calculated inter-layer approximate reflection coefficient r and two-way travel time Δt are normalized and used as an additional physical label. A regression constraint loss is designed: a small regression subnetwork is added after the shared coding branch to attempt to predict r and Δt from F_c, and the mean squared error (MSE) between the predicted value and the true physical label is calculated as the physical consistency loss L_phys. This forces F_c to contain information related to these physical quantities. 3) Generate domain-specific semantic descriptions: During the training phase, the frequency band difference feature vectors F_d of all samples are grouped according to their respective frequency band labels (e.g., 500MHz, 1GHz, 2GHz). The mean vector of F_d within each frequency band group is calculated. The differences between these mean vectors constitute the most basic domain-specific semantic description, which characterizes the "anchor" positions of different frequency bands in the feature space. 4) Construct a cross-domain disease representation information set: For each sample, its intrinsic disease feature vector F_c, frequency band difference feature vector F_d, and the domain-specific semantic description of its respective frequency band (e.g., the mean vector of that frequency band) are concatenated to form the final representation information set.

[0090] In alternative or modified implementations, the dual-branch network can be designed with a partially shared and partially independent parameter architecture to reduce model complexity. Physical consistency constraints can be implemented using contrastive learning instead of regression: requiring that for sample pairs with similar r and Δt values, the distance of their F_c in the feature space should be less than that of sample pairs with significantly different values. Generating domain-discrepancy semantic descriptions can be done in a more dynamic way, such as training a domain classifier that takes F_d as input and outputs the probability distribution of its frequency bands, using this probability distribution as the domain-discrepancy semantic description.

[0091] In some embodiments, based on the cross-domain disease representation information set, it is determined whether the current target domain sample is a labeled adapted sample or an unlabeled migration sample; when the target domain sample is a labeled adapted sample, shared feature refinement and disease category boundary correction based on category supervision are performed to generate supervised identification results; when the target domain sample is an unlabeled migration sample, target domain feature alignment and category semantic compensation based on pseudo-label iteration are performed to generate migration identification results; the generated supervised identification results or migration identification results are associated and encapsulated with the corresponding disease location markers and confidence markers to generate a disease identification result information set.

[0092] Labeled and adapted samples can refer to a small number of radar data samples in the current new project (target domain) that have been labeled with precise disease categories and location information by manual or high-precision auxiliary means.

[0093] Unlabeled migration samples can refer to radar data samples that constitute the vast majority of the current new project (target domain) and do not have artificial disease labels.

[0094] Specifically, in actual engineering deployments, new projects typically have only a very few key locations with precise annotations (annotated adaptation samples), while massive amounts of data are unlabeled (unlabeled transfer samples). Traditional methods, if fine-tuned with only a small number of labeled samples, are prone to overfitting; if labeled samples are completely ignored, the transfer effect is poor. The necessity of this step lies in proposing a hybrid supervision strategy that "divide and conquers" data of different properties in the target domain: using a small number of valuable annotations to directly correct the model's decision boundary; and for a large amount of unlabeled data, using a safe pseudo-label self-training technique for progressive alignment, thereby achieving optimal cross-domain model performance improvement with the lowest annotation cost.

[0095] In the specific analysis process, the system first determines the sample type: reads the target domain sample metadata and checks whether a disease category label field exists. If it exists, it is determined as an "annotated and adapted sample" and enters the supervised learning process; otherwise, it is determined as an "unlabeled transfer sample" and enters the transfer learning process. 1) For annotated and adapted sample: the sample and its true label are input into the network along with the source domain labeled sample. During training, in addition to the regular classification cross-entropy loss, an additional center loss is introduced. This loss calculates the distance from the "disease intrinsic features" of all samples (including source domain and current target domain samples) under the same disease category to their category center and minimizes it, thereby achieving shared feature refinement. At the same time, combined with the calculated physical constraint parameter set, the physical parameter clustering interval of the same type of disease is analyzed, which is used as the physical prior boundary. In the decision layer of the classifier, the classification hyperplane close to the physical prior boundary is fine-tuned according to the distribution of the target domain samples to achieve disease category boundary correction. Finally, a supervised recognition result with high confidence is output. 2) For unlabeled transfer samples: the pseudo-label iteration process is started. First, the current model is used to predict all unlabeled samples. Samples with prediction confidence scores higher than a preset threshold (e.g., 0.9) are selected, and pseudo-labels are generated using their predicted categories. These pseudo-labeled samples are then merged with the source domain data, and target domain feature alignment is achieved by minimizing the maximum mean difference (MMD) between the distributions of the two "intrinsic disease features". For category semantic compensation, the physical parameters (e.g., r, Δt) of these pseudo-labeled samples are calculated. If their parameter values ​​significantly deviate from the physical prior boundary of the disease class, the weight of the pseudo-label in the loss calculation is significantly reduced, or it is directly removed. This prediction-selection-alignment-compensation process is iterated until the model converges, outputting the transfer recognition results. Finally, all recognition results are encapsulated with location and confidence information to generate a disease recognition result information set.

[0096] In alternative or modified implementations, sample type determination can be combined with data collection logs, automatically labeling data from "verification boreholes" or "manual inspection sections" as "labeled and adapted samples." For labeled samples, shared feature refinement can employ contrastive learning methods, constructing positive and negative sample pairs for training. Disease category boundary correction can directly utilize Bayesian optimization to search for the optimal classifier bias parameters. For unlabeled samples, the preset threshold for generating pseudo-labels can employ an adaptive strategy, dynamically adjusting based on the average confidence of each category in each iteration. Target domain feature alignment can utilize Domain Adversarial Neural Networks (DANNs), using a domain discriminator to drive feature distribution alignment. In addition to using physical priors, category semantic compensation can also introduce temporal and spatial continuity constraints (disease categories at adjacent measurement points are usually consistent) to correct isolated erroneous pseudo-labels.

[0097] In some embodiments, based on the cross-domain disease characterization information set, initial disease category prediction is performed on the target domain samples, and target domain samples that meet the confidence threshold condition are selected as pseudo-label samples; the pseudo-label samples are input into the network with the source domain labeled and adapted samples, and the maximum mean difference algorithm is used to achieve feature distribution convergence between the target domain and the source domain; during the feature distribution convergence process, the physical consistency loss of the interlayer approximate reflection coefficient and the two-way travel time is introduced to correct and compensate for the category semantic relationship of the pseudo-label samples; the iteratively updated target domain features and category semantic relationship are jointly solidified to complete the target domain feature alignment and category semantic compensation.

[0098] A pseudo-labeled sample can refer to a sample in the "unlabeled migration sample" that is predicted as a certain disease category by the current iteration of the identification model with a confidence level higher than a preset threshold.

[0099] The maximum mean difference algorithm can be a nonparametric statistical method for measuring the difference between two probability distributions P and Q. In the reproducing kernel Hilbert space (RKHS), MMD measures the difference by calculating the distance between the means of the sample features of the two distributions. The smaller the value, the closer the distributions are.

[0100] Physical consistency loss can be a custom loss function term used to constrain model predictions or feature representations to be consistent with known physical laws.

[0101] Specifically, in pseudo-label-based self-training, the initial model's predictions of unlabeled target domain data will inevitably contain errors. If these erroneous predictions are directly used as labels for training, the errors will accumulate and amplify, eventually leading to model collapse—this is known as "confirmation bias." The necessity of this step lies in introducing a dual safety mechanism: firstly, MMD is used for overall feature distribution alignment to prevent the model from over-adapting to target domain noise; secondly, a novel physical consistency loss is introduced, utilizing the physical laws of electromagnetic reflection (r and Δt) as a "verifier" to detect and correct pseudo-labels that clearly violate physical common sense from a semantic level, thereby ensuring that the self-training process evolves stably and reliably in the correct direction.

[0102] In the specific analysis process, the system performs iterative optimization. 1) Initial prediction and screening: The current model is used to perform forward propagation on all "unlabeled migration samples" in the target domain to obtain the disease category prediction probability distribution of each sample. Screening is performed according to a preset confidence threshold condition: the preset condition is to retain only samples whose maximum predicted probability exceeds the threshold T_high (e.g., 0.9) and whose difference between the maximum probability and the second largest probability exceeds the threshold T_margin (e.g., 0.3) to ensure that the pseudo-labels have high reliability. The samples that meet the conditions and their pseudo-category labels constitute the pseudo-label sample set. 2) Feature distribution convergence based on MMD: A domain alignment network is constructed (shared coding branches can be reused or fine-tuned). The pseudo-label sample set and the labeled and adapted samples of the source domain are input into the network to extract their "intrinsic disease features". The MMD value between these two sets of features is calculated as the distribution difference loss L_mmd. By minimizing L_mmd through backpropagation, the feature distribution of the target domain is driven to move closer to the source domain. 3) Physical consistency correction and compensation: The physical consistency loss is calculated in parallel in the same training batch. For each sample in the pseudo-label sample set, based on its pseudo-class, find the statistical distribution (such as mean and variance) of r and Δt in the source domain for that class. Calculate the Mahalanobis distance (or standardized Euclidean distance) between the sample's own r and Δt values ​​and the statistical center corresponding to its pseudo-class, and use this distance as the physical consistency loss L_phys. The larger L_phys is, the more mismatched the physical parameters of the sample are with its pseudo-class, and it may be an incorrect pseudo-label. The final total loss is L_total = L_mmd + λ * L_phys, where λ is the balance coefficient. By optimizing the total loss, the model "penalizes" physically unreasonable pseudo-label samples while bringing the feature distribution closer, thereby achieving the correction and compensation of category semantic relationships. 4) Iteration and Consolidation: Repeat the above prediction-screening-alignment-compensation steps, updating the model after each iteration. When the MMD value between the target domain feature distribution and the source domain is lower than a preset threshold or the model performance tends to stabilize, stop the iteration and consolidate the target domain features extracted by the final model and the optimized category discrimination relationship.

[0103] In alternative or modified implementations, the confidence threshold condition can employ an adaptive strategy, such as dynamically adjusting T_high based on the average predicted confidence for each class. MMD computation can utilize multi-kernel MMD (MK-MMD) to better capture nonlinear distributional differences. The domain alignment network can be an independent adversarial discriminator, achieving distribution alignment through the concept of Generative Adversarial Networks (GANs). The physical consistency loss can be designed not based on class statistical centers, but rather as a physical parameter predictor, requiring the model to predict r and Δt from "disease intrinsic features," and using the error between the predicted and true values ​​(from radar data) as the loss. This allows physical constraints to act more directly on feature learning itself. The balance coefficient λ can also be dynamically adjusted based on the training epochs.

[0104] In some embodiments, based on the cross-domain disease characterization information set, the source domain samples and the labeled and adapted samples are aggregated by category center to construct shared feature prototypes corresponding to different disease categories; combined with disease physical constraint analysis information, the distribution differences of each disease category in interlayer approximate reflection coefficient, two-way travel time and frequency band energy ratio are analyzed to form physical prior boundaries; based on the shared feature prototypes and physical prior boundaries, the discrimination boundary between adjacent disease categories is shrunk or expanded to reduce category aliasing under cross-frequency band conditions; the category discrimination result after boundary correction is output as the supervised identification result.

[0105] A shared feature prototype can be a vector in the feature space that represents the central location of all samples (including the source and target domains) of a certain disease category.

[0106] The physical prior boundary can be an empirical range or threshold determined by historical data statistics or physical models in the space formed by physical parameters (such as the interlayer approximate reflection coefficient r and the two-way travel time Δt) to distinguish different disease categories.

[0107] The discrimination boundary can refer to a hyperplane or complex surface used to separate different disease categories in the feature space or decision space constructed by the classification model (such as a classifier).

[0108] Specifically, traditional model fine-tuning using a small number of labeled samples in the target domain typically involves overall adjustments across the entire feature space or parameter space, lacking specificity and prone to local overfitting or underfitting due to limited data. The necessity of this step lies in providing a structured and interpretable correction strategy: stabilizing feature centers by constructing a "category prototype" that integrates information from both domains; constraining decision boundaries by defining a "theoretically feasible domain" using physical parameters; and finally, precisely and appropriately shrinking or expanding only the category boundaries prone to confusion, thereby minimizing the risk of category confusion across frequency bands with minimal adjustment cost.

[0109] In the specific analysis process, the system executes the following steps: 1) Constructing a shared feature prototype: For each disease category, all samples of that category in the source domain and the labeled and adapted samples of that category in the target domain are collected. These are input into the shared coding branch to extract the intrinsic features of the disease. Then, the mean vector of these feature vectors is calculated as the shared feature prototype for that disease category. This prototype integrates the consensus features of the historical (source domain) and current (target domain). 2) Forming physical prior boundaries: For the same batch of samples, the corresponding r and Δt values ​​are extracted. For each disease category, the statistical range of its r and Δt is calculated (e.g., mean ± 2 standard deviations). The statistical ranges of each category are plotted on the r-Δt two-dimensional plane, and their overlapping or adjacent areas constitute the physical prior boundary regions that need to be focused on. 3) Discrimination boundary correction: In the decision layer of the classifier (usually the last fully connected layer), the category pairs that need correction are analyzed. For example, the category pairs "interlayer gap" and "poor interlayer contact" that overlap in the physical prior boundary region are identified. The correction method adds a prototype margin loss to the objective function of model training: for samples from the two classes, not only are their features required to be correctly classified, but the distance from their feature vectors to their own class prototypes is also required to be less than the distance to the other class prototype minus a preset boundary margin m. By optimizing this loss, the model is driven to "increase" the distance between the two easily confused class prototypes in the feature space, which is equivalent to expanding the discriminative boundary between them. For class pairs with high physical parameter discriminative power and low likelihood of confusion, no adjustment is needed or the boundary can be moderately shrunk to enhance generalization. Finally, combining prototype learning and boundary margin constraints, the corrected supervised recognition result is output.

[0110] In alternative or modified implementations, the computation of shared feature prototypes can avoid using simple means and instead employ iteratively updated centroids or prototype vectors learnable during training. Physical prior boundaries can be learned directly in the r-Δt space using a Support Vector Machine (SVM) instead of based on statistical ranges, or by fitting the distribution of each class using a Gaussian Mixture Model (GMM). Discriminant boundary correction can be achieved without modifying the loss function, by directly optimizing the classification layer weights: fixing the feature extraction network and fine-tuning only the weights and biases of the classification layer; calculating the inter-class covariance matrix using labeled samples from the target domain; and adjusting the direction and position of the decision hyperplane accordingly to maximize the class margin in the target domain. The boundary margin margin m can be set as an adaptive parameter, dynamically adjusted according to the degree of class confusion.

[0111] In some embodiments, based on the recognition confidence of longitudinal and horizontal perspectives under different frequency bands, the multi-view results in the disease identification result information set are adaptively weighted and fused to generate a joint identification result of the target disease; based on the joint identification result, the longitudinal range, lateral range, and interlayer depth of the disease are jointly located to generate a disease spatial distribution image; simultaneously, based on the interlayer approximate reflection coefficient, two-way travel time, and category confidence, the severity level of the disease is calculated, and a corresponding disease warning mark is formed; integrating the joint identification result, the disease spatial distribution image, and the disease warning mark, a multi-frequency band domain adaptive disease identification report is generated.

[0112] The joint identification result can refer to the unique and final judgment result on the disease category, location and extent formed by integrating information from multiple independent identification sources (such as different frequency bands and different perspectives).

[0113] Disease severity level refers to the classification and location of diseases, and the evaluation of the harm or urgency of diseases based on quantitative indicators (such as reflection coefficient, two-way travel time, spatial range, and identification confidence) to support maintenance decisions.

[0114] Specifically, traditional methods typically use fixed-weighted averaging or simple voting to assess identification results from different perspectives or frequency bands. However, under multi-frequency conditions, the significance and confidence level of the same disease can vary drastically with different perspectives (e.g., the longitudinal profile is clearer at low frequencies, while the horizontal plane details are richer at high frequencies). Static fusion cannot adapt to this dynamic change, easily leading to the "dilution" of low-confidence perspectives or interference with the final judgment. The necessity of this step lies in ensuring that the final report not only provides a unified identification conclusion through confidence-driven adaptive weighted fusion and joint analysis of physical parameters and spatial information, but also accurately depicts the spatial morphology of the disease and quantifies its severity. This achieves a leap from "presence or absence of disease" to "comprehensive assessment of disease status," directly serving graded maintenance.

[0115] In the specific analysis process, the system receives a set of disease identification results. Each target disease in this set may contain multiple identification results (including category, location, and confidence level) from different frequency bands, longitudinal section views, and horizontal viewpoints. 1) Generate joint identification results: For the same target disease, the identification results from all views and frequency bands are summarized. A confidence-based weighted voting method is used for fusion: the category confidence level of each identification result is used as its weight, and a weighted vote is calculated for each possible disease category. The category with the highest weighted vote is determined as the final disease category in the joint identification results. 2) Generate disease spatial distribution image: Using the disease categories in the joint identification results, the location information of the disease under different views is traced back. For the longitudinal section view, the start and end scan lines (horizontal X) and depth range (vertical Z) of the disease are extracted; for the horizontal slice view, the contour coordinates (X, Y) of the disease on the horizontal plane are extracted. By combining coordinate transformation and layer calibration information (from preprocessing), these two sets of two-dimensional information are fused, and interpolation and contour fitting are performed in three-dimensional space to generate a three-dimensional spatial distribution image of the disease, which is displayed in the form of isosurfaces or voxels. 3) Calculate the severity level of the disease and form early warning markers: Establish a severity assessment model. The model takes multiple indicators as input: a. The absolute value of the interlayer approximate reflection coefficient r (reflecting the dielectric difference intensity); b. The abnormal amplitude of the two-way travel time Δt of the disease area (reflecting the abnormal thickness of the structure); c. The equivalent volume of the disease in three-dimensional space (from the spatial distribution image); d. The final fusion confidence of the joint identification results. The model scores each indicator and calculates the weighted sum using a preset scoring rule library based on historical expert experience or regression analysis to obtain a total score. According to the preset threshold range of the total score, it is mapped to the corresponding "slight", "moderate" or "severe" level, and automatically attached with the corresponding color (such as green, yellow, red) or text early warning marker. Finally, by integrating the above three outputs, a structured multi-band domain adaptive disease identification report is generated.

[0116] In alternative or modified implementations, adaptive weighted fusion can employ Dempster-Shafer (DS) evidence theory, treating each identification result as a piece of evidence for uncertainty reasoning and synthesis. The generation of spatial distribution images of the disease can utilize 3D Kriging interpolation or a deep learning image generation network, directly synthesizing 3D volume data from multi-view contours. The severity assessment model can be replaced with a trained Gradient Boosting Decision Tree (GBDT) or a small neural network, directly mapping end-to-end from raw features (r, Δt, volume, confidence) to severity levels. The scoring rule base can be constructed and weighted using the Analytic Hierarchy Process (AHP) combined with expert scoring. Early warning markers can be linked to a specific maintenance measure recommendation library.

[0117] Figure 3This is a schematic diagram of the structure of a domain adaptive disease identification system based on multi-band ground-penetrating radar provided in an embodiment of this application, as shown below. Figure 3 As shown, the domain adaptive disease identification system 300 based on multi-band ground penetrating radar in this embodiment includes: a cross-domain disease characterization module 301, a disease identification module 302, and a result output module 303.

[0118] The cross-domain disease characterization module 301 is used to acquire a multi-band ground-penetrating radar detection dataset, and based on the multi-band ground-penetrating radar detection dataset, to perform cross-band response difference deconstruction, disease physical constraint extraction and multi-view characterization construction to generate a cross-domain disease characterization information set. The disease identification module 302 is used to perform source domain-target domain collaborative-driven shared feature alignment and disease semantic discrimination based on the cross-domain disease representation information set, and generate a disease identification result information set. The result output module 303 is used to perform location output and severity characterization of road surface defects based on the defect identification result information set, and generate and output a multi-frequency domain adaptive defect identification report.

[0119] The system in this embodiment can be used to execute the methods of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

Claims

1. A domain-adaptive disease identification method based on multi-band ground-penetrating radar, characterized in that, include: A multi-band ground-penetrating radar detection dataset is obtained. Based on the multi-band ground-penetrating radar detection dataset, cross-band response difference deconstruction, physical constraint extraction of diseases and multi-view characterization are performed to generate a cross-domain disease characterization information set. Based on the cross-domain disease representation information set, source domain-target domain collaboratively driven shared feature alignment and disease semantic discrimination are performed to generate a disease identification result information set; Based on the information set of the disease identification results, the system performs location output and severity characterization of road surface disease targets, and generates and outputs a multi-frequency domain adaptive disease identification report.

2. The method according to claim 1, characterized in that, The process of generating the cross-domain disease characterization information set includes: The multi-band ground-penetrating radar detection dataset includes source domain detection data and target domain detection data collected under different center frequencies and different bandwidth modes; Based on the multi-band ground-penetrating radar detection dataset, radar image characterization information is generated through echo preprocessing and multi-view reconstruction. Based on the multi-band ground-penetrating radar detection dataset, physical constraint analysis information of the disease is generated through interlayer interface echo analysis and time-frequency response deconstruction. By combining the radar image characterization information and the physical constraint analysis information of the disease, and combining the domain attribute difference analysis, cross-frequency band shared feature extraction and domain offset characterization decoupling are performed to generate the cross-domain disease characterization information set.

3. The method according to claim 2, characterized in that, The process of generating the radar image characterization information includes: Based on the source domain detection data and the target domain detection data, zero correction, DC removal, background removal, gain compensation and bandpass filtering are performed to obtain the initial purified echo data. Based on the initial purified echo data, A-scan echo sequences, B-scan longitudinal section images, and horizontal slice images are constructed as the basic structural representations for disease identification. The basic structure characterization is subjected to amplitude normalization, layer tracking and anomalous echo enhancement processing to generate a set of comparable multi-view radar image features; The radar image characterization information is generated by summarizing the A-scan echo sequence, the B-scan longitudinal section image, the horizontal slice image, and the multi-view radar image feature set to characterize the significant differences in disease severity across different frequency bands.

4. The method according to claim 3, characterized in that, The process of generating the physical constraint analysis information of the disease includes: Based on the interlayer interface reflection waveforms in the source domain detection data and the target domain detection data, the interface peak amplitude, peak polarity and adjacent peak spacing are extracted as basic waveform clues of the interface response. Simultaneously, the interlayer interface reflection waveform is subjected to two-way travel time measurement, main frequency offset analysis and frequency band energy distribution analysis to identify the corresponding propagation path differences and frequency response differences. Based on the interface peak amplitude, the wave crest polarity, the two-way travel time, and the frequency band energy distribution, the interlayer approximate reflection coefficient is calculated and a set of physical constraint parameters characterizing the interlayer contact state is constructed. The basic waveform clues are associated with and mapped to the physical constraint parameter group, and consistency analysis is performed to generate the physical constraint analysis information of the disease.

5. The method according to claim 4, characterized in that, The decoupling of cross-band shared feature extraction and domain offset representation includes: The multi-view features in the radar image representation information are input into the shared coding branch and the domain difference coding branch, and the intrinsic features of the disease and the frequency band difference features are extracted respectively through parallel mapping. The interlayer approximate reflection coefficient, the two-way travel time, and the frequency band energy distribution in the physical constraint analysis information of the disease are introduced into the shared coding branch to impose physical consistency constraints on the intrinsic characteristics of the disease. The distribution shift of frequency band difference characteristics under different frequency bands, different device responses, and different road surface structures is analyzed to generate a domain difference semantic description. Based on the intrinsic features of the disease and the domain-differential semantic description, a cross-domain disease representation information set is constructed, which includes shared disease semantic information and domain-offset semantic information.

6. The method according to claim 5, characterized in that, The process of generating the disease identification result information set includes: Based on the cross-domain disease characterization information set, determine whether the current target domain sample is a labeled and adapted sample or an unlabeled migration sample. When the target domain sample is the labeled and adapted sample, the shared feature refinement and disease category boundary correction based on category supervision are performed to generate supervised identification results; When the target domain sample is the unlabeled transfer sample, target domain feature alignment and category semantic compensation based on pseudo-label iteration are performed to generate transfer recognition results. The generated supervised identification result or the migration identification result is associated and encapsulated with the corresponding disease location marker and confidence marker to generate the disease identification result information set.

7. The method according to claim 6, characterized in that, The target domain feature alignment and category semantic compensation based on pseudo-label iteration includes: Based on the cross-domain disease characterization information set, an initial disease category prediction is performed on the target domain samples, and target domain samples that meet the confidence threshold condition are selected as pseudo-label samples. The pseudo-labeled samples are aligned with the labeled and adapted samples in the source domain and the network is then aligned. The maximum mean difference algorithm is used to achieve convergence of the feature distribution between the target domain and the source domain. During the convergence process of the feature distribution, the physical consistency loss of the interlayer approximate reflection coefficient and the two-way travel time is introduced to correct and compensate for the category semantic relationship of the pseudo-label samples. The updated target domain features are jointly solidified with the category semantic relationship to complete the target domain feature alignment and category semantic compensation.

8. The method according to claim 6, characterized in that, The process of refining shared features and correcting disease category boundaries based on category supervision includes: Based on the cross-domain disease characterization information set, the source domain samples and the labeled and adapted samples are aggregated by category center to construct a shared feature prototype corresponding to different disease categories. Based on the physical constraint analysis information of the disease, the distribution differences of each disease category in terms of interlayer approximate reflection coefficient, two-way travel time and frequency band energy ratio are analyzed to form physical prior boundaries; Based on the shared feature prototype and the physical prior boundary, the discrimination boundary between adjacent disease categories is shrunk or expanded to reduce category aliasing under cross-frequency band conditions. The category discrimination result after boundary correction is output as the supervised recognition result.

9. The method according to claim 6, characterized in that, The generation process of the multi-band domain adaptive disease identification report includes: Based on the recognition confidence of longitudinal and horizontal perspectives under different frequency bands, the multi-perspective results in the disease identification result information set are adaptively weighted and fused to generate joint identification results of the target disease. Based on the joint identification results, the longitudinal range, lateral range and interlayer depth of the disease are jointly located to generate a spatial distribution image of the disease. Simultaneously, based on the interlayer approximate reflection coefficient, the two-way travel time, and the category confidence level, the severity level of the disease is calculated, and a corresponding disease early warning mark is generated; The joint identification results, the spatial distribution image of the disease, and the disease early warning markers are integrated to generate the multi-frequency domain adaptive disease identification report.

10. A domain-adaptive disease identification system based on multi-band ground-penetrating radar, characterized in that, The method applied to any one of claims 1-9 includes: The cross-domain disease characterization module is used to acquire a multi-band ground-penetrating radar detection dataset, and based on the multi-band ground-penetrating radar detection dataset, to perform cross-band response difference deconstruction, disease physical constraint extraction and multi-view characterization construction to generate a cross-domain disease characterization information set. The disease identification module is used to perform source domain-target domain collaborative-driven shared feature alignment and disease semantic discrimination based on the cross-domain disease representation information set, and generate a disease identification result information set; The result output module is used to perform location output and severity characterization of road surface defects based on the defect identification result information set, and generate and output a multi-frequency domain adaptive defect identification report.