Multi-mode large model ground penetrating radar disease data intelligent interpretation method and device
Through three fine-tunings of the multimodal large model, combined with ground-penetrating radar physics knowledge and industry standards, the efficiency and consistency issues of ground-penetrating radar data interpretation were resolved, achieving high-precision automatic interpretation of defects and report generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGAN UNIV
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-08
AI Technical Summary
Existing ground-penetrating radar data interpretation methods rely on manual interpretation, which is inefficient and yields inconsistent results. They also lack multimodal information interaction and industry knowledge guidance, making it difficult to jointly infer the location, outline, and physical properties of damage.
A multimodal large model is constructed, and through three fine-tuning processes—target detection, semantic segmentation, and image description—combined with ground-penetrating radar physics knowledge and industry standards, structured descriptive text is generated.
It achieves high-precision and interpretable automatic interpretation, generating comprehensive interpretation reports that include disease type, spatial location, and physical characteristics, thereby improving interpretation efficiency and reliability.
Smart Images

Figure CN121999375A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer intelligent computing and information processing technology, specifically relating to a method and device for intelligent interpretation of multimodal large-model ground-penetrating radar defect data. Background Technology
[0002] With the continuous expansion of my country's transportation infrastructure, the early identification and precise treatment of underground road defects have become crucial for ensuring public safety and extending the lifespan of infrastructure. Ground-penetrating radar (GPR), as a rapid and non-destructive underground detection technology, can obtain profile images of underground media by analyzing radar wave reflection characteristics and is widely used in road defect detection, municipal engineering surveys, and other fields. However, the current interpretation of GPR data in the industry mainly relies on manual interpretation, which is inefficient and the results are greatly affected by personal experience, making it difficult to guarantee consistency and reproducibility.
[0003] In recent years, intelligent computing technologies, represented by machine learning and artificial intelligence, have provided new pathways for automated image interpretation. Existing research mostly employs convolutional neural network-based target detection or semantic segmentation models to achieve preliminary localization and contour recognition of defects. These methods are essentially single-task applications of computing devices based on specific mathematical models in image processing, focusing on the extraction and classification of low-level features in radar images, and have not yet formed a collaborative intelligent system integrating multi-level information. Although they can replace some manual operations to a certain extent, their model architecture is singular, tasks are isolated, and they lack knowledge guidance and multimodal information interaction capabilities.
[0004] Furthermore, current methods are mostly limited to learning the mapping from data to labels, failing to effectively incorporate prior industry knowledge and physical constraints, and also failing to construct a complete interpretation chain of "perception-reasoning-description" that conforms to human cognitive logic. Existing technologies struggle to achieve joint inference between lesion location, contour morphology, and physical property parameters, and are unable to automatically generate structured descriptive text that conforms to industry standards, thus limiting their reliability and practicality in complex scenarios. Summary of the Invention
[0005] To address the aforementioned problems in existing technologies, this invention provides a method and apparatus for intelligent interpretation of ground-penetrating radar (GPR) defect data using a multimodal large-scale model. By constructing a multi-task collaborative model integrating visual perception, knowledge reasoning, and text generation, it not only utilizes machine learning and artificial intelligence techniques for feature learning and pattern recognition but also introduces a model based on GPR physics knowledge to encode and reason about the physical mechanisms of defects and industry standards. This enables high-precision, interpretable, and automatic interpretation from radar images to structured descriptive information based on a specific computational model. This promotes the deep integration and systematic application of intelligent computing technology in professional detection fields.
[0006] The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides an intelligent interpretation method for ground-penetrating radar defect data using a multimodal large model, comprising: Step 1: Using ground-penetrating radars with different center frequencies, differentiated data collection is carried out for different types of underground diseases to obtain raw radar data; Step 2: After preprocessing the original radar data, crop it to a preset size and label the disease targets in the image to construct a target detection dataset including multiple image data. Step 3: Convert the image data in the target detection dataset into a serialized format adapted to a multimodal large model; Step 4: Perform the first fine-tuning of the multimodal large model using the object detection dataset; Step 5: Based on the target detection results of the target detection dataset of the multimodal large model after the first fine-tuning, the image data is binarized using the adaptive threshold optimization threshold segmentation technique to generate a binarized label mask for the semantic segmentation task. Step 6: Convert the image data and the corresponding binarized label mask into a serialized format, construct a semantic segmentation dataset, and use the semantic segmentation dataset to perform a second fine-tuning of the multimodal large model; Step 7: Based on the semantic segmentation results of the semantic segmentation dataset of the multimodal large model after the second fine-tuning, perform text description annotation on the disease targets, construct an image description dataset, and use the image description dataset to perform a third fine-tuning on the multimodal large model; Step 8: Utilize the multimodal large model after three fine-tuning steps to achieve intelligent interpretation and analysis of underground diseases and generate a comprehensive interpretation report.
[0007] This invention also provides an intelligent interpretation device for multimodal large-scale ground-penetrating radar (GPR) defect data, applicable to the intelligent interpretation method for multimodal large-scale GPR defect data described in any of the above embodiments, comprising: The data acquisition module utilizes ground-penetrating radars with different center frequencies to collect differentiated data for different types of underground diseases and obtain raw radar data. The data preprocessing module is used to preprocess the original radar data and crop it into images of a preset size, and to annotate the disease targets in the images to construct a target detection dataset including multiple image data; and to convert the image data in the target detection dataset into data in a serialized format adapted to a multimodal large model. The model fine-tuning module performs a first fine-tuning of the multimodal large model using the target detection dataset. Based on the target detection results of the multimodal large model after the first fine-tuning, it performs binarization processing on the image data using an adaptive threshold optimization threshold segmentation technique to generate a binary label mask for the semantic segmentation task. The image data and the corresponding binary label mask are converted into a serialized format to construct a semantic segmentation dataset, which is then used to perform a second fine-tuning of the multimodal large model. Based on the semantic segmentation results of the multimodal large model after the second fine-tuning, text descriptions are annotated for the disease targets to construct an image description dataset, which is then used to perform a third fine-tuning of the multimodal large model. The interpretation and analysis module utilizes a multimodal large model that has undergone three fine-tuning steps to achieve intelligent interpretation and analysis of underground diseases and generate a comprehensive interpretation report.
[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: The intelligent interpretation method for ground-penetrating radar (GPR) defect data using a multimodal large-scale model of this invention constructs a multi-task collaborative learning framework. By sequentially performing target detection, semantic segmentation, and image description fine-tuning on the same multimodal large-scale model, the model can simultaneously output the location of defects, pixel-level contours, and industry-standard language descriptions. This achieves layer-by-layer fusion and expression from low-level image features to high-level semantic information, solving the problem of information fragmentation in existing single-task models. The multimodal large-scale model after these three fine-tuning steps can automatically generate a comprehensive interpretation report containing defect categories, spatial locations, geometric features, and physical features. This can assist professionals in efficiently and accurately analyzing underground defect data, significantly improving interpretation efficiency and reliability, and providing objective basis for road maintenance decisions.
[0009] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of an intelligent interpretation method for multimodal large-scale ground-penetrating radar defect data provided in an embodiment of the present invention; Figure 2 This is a flowchart of an intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the fine-tuning PaliGemma model provided in an embodiment of the present invention; Figure 4This is a flowchart of key parameters for Bayesian optimization of threshold segmentation provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the comprehensive interpretation report provided in an embodiment of the present invention. Detailed Implementation
[0011] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the following describes in detail, with reference to the accompanying drawings and specific embodiments, a method and apparatus for intelligent interpretation of multimodal large-scale ground-penetrating radar defect data proposed according to the present invention.
[0012] The foregoing and other technical contents, features, and effects of the present invention will be clearly presented in the following detailed description of specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and concrete understanding can be gained of the technical means and effects adopted by the present invention to achieve its intended purpose. However, the accompanying drawings are for reference and illustration only and are not intended to limit the technical solutions of the present invention.
[0013] Firstly, embodiments of the present invention provide an intelligent interpretation method for ground-penetrating radar (GPR) defect data based on a multimodal large model. This method enables collaborative analysis and intelligent interpretation of GPR data based on a multimodal large model, ultimately generating a structured and highly readable comprehensive interpretation report of underground defects, providing complete technical support for accurate identification and scientific maintenance decisions regarding road defects.
[0014] Please refer to the above. Figure 1 and Figure 2 , Figure 1 This is a schematic diagram of an intelligent interpretation method for multimodal large-scale ground-penetrating radar defect data provided in an embodiment of the present invention; Figure 2 This is a flowchart of a method for intelligent interpretation of multimodal large-scale ground-penetrating radar defect data provided in an embodiment of the present invention.
[0015] The intelligent interpretation method for multimodal large-scale ground-penetrating radar defect data in this embodiment may include the following steps: Step 1: Using ground-penetrating radars with different center frequencies, differentiated data collection is carried out for different types of underground diseases to obtain raw radar data.
[0016] In this embodiment, the ground-penetrating radars with different center frequencies include: a ground-penetrating radar with a center frequency of 400MHz and a ground-penetrating radar with a center frequency of 900MHz. Specifically, for cavitation-type defects, the radar with a center frequency of 400MHz is used because its low-frequency characteristics allow for detection over a deeper area; for crack-type defects, the radar with a center frequency of 900MHz is used because its high-frequency characteristics provide higher resolution.
[0017] For example, during the measurement process, for road scenarios, the dielectric constant is generally set to 9, and the radar data acquisition mode is a continuous measurement model with a benchmark to mark suspected areas. The number of sampling points for a single-channel A-scan is 8192, and the sampling rate is set to 250KHz.
[0018] It is understandable that GPS (Global Positioning System) coordinates and road surface images are recorded simultaneously during the data collection process to provide spatial information for subsequent report generation.
[0019] Step 2: After preprocessing the original radar data, crop it to a preset size and label the disease targets in the images to construct a target detection dataset that includes multiple image data.
[0020] In an optional embodiment, step 2 includes: Step 2.1: Perform direct wave removal, bandwidth filtering, signal gain adjustment, and background filtering on the raw radar data in sequence.
[0021] In this embodiment, direct wave removal can be achieved using a time-threshold method, directly truncating signals within a 0ns-1.5ns time window to eliminate strong direct wave and surface-coupled wave interference between the radar antenna and the ground. Bandpass filtering, since ground-penetrating radar is a broadband antenna, needs to be determined based on the signal distribution at the center frequency of the radar antenna. Signal gain is typically used to compensate for the energy attenuation of electromagnetic waves with increasing depth, employing automatic gain control (AGC) or an exponential gain function. A gain of 10dB to 30dB is dynamically applied based on the estimated target depth range to enhance weakly reflected signals at depth. Background filtering uses an average background cancellation method. The average value of all A-scan signals across the entire B-scan profile is calculated as a "background noise model," and then this background model is subtracted from each A-scan signal, thus suppressing clutter from the in-phase axis noise system.
[0022] Step 2.2: Crop the processed radar data into an image of a preset size.
[0023] Since the B-scan signal of ground penetrating radar is usually composed of multiple A-scans, it is usually a rectangular image. In order for deep learning models to better understand the features of the image, it is usually necessary to crop it into an image of the same size, for example, to an image with a pixel size of 512×512 or 640×640.
[0024] Step 2.3: Use LabelImg (image annotation tool) software to annotate the disease targets in the image and construct the target detection dataset.
[0025] In this embodiment, the annotation categories include "holes" and "cracks". It should be noted that for disputed targets, verification can be made by combining drilling verification and on-site inspection with industrial endoscopy, and finally, a target detection dataset conforming to the VOC format (Visual Object Classes, a file format for target detection and image annotation) can be constructed.
[0026] Step 3: Convert the image data in the object detection dataset into a serialized format that is compatible with multimodal large models.
[0027] In this embodiment, the multimodal large model uses PaliGemma's vision-language model, whose input requirement is JSON Lines (Json1) format. The image data needs to be converted to JSON1 format. In this embodiment, each JSON1 data entry includes three key fields: image filename, task prefix, and task output content. "image" represents the image filename, "prefix" represents the task prefix, and "suffix" represents the task output content. The task prefix includes: an object detection task prefix, represented as "...". <det>", a prefix for semantic segmentation tasks, denoted as " <segment>", the image description task prefix, is represented as " <cap>".
[0028] In this embodiment, during data format conversion, the target detection task follows a normalized coordinate format, converting VOC format annotations to a normalized coordinate format in the order [Y_min, X_min, Y_max, X_max], and uniformly normalizing them to a scale of 1024×1024. Other sizes require corresponding conversions. In the semantic segmentation task, the target region is defined by a bounding box, with pixel values within the box ranging from 0 to 128. By uniformly constructing annotation data suitable for different tasks, a multimodal dedicated dataset for interpreting ground-penetrating radar defects is formed.
[0029] For example, a JSON record of object detection data is: {"image": "img_001.jpg", "prefix": " <det>", "suffix": "void 0.12 0.34 0.45 0.67"}, where void represents the void category.
[0030] Step 4: Perform the first fine-tuning of the multimodal large model using the object detection dataset.
[0031] In this embodiment, after the first fine-tuning, the multimodal large model obtains model weights with target detection capabilities.
[0032] Specifically, the JSON data generated in step 3 is loaded into the PaliGemma model. PaliGemma, as a vision-language multimodal model, divides its input processing into two paths: visual and text. In the visual encoder path, the image is resized to a 224×224×1 grayscale image and divided into 14×14 image patches, each patch being 16×16 in size. Each image patch is converted into a D-dimensional embedding vector through a linear projection layer. After adding positional encoding, the input is processed by a Vision Transformer (ViT), which consists of multiple stacked Transformer encoder layers. This visual encoder extracts global contextual features of the image through a self-attention mechanism and finally converts them into visual tokens (minimum processing units). In the text path, the input task prefix, such as "...", is used. <det>"The visual tokens are converted into text tokens by a multilingual tokenizer. These text tokens are then converted into text embedding vectors via a word embedding layer. After positional encoding, they are input into the embedding layer of the language model decoder. During feature alignment, the visual tokens are fed into a lightweight perceptron-resampler. This module primarily uses a set of learnable query vectors to perform cross-attention calculations on long visual sequences, compressing and converting them into a fixed number of tokens of length, achieving spatial alignment with the tokens in the text feature space. During multimodal decoding, the visual tokens and text tokens are concatenated to form a unified multimodal sequence token, which is then uniformly input into the PaliGemma multimodal decoder to decode the token output. The model predicts the output sequence in an autoregressive manner."
[0033] In this embodiment, the first fine-tuning employs a parameter-efficient fine-tuning strategy: during model fine-tuning, all parameters of the visual encoder of the multimodal large model are frozen, and the attention mechanism-related layers in the decoder of the multimodal large model are fine-tuned. These attention mechanism-related layers include the attention output projection layer (attn_vec_einsum / w), the key-value joint projection layer (kv_einsum / w), and the query projection layer (q_einsum / w). The total number of parameters in these layers is 169.9M, accounting for 5.8% of the total PaliGemma parameters. This allows the model to quickly adapt to ground-penetrating radar missions while perfectly preserving the general visual features obtained during the fine-tuning phase.
[0034] In terms of precision settings, to balance performance and memory efficiency, trainable layers use float32 precision, while frozen layers maintain float16 precision. The actual number of updated parameters accounts for approximately 12% of the entire model. This method significantly reduces the demand for computing resources. At the same time, it utilizes the JAX framework to accelerate and optimize the fine-tuning process. By combining the JAX framework's features of just-in-time compilation and automatic differentiation calculation, the process of calculating the graph matrix is optimized, enabling rapid model iteration.
[0035] During the model fine-tuning stage, the loss function is the cross-entropy loss, and the specific loss function is shown in equation (1).
[0036] (1); in, For loss function, For batch size, For the first The sequence length of each sample. For the first Each sample is located at Real word units, For position All previous morphemes, For model parameters The predicted probability distribution at that time For the input context. During the model fine-tuning phase, the optimizer employs stochastic gradient descent (SGD) in conjunction with a cosine degradation strategy to dynamically adjust the learning rate. This method effectively guides the model parameters to converge to the global optimum. Specifically, the learning rate is dynamically set to smoothly decay to zero based on the initial value. That is, the learning rate is dynamically adjusted from the initial value of 0.0005 to 0 as the model loss changes, until the model converges. The SGD optimizer parameter update formula is shown in equation (2).
[0037] (2); In the formula, It is the first Step model parameters, It is the gradient of the loss function with respect to the parameters. It is the first The learning rate of each step.
[0038] The cosine annealing learning rate formula is shown in equation (3): (3); In the formula, For the current step (the first step) The learning rate of (steps). This is the upper limit of the learning rate, usually set as the initial learning rate for fine-tuning. This is the lower limit of the learning rate, usually set to 0 or a very small value. This represents the number of training steps or cycles currently executed. This refers to the predetermined total number of training steps or total number of cycles.
[0039] In this embodiment, an early stopping mechanism is enabled during model fine-tuning to prevent overfitting. This mechanism effectively prevents overfitting and underfitting during training. During model training, the batch size for the object detection task is 32, the initial learning rate is set to 0.0003, and the cosine annealing strategy has a 10% warm-up ratio. The training epoch is set to 100 rounds to help the model fully extract and understand image information.
[0040] Step 5: Based on the target detection results of the target detection dataset of the multimodal large model after the first fine-tuning, the image data is binarized using the adaptive threshold optimization threshold segmentation technique to generate a binarized label mask for the semantic segmentation task.
[0041] In this embodiment, an adaptive threshold optimization threshold segmentation technique is proposed. This technique can automatically optimize key threshold segmentation parameters to help the model complete threshold segmentation. Specifically, the process of Bayesian optimization of key threshold segmentation parameters is as follows: Figure 4 As shown, Figure 4 This is a flowchart of key parameters for Bayesian optimization of threshold segmentation provided in an embodiment of the present invention.
[0042] In an optional embodiment, step 5 includes: Step 5.1: Perform object detection on the image data in the object detection dataset based on the first fine-tuned multimodal large model, output the object detection bounding box corresponding to each image data, and extract the target region from the image data based on the object detection bounding box.
[0043] Step 5.2: Based on the preset objective function and Bayesian optimization termination condition, the key parameters of threshold segmentation are automatically optimized using the Bayesian optimization method.
[0044] In this embodiment, the objective function is shown in equation (4): (4); In the formula, Let be the objective function. As a key parameter, For grayscale contrast, For edge smoothness, For model connectivity, The weight for grayscale contrast. As the weight for edge smoothness, These are the weights for model connectivity.
[0045] Background noise in real GPR (Ground-Penetrating Radar) signals is complex. Higher connectivity weights help suppress clutter because they favor structurally continuous, coherent regions and filter out minor, isolated noise. Therefore, connectivity is prioritized in the evaluation of segmentation label creation to improve segmentation accuracy in noisy environments. The weight is set to 0.2. The weight is set to 0.1. The weight is set to 0.7.
[0046] The specific formula for Bayesian optimization is shown in equation (5). In this embodiment, the termination condition for Bayesian optimization is that the value of the objective function is greater than or equal to 0.85 or the number of optimization iterations is 50 rounds.
[0047] (5); In the formula, Indicates the first The new parameter point that the wheel will select. This means finding the parameter that maximizes the value of the subsequent function. Represents the search space. This represents the data acquisition function.
[0048] The expression for the Expected Improvement (EI) acquisition function is given by equation (6): (6); In equation (6) As an intermediate variable, it is represented as: (7); In the formula, and For Gaussian processes in The posterior mean and standard deviation at the given location. This is the current optimal observation value. To explore the use of equilibrium parameters, and These are the standard normal distribution function and probability density function.
[0049] Step 5.3: Based on the key parameters obtained from the optimization, generate a binary mask for the target region using an adaptive threshold segmentation method.
[0050] In this embodiment, the specific formula for adaptive threshold segmentation is shown in equation (8): (8); In the formula, For Gaussian kernel function, This represents the convolution operation. The key parameter is the Gaussian kernel function, as shown in equation (9): (9); According to the threshold Binarize the target region; create a binarization mask for the target region.
[0051] Step 5.4: Overlay the corresponding region in the image data with the binarized mask of the target region, and set the pixels of the non-target region in the image data as the background to obtain the binarized label mask.
[0052] Specifically, the binarized mask is placed back into the corresponding region of the original image data, and the pixels in the non-target region are set to 0, which is used as the background. Finally, a pixel-level semantic segmentation mask label with the same size as the original image is obtained, which is the binarized label mask, and used for subsequent model fine-tuning.
[0053] Step 6: Convert the image data and the corresponding binarized label mask into a serialized format, construct a semantic segmentation dataset, and use the semantic segmentation dataset to perform a second fine-tuning of the multimodal large model.
[0054] Specifically, the original image and the corresponding binarized mask image are converted to JSON format, and the "prefix" field in each record is set to "". <segment>The "suffix" field represents the filename (or encoded pixel sequence) of the masked image, forming a semantic segmentation dataset. This semantic segmentation data is used to fine-tune PaliGemma a second time. The second fine-tuning strategy is the same as the first, and will not be elaborated here. The batch size for the semantic segmentation task is set to 32, allowing the model to learn the mapping from images to pixel-level masks. After this second fine-tuning, the multimodal large model acquires model weights with semantic segmentation capabilities.
[0055] Step 7: Based on the semantic segmentation results of the semantic segmentation dataset of the multimodal large model after the second fine-tuning, perform text description annotation on the disease targets, construct an image description dataset, and use the image description dataset to perform a third fine-tuning of the multimodal large model.
[0056] Specifically, the multimodal large model, after a second fine-tuning, is first used to infer the images in the semantic segmentation dataset to obtain a segmentation mask. Then, radar experts use the segmentation mask and the original images, combined with the industry standard T / CHTS10160-2024, to perform textual description annotations on the diseased targets.
[0057] Textual descriptions fall into two main categories: geometric feature descriptions and physical feature descriptions. Geometric feature descriptions include the interpretation of the target's hyperbolic shape, including regular, irregular, and flat-topped hyperbolas, as well as the representation of the target's location, for example, located in the upper part of the image at a depth of approximately 1.2m. The criteria for determining the hyperbolic shape can be based on the radar imaging equations and achieved by fitting a mathematical model. The target location can be represented by a large model, such as ChatGPT-4, which converts coordinate information into natural language.
[0058] The physical characteristics description includes the signal distribution range, signal amplitude strength, signal amplitude attenuation degree, and the presence of multiples. This information is obtained by extracting the A-scan signal corresponding to the affected area and performing time-frequency domain analysis. For example, a physical characteristic description of a strong reflected signal amplitude and the presence of obvious multiples indicates that the reflection coefficient of the cavity interface is large.
[0059] Create a JSON record containing each image and its corresponding text description, with "prefix" set to "". <cap>The image description dataset is composed of "" and "suffix", which are text description strings. PaliGemma is then fine-tuned for the third time using this dataset. The tuning strategy is the same as the previous two, so it will not be elaborated here. However, because the text decoder in the image description task consumes a significant amount of GPU memory, the batch size is set to 16. After this third fine-tuning, the multimodal large model can automatically generate feature descriptions conforming to industry standards based on the input ground-penetrating radar images.
[0060] The process of three fine-tunings of the PaliGemma model in this embodiment is as follows: Figure 3 As shown, Figure 3 This is a schematic diagram of the fine-tuning PaliGemma model provided in an embodiment of the present invention.
[0061] Step 8: Utilize the multimodal large model after three fine-tuning steps to achieve intelligent interpretation and analysis of underground diseases and generate a comprehensive interpretation report.
[0062] Specifically, step 8 includes: Step 8.1 Obtain the ground-penetrating radar data to be tested.
[0063] In this embodiment, ground-penetrating radar with center frequencies of 400MHz and 900MHz can be used to collect data.
[0064] Step 8.2: After performing direct wave removal, bandwidth filtering, signal gain adjustment, and background filtering on the ground-penetrating radar data to be tested, the image is cropped to a preset size.
[0065] In this embodiment, the processed data can be cropped into an image with a pixel size of 640×640.
[0066] Step 8.3: Convert the image data obtained in Step 8.2 into a serialized format adapted to the multimodal large model.
[0067] In this embodiment, the image data obtained in step 8.2 is converted into JSON format adapted to the multimodal large model, and records are constructed for the object detection task, semantic segmentation task and image description task respectively, and output sequentially using the multi-task capability of the model.
[0068] Step 8.4: Input the serialized data into the multimodal large model after three fine-tunings for intelligent interpretation and analysis of underground diseases, and generate a comprehensive interpretation report, which includes disease category, spatial location, geometric feature description and physical feature description.
[0069] In this embodiment, the model automatically outputs target detection results (disease category and coordinates), semantic segmentation masks, and image description text. The outputs of these three tasks are integrated, and a comprehensive interpretation report is generated through information fusion and structured template filling. The report includes: disease category, spatial location (such as image coordinates and GPS coordinates), geometric feature description, physical feature description, disease distribution diagram, segmentation mask image, etc. The report supports exporting to an editable document format and integrates road surface information captured by a 3D ground-penetrating radar following a vehicle camera and disease excavation verification data, providing comprehensive support for road maintenance decisions. Please refer to [link to relevant documentation]. Figure 5 , Figure 5 This is a schematic diagram of the comprehensive interpretation report provided in an embodiment of the present invention.
[0070] The intelligent interpretation method for ground-penetrating radar (GPR) defect data using a multimodal large-scale model, as described in this invention, constructs a multi-task collaborative learning framework. By sequentially performing target detection, semantic segmentation, and image description fine-tuning on the same multimodal large-scale model, the model can simultaneously output the defect location, pixel-level contours, and industry-standard language descriptions. This achieves layer-by-layer fusion and expression from low-level image features to high-level semantic information, solving the problem of information fragmentation in existing single-task models. The multimodal large-scale model, after these three fine-tuning steps, can automatically generate a comprehensive interpretation report containing defect categories, spatial locations, geometric features, and physical features. This can assist professionals in efficiently and accurately analyzing underground defect data, significantly improving interpretation efficiency and reliability, and providing objective basis for road maintenance decisions.
[0071] The intelligent interpretation method for ground-penetrating radar defect data of multimodal large models in this invention utilizes Bayesian optimization to automatically search for key parameters for threshold segmentation during model fine-tuning. By comprehensively considering grayscale contrast, edge smoothness, and connectivity, it can generate high-quality binary masks in complex noise environments, significantly reducing the workload of manually annotating semantic segmentation data while ensuring the accuracy and consistency of labels.
[0072] Secondly, the present invention provides an intelligent interpretation device for multimodal large-scale ground-penetrating radar (GPR) defect data, applicable to the intelligent interpretation method for multimodal large-scale GPR defect data provided in the first aspect. The device comprises: The data acquisition module utilizes ground-penetrating radars with different center frequencies to collect differentiated data for different types of underground diseases and obtain raw radar data. The data preprocessing module is used to preprocess the raw radar data and crop it into images of a preset size, and to annotate the disease targets in the images to construct a target detection dataset that includes multiple image data; and to convert the image data in the target detection dataset into data in a serialized format adapted to a multimodal large model. The model fine-tuning module performs the first fine-tuning of the multimodal large model using the object detection dataset. Based on the object detection results of the multimodal large model after the first fine-tuning, it performs binarization processing on the image data using an adaptive threshold optimization threshold segmentation technique to generate a binary label mask for the semantic segmentation task. The image data and the corresponding binary label mask are converted into a serialized format to construct a semantic segmentation dataset, which is then used to perform the second fine-tuning of the multimodal large model. Based on the semantic segmentation results of the multimodal large model after the second fine-tuning, it performs text description annotation on the disease targets to construct an image description dataset, which is then used to perform the third fine-tuning of the multimodal large model. The interpretation and analysis module utilizes a multimodal large model that has undergone three fine-tuning steps to achieve intelligent interpretation and analysis of underground diseases and generate a comprehensive interpretation report.
[0073] For details regarding the intelligent interpretation device for ground-penetrating radar defect data of the multimodal large model and its corresponding beneficial effects, please refer to the relevant content of the intelligent interpretation method for ground-penetrating radar defect data of the multimodal large model provided in the first aspect, which will not be repeated here.
[0074] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that an article or apparatus comprising a list of elements includes not only those elements but also other elements not expressly listed. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or apparatus that includes said element. Terms such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect. The orientations or positional relationships indicated by terms such as "upper," "lower," "left," and "right" are based on the orientations or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0075] In the description of this specification, the references to "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0076] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.< / cap> < / segment> < / det> < / det> < / cap> < / segment> < / det>
Claims
1. A method for intelligent interpretation of ground-penetrating radar defect data in a multimodal large model, characterized in that, include: Step 1: Using ground-penetrating radars with different center frequencies, differentiated data collection is carried out for different types of underground diseases to obtain raw radar data; Step 2: After preprocessing the original radar data, crop it to a preset size and label the disease targets in the image to construct a target detection dataset including multiple image data. Step 3: Convert the image data in the target detection dataset into a serialized format adapted to a multimodal large model; Step 4: Perform the first fine-tuning of the multimodal large model using the object detection dataset; Step 5: Based on the target detection results of the target detection dataset of the multimodal large model after the first fine-tuning, the image data is binarized using the adaptive threshold optimization threshold segmentation technique to generate a binarized label mask for the semantic segmentation task. Step 6: Convert the image data and the corresponding binarized label mask into a serialized format, construct a semantic segmentation dataset, and use the semantic segmentation dataset to perform a second fine-tuning of the multimodal large model; Step 7: Based on the semantic segmentation results of the semantic segmentation dataset of the multimodal large model after the second fine-tuning, perform text description annotation on the disease targets, construct an image description dataset, and use the image description dataset to perform a third fine-tuning on the multimodal large model; Step 8: Utilize the multimodal large model after three fine-tuning steps to achieve intelligent interpretation and analysis of underground diseases and generate a comprehensive interpretation report.
2. The intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model according to claim 1, characterized in that, The various ground-penetrating radars with different center frequencies include: a ground-penetrating radar with a center frequency of 400MHz and a ground-penetrating radar with a center frequency of 900MHz; wherein, the ground-penetrating radar with a center frequency of 400MHz scans for cavitation-type defects, and the ground-penetrating radar with a center frequency of 900MHz scans for crack-type defects.
3. The intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model according to claim 1, characterized in that, Step 2 includes: Step 2.1: Perform direct wave removal, bandwidth filtering, signal gain adjustment, and background filtering on the raw radar data in sequence; Step 2.2: Crop the processed radar data into an image of a preset size; Step 2.3: Use LabelImg software to label the disease targets in the image and construct the target detection dataset.
4. The intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model according to claim 1, characterized in that, In step 3, the serialization format is JSON format, which includes three key fields: image file name, task prefix, and task output content. Here, "image" represents the image file name, "prefix" represents the task prefix, and "suffix" represents the task output content. The task prefix includes: the object detection task prefix, represented as " <det>", a prefix for semantic segmentation tasks, denoted as " <segment>", the image description task prefix, is represented as " <cap> "。< / cap> < / segment> < / det> 5. The intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model according to claim 1, characterized in that, The first, second, and third fine-tuning of the multimodal large model all employ a parameter-efficient fine-tuning strategy, which includes: During model fine-tuning, all parameters of the visual encoder of the multimodal large model are frozen, and the attention mechanism-related layers in the decoder of the multimodal large model are fine-tuned. The attention mechanism-related layers include an attention output projection layer, a key-value joint projection layer, and a query projection layer. The trainable layers use float32 precision, while the frozen layers maintain float16 precision. The JAX framework is used to accelerate and optimize the fine-tuning process.
6. The intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model according to claim 1, characterized in that, Step 5 includes: Step 5.1: Perform target detection on the image data in the target detection dataset based on the first fine-tuned multimodal large model, output the target detection bounding box corresponding to each image data, and extract the target region from the image data based on the target detection bounding box; Step 5.2: Based on the preset objective function and Bayesian optimization termination condition, the key parameters of threshold segmentation are automatically optimized using the Bayesian optimization method; Step 5.3: Based on the key parameters obtained through optimization, generate a binary mask for the target region using an adaptive threshold segmentation method; Step 5.4: Cover the corresponding area in the image data with the binarized mask of the target area, and set the pixels of the non-target area in the image data as the background to obtain the binarized label mask.
7. The intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model according to claim 6, characterized in that, The objective function is expressed as: ; In the formula, Let be the objective function. As a key parameter, For grayscale contrast, For edge smoothness, For model connectivity, The weight for grayscale contrast. As the weight for edge smoothness, These are the weights for model connectivity.
8. The intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model according to claim 1, characterized in that, In step 7, the text description includes both geometric feature description and physical feature description; The geometric feature description includes the interpretation of the target hyperbolic shape and the characterization of the target position; The physical characteristics described include the signal distribution range, signal amplitude strength, signal amplitude attenuation degree, and the presence of multiple waves.
9. The intelligent interpretation method for ground-penetrating radar defect data of a multimodal large model according to claim 1, characterized in that, Step 8 includes: Step 8.1 Acquire the ground-penetrating radar data to be tested; Step 8.2: After sequentially performing direct wave removal, bandwidth filtering, signal gain adjustment, and background filtering on the ground-penetrating radar data to be tested, the image is cropped to a preset size; Step 8.3: Convert the image data obtained in Step 8.2 into a serialized format adapted to the multimodal large model; Step 8.4: Input the serialized data into the multimodal large model after three fine-tunings for intelligent interpretation and analysis of underground diseases, and generate the comprehensive interpretation report, which includes disease category, spatial location, geometric feature description and physical feature description.
10. A multimodal large-scale ground-penetrating radar defect data intelligent interpretation device, characterized in that, The intelligent interpretation method for ground-penetrating radar defect data applicable to any one of claims 1-9 includes: The data acquisition module utilizes ground-penetrating radars with different center frequencies to collect differentiated data for different types of underground diseases and obtain raw radar data. The data preprocessing module is used to preprocess the original radar data and crop it into images of a preset size, and to annotate the disease targets in the images to construct a target detection dataset including multiple image data; and to convert the image data in the target detection dataset into data in a serialized format adapted to a multimodal large model. The model fine-tuning module performs a first fine-tuning of the multimodal large model using the target detection dataset. Based on the target detection results of the multimodal large model after the first fine-tuning, it performs binarization processing on the image data using an adaptive threshold optimization threshold segmentation technique to generate a binary label mask for the semantic segmentation task. The image data and the corresponding binary label mask are converted into a serialized format to construct a semantic segmentation dataset, which is then used to perform a second fine-tuning of the multimodal large model. Based on the semantic segmentation results of the multimodal large model after the second fine-tuning, text descriptions are annotated for the disease targets to construct an image description dataset, which is then used to perform a third fine-tuning of the multimodal large model. The interpretation and analysis module utilizes a multimodal large model that has undergone three fine-tuning steps to achieve intelligent interpretation and analysis of underground diseases and generate a comprehensive interpretation report.
Citation Information
Patent Citations
Follow-up type three-dimensional detection advanced forecasting detection method, tunnel drilling and digging equipment and application
CN118897325A
Autonomous Vehicle Sensor Fusion Using Multimodal Series Transformation with Neural Upsampling and Error Resilience
US20250389565A1