Bridge disease recognition and maintenance method and system based on multi-modal large model

By combining a multimodal large model with YOLOv5 and Llama3, efficient and automated identification and maintenance of bridge defects have been achieved, solving the problems of subjective error and environmental adaptability in the identification of bridge defects in existing technologies, and providing professional maintenance solutions.

CN120953801APending Publication Date: 2025-11-14NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511055205.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing bridge defect identification methods rely on expert experience, which are subject to subjective errors, costly, and difficult to implement quickly in different environments. Deep learning methods do not perform well in complex environments.

Method used

A bridge defect identification method based on a multimodal large model is adopted. A damage detection model is constructed using YOLOv5, and a multimodal fusion prompt generator and a zero-shot segmenter are combined. Defect detection and segmentation are performed by fusing image features and labeled features, and maintenance plans are generated using Llama3.

Benefits of technology

It improves the accuracy and automation of bridge defect identification, can adapt to different bridge defect scenarios in complex environments, and provides professional maintenance solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953801A_ABST
    Figure CN120953801A_ABST
Patent Text Reader

Abstract

The invention discloses a bridge disease recognition and maintenance method and system based on a multi-modal large model. The method comprises the steps of S1, building a general scene bridge disease training annotation data set based on historical bridge disease images and artificial disease annotation; s2, constructing a damage detection model based on YOLOv5, and training the damage detection model by using a general scene bridge disease training annotation data set to obtain a detection result of the damage detection model; s3, pre-training the general scene bridge disease training annotation data set to obtain pre-processed data; s4, inputting the detection result and the preprocessed data into a segmentation model for training to obtain a segmentation mask graph, and calculating the number and area of bridge diseases based on the segmentation mask graph; and S5, based on the number and the area of the bridge diseases, using an Llama3 model to generate a maintenance scheme, and completing the bridge disease identification and maintenance method and system based on the multi-modal large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bridge operation and maintenance technology, specifically to a method and system for bridge defect identification and maintenance based on a multimodal large model. Background Technology

[0002] With the increasing service life of bridges and the growing traffic pressure, the safety and durability of bridges, as crucial infrastructure, are receiving increasing attention. On the one hand, current identification, diagnosis, and maintenance methods are largely based on expert experience, which not only introduces subjective errors but also limits the scope of maintenance and incurs high costs. On the other hand, the application of various technologies in the field of artificial intelligence, especially computer vision, in bridge defect identification is relatively rudimentary. While these methods can effectively improve the automation level of bridge maintenance, they are still affected by the quality of the scene data used for training, making rapid application difficult in diverse scenarios. Therefore, how to provide a bridge defect maintenance method and system to automate and systematize bridge maintenance work is an urgent problem to be solved in this field.

[0003] Previously, the detection and segmentation of bridge defects mainly relied on manual visual inspection by experts. This method suffered from poor real-time performance, incomplete inspection scope, traffic disruption, and excessive costs, failing to meet current inspection needs. In recent years, the development of computer technology, especially deep learning methods represented by computer vision, has enabled a certain degree of automation in bridge maintenance. However, in real-world environments, bridge defect identification is easily affected by environmental factors. When faced with different backgrounds and interference, it often fails to achieve good detection and segmentation results. Moreover, although current deep learning detection methods have achieved a certain degree of automation, researchers still need to make some hyperparameter adjustments to the detection results. Summary of the Invention

[0004] To address the above technical problems, this invention provides a bridge defect identification and maintenance method based on a multimodal large model, the method comprising:

[0005] Step S1: Construct a training and annotation dataset for general-scenario bridge defects based on historical bridge defect images and manual defect annotations;

[0006] Step S2: Construct a damage detection model based on YOLOv5, and pre-train the damage detection model using the general scenario bridge disease training and annotation dataset to obtain the detection results of the damage detection model;

[0007] Step S3: Pre-train the general scenario bridge defect training and annotation dataset to obtain pre-trained data;

[0008] Step S4: Input the detection results and the pre-trained data into the segmentation model for training to obtain a segmentation mask image. Calculate the number and area of ​​bridge defects based on the segmentation mask image.

[0009] Step S5: Based on the number and area of ​​bridge defects, use the Llama3 model to generate a maintenance plan, thus completing the bridge defect identification and maintenance method based on a multimodal large model.

[0010] Optionally, in step S1, the historical bridge defect images include exposed reinforcement, spalling, damage, and pitting.

[0011] Optionally, in step S4, the segmentation model includes a multimodal fusion cue generator and a zero-sample segmenter;

[0012] The multimodal fusion prompt generator is used to fuse image features and labeled features to obtain a fused image. After fusion, the image features are convolved to obtain a constraint mask for the region to be segmented.

[0013] The zero-shot segmenter is used to crop the input image based on the constraint mask to obtain a test image, and then uses the zero-shot segmenter to segment the test image to obtain a binary image of image defect segmentation.

[0014] Optionally, the zero-sample segmenter includes an image encoder, a cue encoder, a mask decoder, and a segmentation enhancer;

[0015] The image encoder is used to re-encode the test image;

[0016] The prompt encoder is used to obtain point prompts for randomly distributed points based on the re-encoded test image;

[0017] The mask decoder is used to decode the re-encoded test image based on the dot cue to obtain a decoded image;

[0018] The segmentation enhancer is used to convolve and sum the image features of different depths in the decoded image, and then multiply the sum by point with the decoded image to obtain a binary image of image lesion segmentation.

[0019] This invention also discloses a bridge defect identification and maintenance system based on a multimodal large model, the system comprising:

[0020] The data acquisition and annotation module is used to build a training and annotation dataset for general-scenario bridge defects based on historical bridge defect images and manual defect annotations.

[0021] The damage detection module is used to build a damage detection model based on YOLOv5, and to pre-train the damage detection model using the general scenario bridge disease training and annotation dataset to obtain the detection results of the damage detection model.

[0022] The data pre-training module is used to pre-train the training and annotation dataset of the general scenario bridge defects to obtain pre-trained data.

[0023] The segmentation model training module is used to input the detection results and the pre-trained data into the segmentation model for training, to obtain a segmentation mask image, and to calculate the number and area of ​​bridge defects based on the segmentation mask image;

[0024] The defect identification module is used to generate maintenance plans using the Llama3 model based on the number and area of ​​bridge defects, thus completing the bridge defect identification and maintenance method based on a multimodal large model.

[0025] Optionally, in the data acquisition and annotation module, historical bridge defect images include exposed reinforcement, spalling, damage, and pitting.

[0026] Optionally, in the segmentation model training module, the segmentation model includes a multimodal fusion cue generator and a zero-shot segmenter;

[0027] The multimodal fusion prompt generator is used to fuse image features and labeled features to obtain a fused image. After fusion, the image features are convolved to obtain a constraint mask for the region to be segmented.

[0028] The zero-shot segmenter is used to crop the input image based on the constraint mask to obtain a test image, and then uses the zero-shot segmenter to segment the test image to obtain a binary image of image defect segmentation.

[0029] Optionally, the zero-sample segmenter includes an image encoder, a cue encoder, a mask decoder, and a segmentation enhancer;

[0030] The image encoder is used to re-encode the test image;

[0031] The prompt encoder is used to obtain point prompts for randomly distributed points based on the re-encoded test image;

[0032] The mask decoder is used to decode the re-encoded test image based on the dot cue to obtain a decoded image;

[0033] The segmentation enhancer is used to convolve and sum the image features of different depths in the decoded image, and then multiply the sum by point with the decoded image to obtain a binary image of image lesion segmentation.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] To address the shortcomings of current bridge defect identification and maintenance methods and systems, this invention proposes a bridge defect identification and maintenance method and system based on a multimodal large-scale model. The SAM-based visual large-scale model provides multimodal cues, enabling the model to adapt to different bridge defect scenarios with only a small number of samples, and enhancing the detection accuracy of different types of targets in complex environments, thereby improving recognition precision. The CLIP visual large-scale model can utilize textual information to focus on key features in the image, improving the defect segmentation effect.

[0036] Specifically, this invention designs a bridge defect detection method and system based on a multimodal model. After acquiring image information and performing preliminary target detection using a pre-trained YOLO detector, the region suspected of containing the defect is cropped and transmitted to a segmentation model. This segmentation model combines SAM and CLIP models, using the bounding box position, defect type, and morphological description of the defect in the cropped image region as cues for feature extraction. This leads to the segmentation of defect-containing regions, which can be used to assess the damage level. Simultaneously, by calculating the maximum closure of the segmentation, learnable evaluation conditions are provided to the detector for fine-tuning its hyperparameters and retraining the parameters. Finally, the detection results are saved and recorded. Relevant professional knowledge and records are provided to a script for report generation. Furthermore, through RAG technology, Llama 3 can acquire professional knowledge of bridge maintenance, providing specialized solutions based on the report data. Attached Figure Description

[0037] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating the method steps of the bridge defect identification and maintenance method based on a multimodal large model according to an embodiment of the present invention.

[0039] Figure 2 This is a schematic diagram of the segmentation model according to an embodiment of the present invention;

[0040] Figure 3 This is a schematic diagram of a multimodal fusion prompt generator according to an embodiment of the present invention;

[0041] Figure 4 This is a schematic diagram of the zero-sample segmenter according to an embodiment of the present invention;

[0042] Figure 5This is a schematic diagram of the Adapter structure according to an embodiment of the present invention;

[0043] Figure 6 This is a schematic diagram of the segmentation enhancer according to an embodiment of the present invention. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] Example 1

[0047] A bridge defect identification and maintenance method based on a multimodal large model, such as Figure 1 As shown, the method includes:

[0048] Step S1: Construct a training and annotation dataset for general-scenario bridge defects based on historical bridge defect images and manual defect annotations.

[0049] A training and annotation dataset for bridge defects in general scenarios was constructed, including bridge defect images and corresponding manually annotated defects, and a professional knowledge base related to bridge maintenance was established. Data sources include historical records and maintenance images from historical tasks. Annotations include bounding boxes, semantic segmentation masks, category labels, and category feature descriptions. All annotations were verified for accuracy through manual annotation or machine generation and manual checking.

[0050] In this embodiment, a general scenario bridge defect training and annotation dataset P0 is constructed based on historical bridge defect data and publicly available datasets. Images that need to be identified for the maintenance task are acquired through data acquisition equipment, and a detection dataset P1 and a few-sample annotation dataset P2 are established. The training and annotation dataset P0 includes bridge defect images and corresponding manual defect annotations, while the few-sample annotation dataset P2 includes a small number of defect images for each category in the maintenance task and corresponding manual annotations.

[0051] Step S2: Construct a damage detection model based on YOLOv5, and pre-train the damage detection model using the general scenario bridge disease training and annotation dataset to obtain the detection results of the damage detection model.

[0052] A bridge defect detector model was constructed, pre-trained using a dataset, and the results of each training round were evaluated using a test set. The bridge defect detector adopted the YOLOv5x architecture, with the dataset divided into training and test sets in an 8:2 ratio. The bridge defect images in the training set served as input to the input layer, and the corresponding manually annotated defect detection boxes served as output to the output layer. This process trained the artificial intelligence model, and the test set was used to find the model weights that yielded the best convergence, thus establishing a defect recognition model.

[0053] The YOLOv5 detector described herein employs an optimized IOU evaluation method provided by this invention to optimize model parameters. The calculation method is shown in the following formula:

[0054] ζ=R WIOU ·GIOU;

[0055]

[0056] Among them, H g and W g ... gt and y gt R represents the x and y coordinates of the center of the truth box, respectively. WIOU This represents a confidence-weighted adjustment for GIOU, combining two different assessment methods to better regress the lesion location.

[0057] Step S3: Pre-train the general scenario bridge defect training and annotation dataset to obtain pre-trained data. Step S4: Input the detection results and the pre-trained data into the segmentation model for training to obtain a segmentation mask image. Calculate the number and area of ​​bridge defects based on the segmentation mask image.

[0058] In step S4, the segmentation model includes a multimodal fusion cue generator and a zero-shot segmenter. The multimodal fusion cue generator is used to fuse image features and labeled features to obtain a fused image. After fusion, the fused image is convolved to obtain a constraint mask for the region to be segmented. The zero-shot segmenter is used to crop the input image based on the constraint mask to obtain a test image. The zero-shot segmenter is used to segment the test image to obtain a binary image of image lesion segmentation.

[0059] The zero-shot segmenter includes an image encoder, a cue encoder, a mask decoder, and a segmentation enhancer. The image encoder is used to re-encode the test image. The cue encoder is used to obtain point cues of randomly distributed points based on the re-encoded test image. The mask decoder is used to decode the re-encoded test image based on the point cues to obtain a decoded image. The segmentation enhancer is used to convolve and sum the image features of different depths in the decoded image, and then multiply the sum by the decoded image point by point to obtain a binary image of image lesion segmentation.

[0060] The segmentation model network is constructed, which includes two parts: a multimodal fusion cue generator and a zero-shot segmenter, with the following structure: Figure 2 As shown, the multimodal fusion cue generator includes an image encoder, a text encoder, and a feature fusion unit, while the zero-shot segmenter includes an image encoder, a cue encoder, a mask decoder, and a segmentation enhancer. Based on the cue generated by the multimodal fusion cue generator, the zero-shot segmenter more accurately captures the damage information of interest and automatically generates damage segments.

[0061] In this embodiment, cue points are randomly generated from the constraint mask to be used as input for the zero-sample segmentation model.

[0062] The multimodal fusion prompt generator, such as Figure 3 As shown. The image encoder extracts features from the input image at different depths; the text encoder extracts the text from the annotations and uses a weighted average method to weight the name and description features, with the name feature weighted at 0.85 and the description feature weighted at 0.15. Both parts use a CLIP VIT-L / 14 structure for initial structure and weights, and an adapter is added to each layer for fine-tuning during training. The feature fusion processor uses text information as the key and value for attention, and the convolutional image feature information as the query for feature fusion, resulting in a constraint mask for the region to be segmented. This mask represents the location information of the region to be segmented.

[0063] The zero-sample segmenter structure is as follows: Figure 4As shown. The image encoder is used to re-encode the image. Relative to the initial preset structure and weights, it is only fine-tuned during training using a learnable adapter. The cue mask consists of two parts: one is the bounding box supervision built into the detector, and the other is point cues randomly distributed based on regions obtained from the multimodal fusion cue generator. The mask decoder maintains the original settings and weights, adding only an adapter structure in the last layer. The feature enhancer combines image features of different depths from the image encoder through convolution, performs addition, and multiplies point-by-point with the output of the mask decoder. In the actual task, after the test image enters the image encoder, it first undergoes a convolutional block embedding encoding, taking 16*16 blocks with a stride of 16, downsampling the feature map, and mapping the channels from 3 to 768. After the addition of positional encoding, the feature map is passed through 16 transformer modules to complete the image encoding, with the first three layers and the last two layers using an adapter for parameter adjustment. Prompt information, such as sampling point and bounding box location information, is encoded using position coordinates, concatenated with a learnable output token, and then input into the mask decoder along with the image encoder. The mask decoder follows a SAM configuration. After mask decoding, it inputs the shallow and deep features from the image encoder, along with the binary image from the mask decoder, into the segmentation enhancer. The structure of the segmentation enhancer is as follows: Figure 6 Using the structure described above, a binary image for image defect segmentation can be obtained.

[0064] The learnable Adapter tuning, after pre-training, can use a small-sample labeled dataset P2 to fine-tune other parameters to make the model more suitable for the application scenario.

[0065] Using the few-sample labeled dataset P2, the segmentation model weights are frozen, and only the Adapter part is retained for training, allowing for adaptive weight fine-tuning of the segmentation model. The Adapter structure is as follows: Figure 5 As shown. During training, other parameters are frozen, and only the adapter settings for the first three layers and the last two layers of the encoder are fine-tuned.

[0066] Input the data from the dataset P1 to be detected into the detection model to obtain the detection results of bridge defects.

[0067] After preprocessing the detection results, the image of each defect is input into the segmentation model to obtain the segmentation mask image of the bridge defects.

[0068] Based on the mask image obtained from the segmentation, the maximum detection bounding box of the segmentation result can be obtained by calculating the maximum closure. This part is used as a pseudo-label and is used together with the detector's detection result to calculate the detection accuracy. This result is used as the optimization value for the detector's hyperparameters for optimization. The optimization method adopts particle swarm optimization, and the hyperparameters of the detector are adjusted according to the segmentation result. The adjusted hyperparameters include the maximum number of detections, the confidence threshold, and the non-maximum suppression threshold.

[0069] The formula for calculating the maximum closure is as follows:

[0070] x0 = min(x|Mask(x,y) = 1)

[0071] x1 = max(x|Mask(x,y) = 1)

[0072] y0 = min(y|Mask(x,y) = 1)

[0073] y1 = max(y|Mask(x,y) = 1)

[0074] The particle swarm optimization method is as follows:

[0075] Initial parameter settings: Maximum number of detectors 300, number of bounding boxes before NMS 3000, detection box confidence threshold 0.01, non-maximum suppression threshold 0.5. Initialization vector is (-5, -100, 0.00005, 0.0001). Parameter variations should conform to their described data type and be within their reasonable range.

[0076] Each batch records the IOU score of the matching.

[0077] Update the individual optimal position pbest and the group optimal value gbest.

[0078] Determine if the IOU >= 92%. If it does, the parameters will not be updated further; otherwise, update the vector and adjust the hyperparameter values ​​according to the following formula.

[0079] v1 = v0 + 2 × rand × (pbest) i -x i )+2×rand×(gbest i -x i )

[0080] Repeat the above steps until the conditions are met.

[0081] The accuracy of the detection is determined using the IOU calculation method described in the preparation stage of this invention.

[0082] The learnable Adapter tuning, after pre-training, can use a small-sample labeled dataset P2 to fine-tune other parameters to make the model more suitable for the application scenario.

[0083] Using the general scenario bridge defects training and annotation dataset P0, the dataset was preprocessed and divided into training and test sets in an 8:2 ratio, with corresponding text annotations added to the data. New modules of the segmentation model were pre-trained based on the training set, and the best-performing weights from the test set were selected to obtain the segmentation model.

[0084] The preprocessing operation is as follows: the target box in the image is enlarged by 0.5 times and cropped according to its position; the labeled part is processed in the same way.

[0085] The text annotation includes: adding the category name corresponding to the preprocessed image to the Prompt template "a photo of"; and generating descriptive terms for the characteristics corresponding to the disease category using Llama 3. For example, for significant crack disease, the generated disease category template is "a photo of Significant cracks", and the generated descriptive term is "long, jagged crackuneven damage".

[0086] A segmentation model is constructed and pre-trained using the preprocessed dataset P0 to obtain usable model weights.

[0087] Step S5: Based on the number and area of ​​bridge defects, use the Llama3 model to generate a maintenance plan, thus completing the bridge defect identification and maintenance method based on a multimodal large model.

[0088] Documents are loaded and split based on a knowledge base related to bridge maintenance. Embeddings and vector storage are created, and a retrieval tool is created from the vector storage. A defined helper function combines the retrieved documents into a formatted context string. When a user's question is received, the retrieval tool retrieves relevant documents, combines them into the formatted context, and passes the question and context to the language generation function of the larger model to generate a response.

[0089] The number and area of ​​bridge defects in the identification results of the dataset P1 to be detected are statistically analyzed, and a script is written to automatically generate statistical reports and ratings. Based on the statistical results, expert evaluations and suggestions are provided using the Llama 3 model, and the results are generated into a document.

[0090] The Llama 3 model is a personalized large-scale language model generated through retrieval-enhanced knowledge base (RAG) based on a knowledge base of bridge maintenance expertise. The knowledge base originates from publicly available online resources and assessment recommendations for maintenance system design.

[0091] Example 2

[0092] A bridge defect identification and maintenance system based on a multimodal large model, the system comprising:

[0093] The data acquisition and annotation module is used to build a training and annotation dataset for general-scenario bridge defects based on historical bridge defect images and manual defect annotations.

[0094] A training and annotation dataset for bridge defects in general scenarios was constructed, including bridge defect images and corresponding manually annotated defects, and a professional knowledge base related to bridge maintenance was established. Data sources include historical records and maintenance images from historical tasks. Annotations include bounding boxes, semantic segmentation masks, category labels, and category feature descriptions. All annotations were verified for accuracy through manual annotation or machine generation and manual checking.

[0095] In this embodiment, a general scenario bridge defect training and annotation dataset P0 is constructed based on historical bridge defect data and publicly available datasets. Images that need to be identified for the maintenance task are acquired through data acquisition equipment, and a detection dataset P1 and a few-sample annotation dataset P2 are established. The training and annotation dataset P0 includes bridge defect images and corresponding manual defect annotations, while the few-sample annotation dataset P2 includes a small number of defect images for each category in the maintenance task and corresponding manual annotations.

[0096] The damage detection module is used to build a damage detection model based on YOLOv5. The damage detection model is pre-trained using the general scenario bridge disease training and annotation dataset to obtain the detection results of the damage detection model.

[0097] A bridge defect detector model was constructed, pre-trained using a dataset, and the results of each training round were evaluated using a test set. The bridge defect detector adopted the YOLOv5x architecture, with the dataset divided into training and test sets in an 8:2 ratio. The bridge defect images in the training set served as input to the input layer, and the corresponding manually annotated defect detection boxes served as output to the output layer. This process trained the artificial intelligence model, and the test set was used to find the model weights that yielded the best convergence, thus establishing a defect recognition model.

[0098] The YOLOv5 detector described herein employs an optimized IOU evaluation method provided by this invention to optimize model parameters. The calculation method is shown in the following formula, where H... g and W g represents the length and width of the minimum closure of the union, respectively; A and B represent the detection box of the detector and the detection box of the pseudo-label, respectively; and x and y represent the horizontal and vertical coordinates of the box center, respectively.

[0099] ζ=R WIOU ·GIOU;

[0100]

[0101] Among them, H g and W g ... gt and y gt R represents the x and y coordinates of the center of the truth box, respectively. WIOU This represents a confidence-weighted adjustment for GIOU, combining two different assessment methods to better regress the lesion location.

[0102] The data pre-training module is used to pre-train the general scenario bridge disease training annotation dataset to obtain pre-trained data.

[0103] The segmentation model training module is used to input the detection results and the pre-trained data into the segmentation model for training, to obtain a segmentation mask image, and to calculate the number and area of ​​bridge defects based on the segmentation mask image.

[0104] The segmentation model includes a multimodal fusion cue generator and a zero-shot segmenter. The multimodal fusion cue generator is used to fuse image features and labeled features, and the fused features are then convolved to obtain a constraint mask for the region to be segmented. The zero-shot segmenter is used to crop the input image based on the constraint mask to obtain a test image. The zero-shot segmenter is then used to segment the test image to obtain a binary image of the image defect segmentation.

[0105] The zero-shot segmenter includes an image encoder, a cue encoder, a mask decoder, and a segmentation enhancer. The image encoder is used to re-encode the test image. The cue encoder is used to obtain point cues of randomly distributed points based on the re-encoded test image. The mask decoder is used to decode the re-encoded test image based on the point cues to obtain a decoded image. The segmentation enhancer is used to convolve and sum the image features of different depths in the decoded image, and then multiply the sum by the decoded image point by point to obtain a binary image of image lesion segmentation.

[0106] The segmentation model network is constructed, which includes two parts: a multimodal fusion cue generator and a zero-shot segmenter, with the following structure: Figure 2 As shown. The multimodal fusion cue generator includes an image encoder, a text encoder, and a feature fusion unit, while the zero-shot segmenter includes an image encoder, a cue encoder, a mask decoder, and a segmentation enhancer.

[0107] The multimodal fusion prompt generator, such as Figure 3As shown. The image encoder extracts features from the input image at different depths; the text encoder extracts the text from the annotations and uses a weighted average method to weight the name and description features, with the name feature weighted at 0.85 and the description feature weighted at 0.15. Both parts use a CLIP VIT-L / 14 structure for initial structure and weights, and an adapter is added to each layer for fine-tuning during training. The feature fusion processor uses text information as the key and value for attention, and the convolutional image feature information as the query for feature fusion, resulting in a constraint mask for the region to be segmented. This mask represents the location information of the region to be segmented.

[0108] The zero-sample segmenter structure is as follows: Figure 4 As shown. The image encoder is used to re-encode the image. Relative to the initial preset structure and weights, it is only fine-tuned during training through a learnable adapter. The cue mask consists of two parts: one is the bounding box supervision built into the detector, and the other is point cues randomly distributed based on the regions obtained from the multimodal fusion cue generator. In this embodiment, the multimodal cue generator identifies the regions of interest based on the image and text description. However, these regions may contain noise. Therefore, by setting a threshold, regions with potential damage can be identified from the image results inferred by the multimodal cue generator. Random sampling points are then generated in these regions as cues for the zero-shot segmenter. That is, based on the set threshold, high-probability damaged regions and irrelevant regions are divided. Then, random sampling points are generated from the high-probability damaged regions. These points have high-confidence cues, and these point cues, combined with the detection bounding boxes, serve as segmentation cues for the zero-shot segmenter, assisting in segmentation.

[0109] The mask decoder maintains the original settings and weights, adding only an Adapter structure in the last layer. In this embodiment, a small number of trainable parameters are introduced to adjust the pre-trained model, adapting it to a specific task or domain, while keeping most of the original model weights unchanged. This method significantly reduces computational cost and memory usage compared to fully tuning model weights, while still effectively improving model performance. The feature enhancer combines image features of different depths from the image encoder through convolution, performs addition, and multiplies pointwise with the output of the mask decoder. In practical tasks, after the test image enters the image encoder, it first undergoes a convolutional block embedding encoding, taking 16*16 blocks with a stride of 16, downsampling the feature map, and mapping the channels from 3 to 768. After the addition of positional encodings, the feature map is processed through 16 transformer modules to complete image encoding, with the first three layers and the last two layers using an Adapter for parameter adjustment. Prompt information, such as sampling points and detection box positions, is encoded using positional coordinates, concatenated with a learnable output token, and then input into the mask decoder along with the image encoding. The mask decoder follows the SAM configuration. After mask decoding, it inputs the shallow and deep features from the image encoder, along with the binary image decoded from the mask, into the segmentation enhancer. The structure of the segmentation enhancer is as follows: Figure 6 Using the structure described above, a binary image for image defect segmentation can be obtained.

[0110] The learnable Adapter tuning, after pre-training, can use a small-sample labeled dataset P2 to fine-tune other parameters to make the model more suitable for the application scenario.

[0111] Using the few-sample labeled dataset P2, the segmentation model weights are frozen, and only the Adapter part is retained for training, allowing for adaptive weight fine-tuning of the segmentation model. The Adapter structure is as follows: Figure 5 As shown. During training, other parameters are frozen, and only the adapter settings for the first three layers and the last two layers of the encoder are fine-tuned.

[0112] Input the data from the dataset P1 to be detected into the detection model to obtain the detection results of bridge defects.

[0113] After preprocessing the detection results, the image of each defect is input into the segmentation model to obtain the segmentation mask image of the bridge defects.

[0114] Based on the mask image obtained from the segmentation, the maximum detection bounding box of the segmentation result can be obtained by calculating the maximum closure. This part is used as a pseudo-label and is used together with the detector's detection result to calculate the detection accuracy. This result is used as the optimization value for the detector's hyperparameters for optimization. The optimization method adopts particle swarm optimization, and the hyperparameters of the detector are adjusted according to the segmentation result. The adjusted hyperparameters include the maximum number of detections, the confidence threshold, and the non-maximum suppression threshold.

[0115] The formula for calculating the maximum closure is as follows:

[0116] x0 = min(x|Mask(x,y) = 1)

[0117] x1 = max(x|Mask(x,y) = 1)

[0118] y0 = min(y|Mask(x,y) = 1)

[0119] y1 = max(y|Mask(x,y) = 1)

[0120] The particle swarm optimization method is as follows:

[0121] Initial parameter settings: Maximum number of detectors 300, number of bounding boxes before NMS 3000, detection box confidence threshold 0.01, non-maximum suppression threshold 0.5. Initialization vector is (-5, -100, 0.00005, 0.0001). Parameter variations should conform to their described data type and be within their reasonable range.

[0122] Each batch records the IOU score of the matching.

[0123] Update the individual's best position (pbest) and the group's best value (gbest);

[0124] Determine if IOU >= 92%. If it does, the parameters will not be updated. Otherwise, update the vector and change the hyperparameter values ​​according to the following formula.

[0125] v1 = v0 + 2 × rand × (pbest) i -x i )+2×rand×(gbest i —x i )

[0126] Repeat the above steps until the conditions are met.

[0127] The accuracy of the detection is determined using the IOU calculation method described in the preparation stage of this invention.

[0128] Learnable Adapter tuning allows for fine-tuning of the model after pre-training by using a small-sample labeled dataset P2 and freezing other parameters, making the model more suitable for the application scenario.

[0129] Using the general scenario bridge defects training and annotation dataset P0, the dataset was preprocessed and divided into training and test sets in an 8:2 ratio, with corresponding text annotations added to the data. New modules of the segmentation model were pre-trained based on the training set, and the best-performing weights from the test set were selected to obtain the segmentation model.

[0130] The preprocessing operation is as follows: the target box in the image is enlarged by 0.5 times and cropped according to its position; the labeled part is processed in the same way.

[0131] The text annotation includes: adding the category name corresponding to the preprocessed image to the Prompt template "a photo of"; and generating descriptive terms for the characteristics corresponding to the disease category using Llama 3. For example, for significant crack disease, the generated disease category template is "a photo of Significant cracks", and the generated descriptive term is "long, jagged crackuneven damage".

[0132] A segmentation model is constructed and pre-trained using the preprocessed dataset P0 to obtain usable model weights.

[0133] The defect identification module is used to generate maintenance plans using the Llama3 model based on the number and area of ​​bridge defects, thus completing the bridge defect identification and maintenance method based on a multimodal large model.

[0134] Documents are loaded and split based on a knowledge base related to bridge maintenance. Embeddings and vector storage are created, and a retrieval tool is created from the vector storage. A defined helper function combines the retrieved documents into a formatted context string. When a user's question is received, the retrieval tool retrieves relevant documents, combines them into the formatted context, and passes the question and context to the language generation function of the larger model to generate a response.

[0135] The number and area of ​​bridge defects in the identification results of the dataset P1 to be detected are statistically analyzed, and a script is written to automatically generate statistical reports and ratings. Based on the statistical results, expert evaluations and suggestions are provided using the Llama 3 model, and the results are generated into a document.

[0136] The Llama 3 model is a large, personalized language model generated through retrieval-enhanced knowledge base (RAG) based on a knowledge base of bridge maintenance expertise. The knowledge base originates from publicly available online resources and assessment recommendations for maintenance system design.

[0137] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for bridge defect identification and maintenance based on a multimodal large model, characterized in that, The method includes: Step S1: Construct a training and annotation dataset for general-scenario bridge defects based on historical bridge defect images and manual defect annotations; Step S2: Construct a damage detection model based on YOLOv5, and pre-train the damage detection model using the general scenario bridge disease training and annotation dataset to obtain the detection results of the damage detection model; Step S3: Pre-train the general scenario bridge defect training and annotation dataset to obtain pre-trained data; Step S4: Input the detection results and the pre-trained data into the segmentation model for training to obtain a segmentation mask image. Calculate the number and area of ​​bridge defects based on the segmentation mask image. Step S5: Based on the number and area of ​​bridge defects, use the Llama3 model to generate a maintenance plan, thus completing the bridge defect identification and maintenance method based on a multimodal large model.

2. The bridge defect identification and maintenance method based on a multimodal large model according to claim 1, characterized in that, In step S1, the historical bridge defect images include exposed reinforcement, peeling, damage, and pitting.

3. The bridge defect identification and maintenance method based on a multimodal large model according to claim 1, characterized in that, In step S4, the segmentation model includes a multimodal fusion cue generator and a zero-sample segmenter; The multimodal fusion prompt generator is used to fuse image features and labeled features to obtain a fused image. After fusion, the image features are convolved to obtain a constraint mask for the region to be segmented. The zero-shot segmenter is used to crop the input image based on the constraint mask to obtain a test image, and then uses the zero-shot segmenter to segment the test image to obtain a binary image of image defect segmentation.

4. The bridge defect identification and maintenance method based on a multimodal large model according to claim 3, characterized in that, The zero-sample segmenter includes an image encoder, a cue encoder, a mask decoder, and a segmentation enhancer; The image encoder is used to re-encode the test image; The prompt encoder is used to obtain point prompts for randomly distributed points based on the re-encoded test image; The mask decoder is used to decode the re-encoded test image based on the dot cue to obtain a decoded image; The segmentation enhancer is used to convolve and sum the image features of different depths in the decoded image, and then multiply the sum by point with the decoded image to obtain a binary image of image lesion segmentation.

5. A bridge defect identification and maintenance system based on a multimodal large model, the system being used to implement the bridge defect identification and maintenance method based on a multimodal large model as described in any one of claims 1-4, characterized in that, The system includes: The data acquisition and annotation module is used to build a training and annotation dataset for general-scenario bridge defects based on historical bridge defect images and manual defect annotations. The damage detection module is used to build a damage detection model based on YOLOv5, and to pre-train the damage detection model using the general scenario bridge disease training and annotation dataset to obtain the detection results of the damage detection model. The data pre-training module is used to pre-train the training and annotation dataset of the general scenario bridge defects to obtain pre-trained data. The segmentation model training module is used to input the detection results and the pre-trained data into the segmentation model for training, to obtain a segmentation mask image, and to calculate the number and area of ​​bridge defects based on the segmentation mask image; The defect identification module is used to generate maintenance plans using the Llama3 model based on the number and area of ​​bridge defects, thus completing the bridge defect identification and maintenance method based on a multimodal large model.

6. The bridge defect identification and maintenance system based on a multimodal large model according to claim 5, characterized in that, In the data acquisition and annotation module, historical bridge defect images include exposed reinforcement, spalling, damage, and pitting.

7. The bridge defect identification and maintenance system based on a multimodal large model according to claim 6, characterized in that, The segmentation model training module includes a multimodal fusion cue generator and a zero-shot segmenter. The multimodal fusion prompt generator is used to fuse image features and labeled features to obtain a fused image. After fusion, the image features are convolved to obtain a constraint mask for the region to be segmented. The zero-shot segmenter is used to crop the input image based on the constraint mask to obtain a test image, and then uses the zero-shot segmenter to segment the test image to obtain a binary image of image defect segmentation.

8. The bridge defect identification and maintenance system based on a multimodal large model according to claim 7, characterized in that, The zero-sample segmenter includes an image encoder, a cue encoder, a mask decoder, and a segmentation enhancer; The image encoder is used to re-encode the test image; The prompt encoder is used to obtain point prompts for randomly distributed points based on the re-encoded test image; The mask decoder is used to decode the re-encoded test image based on the dot cue to obtain a decoded image; The segmentation enhancer is used to convolve and sum the image features of different depths in the decoded image, and then multiply the sum by point with the decoded image to obtain a binary image of image lesion segmentation.

Citation Information

Cited By

  • Infrastructure apparent disease detection method, system and terminal based on AI large model

    CN122024075A

  • Pavement disease identification method, system and equipment based on semi-supervised learning, and medium

    CN122116155A

  • Road disease identification method, system, device and medium based on semi-supervised learning

    CN122116155B