Bridge multi-mode multi-target disease intelligent identification method and device under complex background and medium
By adopting multimodal multi-objective disease intelligent identification method in bridge detection, the problems of low efficiency and poor accuracy of bridge disease identification in complex backgrounds are solved, efficient and accurate disease identification is achieved, and maintenance costs are reduced.
Patent Information
- Application Number
- CN202510224401.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing bridge detection methods are difficult to effectively identify multimodal multi-target diseases in complex contexts, resulting in low detection efficiency, poor accuracy, and safety hazards and high maintenance costs.
The intelligent identification method of multimodal multi-objective disease of bridges in complex contexts is adopted. By acquiring and labeling bridge image data sets, data expansion and multi-objective detection model training are performed, and the model with the best performance is selected for identification.
It realizes efficient identification of multimodal and multi-target diseases of bridges in complex contexts, improves the accuracy and efficiency of detection, and reduces the burden of human resources and maintenance costs.
Smart Images

Figure CN120071007A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bridge disease identification, and more specifically, to an intelligent identification method, device and medium for multi-modal and multi-objective diseases of bridges under complex backgrounds. Background Art
[0002] As an important part of the road traffic network, the existence of bridges is directly related to the safety and optimization performance of the transportation system. However, due to long-term service and natural factors, bridges will encounter diseases such as weathering, disintegration, cracks and spalling caused by increased loads. If these diseases cannot be detected and treated in the early stage, it will lead to an increase in the loads borne by bridge components, and even cause damage to the basic components, resulting in unpredictable economic losses and casualties.
[0003] At present, most of the detection methods for bridges are divided into two categories. One is manual inspection, and the other is sensor monitoring. The manual detection method is difficult to obtain accurate data of the bridge during operation in real time, so it is impossible to analyze the health status of the bridge structure in time. Secondly, it is difficult to ensure the stability and continuity of the data. In addition, manual detection may also pose a threat to the safety of the staff, increasing the work risk. The sensor monitoring method has high maintenance costs, a large amount of monitoring data, and limitations in its own battery life. Therefore, these two types of detection methods cannot provide effective and timely monitoring for the bridge structure.
[0004] To solve these problems, all parties involved in bridge detection are working hard to adopt automation technologies to reduce the human and material resources required in the inspection process. Automation technologies such as drones, robots, wireless sensor networks, machine learning, and artificial intelligence are used. These technologies cover all aspects of the inspection process, thereby improving the efficiency and accuracy of the inspection, helping to improve the bridge maintenance efficiency, reducing the burden on human resources, saving costs, improving safety, and promoting digital transformation. Deep learning technology is one of the most widely used technologies among them. Deep learning is a machine learning technology based on multi-layer neural networks. It can automatically learn feature representations from massive amounts of data and realize complex non-linear mapping relationships. Bridge diseases have strong non-linear characteristics, which can just realize complex non-linear mapping relationships. Deep learning has extensive and effective applications in the field of image processing, such as classification, object detection, semantic segmentation, instance segmentation, etc., and also shows excellent performance in the identification of bridge appearance diseases.
[0005] Computer vision-based detection methods can be roughly divided into two categories: The first category is the candidate region algorithm represented by R-CNN, also known as the two-stage algorithm; the second category is the regression-based algorithm represented by YOLO, also known as the single-stage algorithm. The single-stage detection algorithm is relatively simple and fast. It directly gives the bounding box and category through one-time prediction. This type of method has a fast detection speed and is suitable for real-time demand scenarios. However, its detection effect is poor when detecting objects with many small targets. The two-stage detection algorithm adopts a step-by-step prediction process. First, it uses prior boxes to propose candidate regions (candidats), and then separately predicts their categories. It has high detection accuracy, but a large amount of computation and low efficiency. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a method, device and medium for intelligent recognition of multi-modal and multi-object diseases of bridges under complex backgrounds, so as to solve the problems of complex backgrounds, unclear features and complex environments in the task of detecting bridge apparent disease targets.
[0007] In the first aspect, the present invention provides a method for intelligent recognition of multi-modal and multi-object diseases of bridges under complex backgrounds, and the method includes:
[0008] Obtain an image data set and a label data set; wherein, the data set includes multiple bridge images with diseases, the bridge images include various textures, lighting conditions and complex backgrounds, and the label data set includes labels for identifying various bridge damages and key bridge components;
[0009] Annotate the image data set based on the label data set to obtain an annotated data set;
[0010] Perform interpolation processing on the annotated data set, and after performing random flipping, adding noise and grayscaling processing, obtain an augmented data set;
[0011] Establish at least two multi-object detection models, and train the multi-object detection models based on the augmented data set to obtain multiple trained multi-object detection models;
[0012] Evaluate the multiple trained multi-object detection models using evaluation metrics, screen out the trained multi-object detection model with the best performance as the disease intelligent multi-object detection model, and realize the intelligent recognition of multi-modal and multi-object diseases of bridges under complex backgrounds based on the disease intelligent recognition.
[0013] Further, the interpolation processing of the annotated data set is performed through the following formula:
[0014]
[0015]
[0016] In the formula, L(x) represents the Lanczos kernel function; x represents the relative offset of the sample position; a represents the size of the Lanczos kernel; f(k) represents the source pixel value in the image; L(x - k)·f(k) represents the Lanczos kernel; L(x - k) represents the form after the translation of the Lanczos kernel function L(x), f(x) represents the interpolated pixel value; k represents the integer position (discrete coordinate) of the pixel in the original image.
[0017] Further, the at least two multi-object detection models include at least two of the YOLOv8 model, the YOLOv11 model, the RT-DETR model, and the Mamba-YOLO model.
[0018] Further, the YOLOv8 model includes a C2f module, a neck network part, and a detection head part. The C2f module includes 3 convolutional modules and n Bottleneck modules; the detection head part adopts a decoupled head structure; when training the YOLOv8 model, an adaptive data augmentation strategy is adopted to automatically adjust the intensity of data augmentation according to the difficulty of object detection.
[0019] Further, the YOLOv11 model includes a backbone network and an SPFF module; the backbone network includes a C3k2 module and a C2PSA module. The C3k2 module includes an outer layer and an inner layer. The outer layer is a C2f structure, and the inner layer is a C3 structure; the C2PSA module includes a position-sensitive attention model and a C2PASA block; the C2PSA module applies position-sensitive attention and a feed-forward network to the input tensor, thereby enhancing the feature extraction and processing capabilities. The SPFF module aggregates multi-scale context information through max-pooling operations with multiple different kernel sizes.
[0020] Further, the RT-DETR model is based on the DETR model. A main chain based on CORS and an effective hybrid encoder are introduced into the DETR model to obtain real-time speed. The RT-DRET model processes multi-scale features through decoupled scale interactions and cross-scale fusions.
[0021] Further, the Mamba-YOLO model applies a state space transformation model based on the state space model to each layer of YOLO to effectively capture global dependencies, and uses the advantages of local convolution to improve the detection accuracy and the model's understanding ability of complex scenes while maintaining real-time performance.
[0022] Further, the evaluation metrics include precision, recall, average precision, and F1 score, which are calculated by the following formulas:
[0023]
[0024] Wherein, recall represents the recall rate, that is, the proportion of positive examples predicted as true (TP) in all samples that are actually positive examples, precision represents the precision, TP represents positive examples predicted as true (predicted as positive examples and actually positive examples), FN represents positive examples predicted as false (predicted as negative examples but actually positive examples), FP represents negative examples predicted as true (predicted as positive examples but actually negative examples), F1-score represents the weighted average of accuracy and recall rate, IOU represents the intersection over union, which is used to quantify the proximity of two bounding boxes (ground truth and prediction), area of overlap represents the intersecting area of the bounding boxes, and area of union represents the union area of the two bounding boxes.
[0025] In a second aspect, the present invention provides a multi-modal multi-object damage intelligent recognition device for bridges under complex backgrounds. The device includes:
[0026] A data acquisition unit configured to acquire an image data set and a label data set; wherein, the data set includes a plurality of bridge images with diseases, the bridge images include various textures, lighting conditions, and complex backgrounds, and the label data set includes labels for identifying various types of bridge damages and key bridge components;
[0027] A data annotation unit configured to annotate the image data set based on the label data set to obtain an annotated data set;
[0028] A data augmentation unit configured to perform interpolation processing on the annotated data set, and after performing random flipping, noise addition, and grayscale processing, obtain an augmented data set;
[0029] A model construction unit configured to establish at least two multi-object detection models, and train the multi-object detection models based on the augmented data set to obtain a plurality of trained multi-object detection models;
[0030] A model screening unit configured to evaluate the plurality of trained multi-object detection models using evaluation metrics, screen out the trained multi-object detection model with the optimal performance as the intelligent multi-object detection model for diseases, and realize the intelligent recognition of multi-modal multi-object diseases of bridges under complex backgrounds based on the intelligent recognition of diseases.
[0031] In a third aspect, the present invention provides a readable storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the method as described above.
[0032] The present invention has at least the following beneficial effects:
[0033] In view of the problems of complex background, unclear features, and complex environment in the target detection task of bridge apparent diseases, the present invention constructs a dataset consisting of 3,996 real images and their annotations, with different features of different textures, illuminations, and complex backgrounds (such as graffiti and occlusion), forming a multi-modal and multi-object bridge dataset in a complex environment. Training is carried out on the constructed dataset based on current advanced target detection models (YOLOv8, YOLOv11, RT-DETR, Mamba-YOLO). By evaluating the real-time performance, accuracy, and adaptability of each model, the YOLOv11 model is finally selected, providing new ideas for future bridge algorithm improvement work. Among them, the YOLOv11 model shows good results in the detection and location of bridge diseases in complex backgrounds, and the detection accuracies of cracks, spalling, exposed reinforcement, and efflorescence are 81.5%, 94.6%, 87.2%, and 83.9% respectively, and the precision, recall, and F1 score are 89.2%, 80%, and 83.5% respectively, with only 6.3 GFLOPS. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Shows the overall flowchart of a method for intelligent identification of multi-modal and multi-object bridge diseases in a complex background according to an embodiment of the present invention.
[0035] Figure 2 Shows a schematic diagram of a bridge image according to an embodiment of the present invention; wherein, (a) the screened picture (b) Crack500 picture;
[0036] Figure 3 Shows the dataset annotation diagram according to an embodiment of the present invention;
[0037] Figure 4 Shows the data augmentation display diagram according to an embodiment of the present invention;
[0038] Figure 5 Shows the training set display diagram according to an embodiment of the present invention;
[0039] Figure 6 Shows the overall architecture diagram of the YOLOv8 model according to an embodiment of the present invention;
[0040] Figure 7 Shows the C2f module structure diagram of the YOLOv8 model according to an embodiment of the present invention;
[0041] Figure 8 Shows the partial processing flowchart of the YOLOv8 model according to an embodiment of the present invention;
[0042] Figure 9Shows the network architecture diagram of the YOLOv11 model according to an embodiment of the present invention;
[0043] Figure 10 Shows the structural diagram of the C3K2 module in the YOLOv11 model according to an embodiment of the present invention;
[0044] Figure 11 Shows the structural diagram of the C2PSA module in the YOLOv11 model according to an embodiment of the present invention;
[0045] Figure 12 Shows the structural diagram of the SPFF module in the YOLOv11 model according to an embodiment of the present invention;
[0046] Figure 13 Shows the architecture diagram of the RT-DETR model according to an embodiment of the present invention;
[0047] Figure 14 Shows the architecture diagram of the MambaYOLO model according to an embodiment of the present invention;
[0048] Figure 15 Shows the flowchart from the preparation of the entire dataset to model training according to an embodiment of the present invention;
[0049] Figure 16 Shows the inference result diagram of bridge disease detection by each model according to an embodiment of the present invention. Detailed implementation manners
[0050] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners. The embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, but shall not be construed as a limitation to the present invention. For the various steps described herein, if there is no necessity for a front-back relationship between them, the order in which they are described as examples herein shall not be construed as a limitation. Those skilled in the art should know that they can adjust the order as long as the logic between them is not destroyed and the entire process cannot be realized.
[0051] The embodiment of the present invention provides a method for intelligent recognition of multi-modal and multi-object bridge diseases in complex backgrounds. Aiming at the problems of complex backgrounds, unclear features, and complex environments in the task of bridge apparent disease target detection, a dataset composed of 3,996 real images and their annotations is constructed, with different features of different textures, illuminations, and complex backgrounds (such as graffiti and occlusion), forming a multi-modal and multi-object bridge dataset in a complex environment. This embodiment trains based on the current advanced object detection models on the constructed dataset, evaluates the performance of the current latest object detection models, and provides new ideas for future bridge algorithm improvement work.
[0052] This method automatically detects damages to bridge images in complex backgrounds in several steps: (1) constructing a dataset of different bridge diseases in complex backgrounds, (2) selecting the latest deep learning model to train the bridge dataset and applying it to actual engineering, and (3) testing and analyzing the performance of the model to detect diseases in bridge images.
[0053] Specifically, Figure 1 FIG. shows the overall flowchart of a multi-modal and multi-objective disease intelligent recognition method for bridges in complex backgrounds according to an embodiment of the present invention. As Figure 1 shown, this multi-modal and multi-objective disease intelligent recognition method for bridges in complex backgrounds includes steps S1 to S5, which are introduced in detail as follows.
[0054] S1. Obtain an image dataset and a label dataset; wherein, the dataset includes multiple bridge images with diseases, the bridge images include various textures, lighting conditions, and complex backgrounds, and the label dataset includes labels for identifying various types of bridge damages and key bridge components.
[0055] The dataset used in this embodiment is composed of the dacl_10k public dataset, the Crack500 dataset, and a bridge dataset taken by mobile phones. This dataset is composed of bridge images with diseases, various textures, lighting conditions, and complex backgrounds (such as graffiti and markings).
[0056] S2. Annotate the image dataset based on the label dataset to obtain an annotated dataset.
[0057] Since the dacl_10k dataset is large, in this embodiment, 1300 bridge disease pictures of four types, namely cracks, spalls, exposed reinforcement, and efflorescence, are selected, and 300 pictures from the Crack500 dataset and pictures of bridge diseases on a certain highway bridge in Chengdu are added, totaling 2001 pictures, as Figure 2 shown. And the labelImg tool is used to perform rectangular box annotation on the pictures to generate label files in YOLO format, as Figure 3 shown.
[0058] S3. Perform interpolation processing on the annotated dataset, and after performing random flipping, noise addition, and grayscale processing, obtain an augmented dataset.
[0059] Step S3 is the step of preprocessing the dataset. The labeled dataset is imported into Roboflow for preprocessing, and Lanczos interpolation is used to convert the images into images of size 640×640. Lanczos interpolation is an image interpolation method based on a finite-term convolution kernel, which is usually used to improve the image quality when scaling (zooming in or out) images. It is calculated through a window function called the Lanczos kernel and is a high-quality interpolation method, especially suitable for resampling high-resolution images.
[0060] Lanczos interpolation is based on the Sinc function. It obtains a smoother and higher-quality image by interpolating between pixels. Specifically, Lanczos interpolation uses a Sinc function window of finite size, usually a 3×3 or 5×5 window. This interpolation method can be used for both image magnification and reduction, and can reduce the blur and artifacts common in many low-quality interpolation methods. The kernel function (also called the window function) of Lanczos interpolation is based on the truncation of the Sinc function:
[0061]
[0062] where: x is the relative offset of the sample position, usually the distance between the target pixel and the source pixel. a is the parameter of the Lanczos kernel, usually a positive integer, and here it takes the value of 4. It determines the "width" of the kernel function, that is, the width of the non-zero part of the kernel function.
[0063] For any pixel position, the formula for Lanczos interpolation can be expressed as:
[0064]
[0065] Subsequently, data augmentation techniques are used to perform multiple operations such as random flipping, adding noise, and grayscaling (that is, the enhanced images are enhanced by two or more methods), and the dataset is expanded to 3996 images, as Figure 4 shown.
[0066] S4. Establish at least two multi-object detection models, train the multi-object detection models based on the expanded dataset, and obtain multiple trained multi-object detection models.
[0067] In this embodiment, Python programming is used to divide the dataset into a training set, a test set, and a validation set in a ratio of 8:1:1, and programming is used to make the quantity of each category in the training set as balanced as possible to prevent underfitting or overfitting problems.
[0068] In this embodiment, four multi-object detection models are adopted, namely the YOLOv8 model, the YOLOv11 model, the RT-DETR model, and the Mamba-YOLO model.
[0069] The YOLOv8 model provides advanced performance in object detection, image classification, and instance segmentation tasks. The overall architecture of the YOLOv8 model is as Figure 6 shown. The backbone network of the YOLOV8 model uses the C2f module. In the neck network part, the YOLOv8 model uses the idea of PAN, as Figure 3 shown. In the detection head part, the Yolov8 model adopts a decoupled head structure. For training data, the YOLOv8 model adopts an adaptive data augmentation strategy, which can automatically adjust the intensity of data augmentation according to the difficulty of object detection.
[0070] The YOLOv11 model continues the structural pattern of the YOLOv8 model. Compared with the YOLOv8 model, the YOLOv11 model adopts an improved backbone network and neck network architecture, which improves the feature extraction ability, enhances the object detection accuracy and complex task performance, and its network architecture is as Figure 9 shown. The backbone network of YOLO11 introduces two new modules, C3k2 and C2PSA. The C3K2 module provides stronger feature extraction ability, especially suitable for complex scenarios and deep feature extraction tasks. The specific model structure is as Figure 10 shown. The C2PSA module (Cross-stage Partial Spatial Attention) includes position-sensitive attention and the C2PASA block. This module applies position-sensitive attention and a feed-forward network to the input tensor, thereby enhancing the feature extraction and processing ability, and its architecture diagram is as Figure 11 shown. In the neck network part, the YOLOv11 model retains the SPFF module (SpatialPyramid Pooling Fast), which aggregates the features of each region of the image through different scales. SPFF aggregates multi-scale context information through max-pooling operations with multiple different kernel sizes, as Figure 12 shown.
[0071] The RT-DETR model is based on the idea of DETR (a framework without NMS), and at the same time introduces a backbone based on CORS and an effective hybrid encoder to achieve real-time speed. RT-DRET effectively processes multi-scale features by decoupling the interaction of scales and cross-scale fusion. The model is highly adaptable and can use different decoder layers without retraining to support flexible adjustment of the inference speed. Its architecture is as Figure 13 shown.
[0072] The MambaYOLO model introduces the ODSSBlock module and applies the SSM structure to the field of object detection, and its architecture is as Figure 14As shown. It is divided into the backbone and neck parts of ODMamba. ODMamba consists of a Simple Stem and a Downsample Block. In the neck part, following the design of PAN-FPN, the ODSSBlock module is used to replace C2f to capture a richer gradient information flow. The backbone first performs downsampling through the backbone module. This architecture applies the state space transformation model based on SSM to each layer of YOLO to effectively capture global dependencies, and utilizes the advantages of local convolution to improve the detection accuracy and the model's understanding ability of complex scenes, while maintaining real-time performance. Its mAP on MSCOCO is 8.1% higher than that of the baseline YOLOv8.
[0073] Each of the above models has unique advantages and disadvantages, making them suitable for different applications. In this embodiment, the above four models are used to train the newly constructed dataset to explore the performance of these models for bridge disease detection in complex backgrounds and environments.
[0074] This embodiment is carried out on the Windows operating system, with the GPU being NVIDIA GeForce RTX 3090 (24GB video memory) and the CPU being Intel(R) Xeon(R) Gold 6152 (10 cores). The development platform is PyCharm 2.3.1, and the programming language is Python. This embodiment is based on the PyTorch deep learning framework developed by Facebook, with the torch version being 2.1.0. In the virtual environment (Python 3.10) created by Anaconda, YOLOv8, YOLOv11, RT-DETR, and Mamba-YOLO are respectively used to train the dataset constructed in this embodiment.
[0075] During the training process, the pre-trained weights used are all downloaded from the official website. When training, the batch size is set to 32, the maximum number of training epochs is 500, the SGD optimizer is used, and the learning rate is set to 0.01. To prevent overfitting, the patience value is set to 50. The Mosaic method is adopted to combine multiple pictures into one picture according to a certain ratio, enabling the model to identify targets within a smaller range, improving the detection performance of small targets, enhancing data diversity and model robustness. Figure 15 Shows the process from the preparation of the entire dataset to model training.
[0076] S5. Use evaluation metrics to evaluate multiple trained multi-object detection models, and select the trained multi-object detection model with the best performance as the intelligent multi-object detection model for diseases, and realize intelligent multi-modal multi-object disease identification for bridges under complex backgrounds based on intelligent disease identification.
[0077] The performance of the model is evaluated based on four parameters, namely precision, recall, mean average precision (mAP), and F1-score. Precision is the proportion of positive samples predicted as positive that are actually positive. In object detection, if the bounding box predicted by the model coincides with the true bounding box, the prediction is considered correct. Recall is the proportion of positive samples that are actually positive and are correctly predicted as positive. In object detection, if the true bounding box coincides with the predicted bounding box, the sample is considered to be correctly recalled. Mean average precision (mAP) is a curve that shows the performance of a classifier by plotting precision against recall. The F1-score is a metric for evaluating the model's ability to identify, classify, and locate different cracks.
[0078] To evaluate the performance of the detection model, the intersection over union (IoU) metric is usually used to quantify the quality of bridge detection. The IoU measures the overlap between the ground truth bounding box and the predicted bounding box. A predefined threshold is usually set to determine the accuracy of the bounding box. If the IoU between the ground truth and the predicted bounding box exceeds the set threshold, the crack detection is considered valid. Conversely, if the IoU is below the threshold, the detection is considered invalid. Usually, the solid box represents the ground truth bounding box, the dashed box represents the predicted bounding box, and the IoU value is the ratio of the intersection area to the union area of the two boxes. The higher the IoU value, the better the consistency between the predicted value and the ground truth bounding box, indicating more accurate detection. The evaluation metrics are shown in equations (3)-(7):
[0079]
[0080] Multi-object testing of diseases:
[0081] Figure 16 These are the inference result graphs of each model for bridge disease detection. In the figure, (a), (b), (c), (d), and (e) show the detection results of the four models, YOLOv8, YOLOv11, RT-DETR, and Mamba-YOLO, for diseases such as exposed reinforcement, spalling, efflorescence, and cracks.
[0082] The discovery of this embodiment is of great significance for on-site structural health assessment. We applied the current mainstream object detection models to the detection of bridge diseases under complex conditions and environments. Through comparison, we found that when there is sufficient computing resources, the RT-DETR model is obviously superior. However, in scenarios where real-time processing is crucial for edge computing devices, this becomes a significant limitation. In this case, YOLOv11 and Mamba-YOLO are more suitable. In the later stage, we can consider using some structural networks of Mamba-YOLO to improve YOLOv11, while ensuring low computing resource consumption and real-time performance, and further making up for the deficiencies in the accuracy and positioning of YOLOv11. Using an efficient method like YOLOv11 can speed up the process of identifying potential problems and damages in the structure. By accurately detecting damages and other signs of deterioration, timely repairs and maintenance can be initiated, thereby minimizing the risk of bridge structure failures and ensuring traffic safety.
[0083] This embodiment proposes a method for automatically detecting concrete structure image damages using the YOLO model based on computer vision, especially version 11 and the Mamba variant model. These models contribute to the cost-effective preliminary health monitoring of structures. The training of these models is carried out on a constructed dataset of bridge diseases with complex backgrounds and environments. The average precision of the YOLOv11 model in detecting cracks, spalling, exposed reinforcement, and efflorescence weathering of bridge apparent diseases in complex backgrounds and environments is 81.5%, 94.6%, 87.2%, and 83.9% respectively, and its accuracy, recall, and F1 score are 89.2%, 80%, and 83.5% respectively. In terms of detection speed, it is superior to the other three models, only 6.3 GFLOPS. The future research scope will modify the model structure to increase the detection accuracy and robustness (such as using Mamba to optimize the network structure). The fine-tuned model will have further potential to be integrated with other evaluation devices and be carried on unmanned aerial vehicle devices, so as to achieve cheap, real-time, and effective preliminary health monitoring of any structure and promote the development of urban low-altitude economy and urban air traffic.
[0084] The embodiment of the present invention also provides a device for intelligent identification of multi-modal and multi-object bridge diseases under complex backgrounds, and the device includes:
[0085] A data acquisition unit configured to acquire an image dataset and a label dataset; wherein, the dataset includes multiple bridge images with diseases, the bridge images include various textures, lighting conditions, and complex backgrounds, and the label dataset includes labels for identifying various types of bridge damages and key bridge components;
[0086] A data annotation unit configured to annotate the image dataset based on the label dataset to obtain an annotated dataset;
[0087] A data augmentation unit, configured to perform interpolation processing on the labeled data set, and after performing random flipping, noise addition, and grayscaling processing, obtain an augmented data set;
[0088] A model construction unit, configured to establish at least two multi-object detection models, and train the multi-object detection models based on the augmented data set to obtain multiple trained multi-object detection models;
[0089] A model screening unit, configured to evaluate the multiple trained multi-object detection models using evaluation metrics, screen out the trained multi-object detection model with the best performance as the intelligent multi-object detection model for diseases, and realize intelligent identification of multi-modal multi-object diseases of bridges under complex backgrounds based on the intelligent identification of diseases.
[0090] It should be noted that the structures of the intelligent multi-modal multi-object disease identification devices for bridges under complex backgrounds described in this embodiment belong to the same inventive concept as the previously described intelligent multi-modal multi-object disease identification method for bridges under complex backgrounds, and achieve the same beneficial effects through the same principle, which will not be elaborated here.
[0091] The embodiment of the present invention also provides a readable storage medium, where the readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in any of the above embodiments.
[0092] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. In addition, in the above specific embodiments, various features can be grouped together to simplify the present invention. This should not be construed as an intention that the features of an unclaimed invention are necessary for any claim. On the contrary, the subject matter of the present invention may be less than all the features of a specific embodiment of the invention. Thus, the following claims are incorporated herein as examples or embodiments in the specific embodiments, where each claim independently serves as a separate embodiment, and it is contemplated that these embodiments can be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of the equivalents of the claims.
Claims
1. A multi-modal and multi-target intelligent identification method for bridge defects under complex background, characterized in that: The method comprises: Acquire an image dataset and a label dataset; wherein the dataset includes a plurality of images of bridges with defects, the bridge images include various textures, lighting conditions, and complex backgrounds, and the label dataset includes labels for identifying various types of bridge damage and key bridge components; Annotating the image dataset based on the label dataset to obtain an annotated dataset; The annotated data set is interpolated, and randomly flipped, denoised, and grayed to obtain an expanded data set; Establishing at least two multi-target detection models, and training the multi-target detection models based on the expanded data set to obtain a plurality of trained multi-target detection models; The multiple trained multi-target detection models are evaluated using evaluation indicators, and the trained multi-target detection model with the best performance is selected as the intelligent multi-target disease detection model. Based on the intelligent disease identification, multi-modal multi-target disease intelligent identification of bridges under complex backgrounds is realized.
2. The multi-modal and multi-target intelligent identification method for bridge defects under complex background according to claim 1 is characterized in that: The annotated data set is interpolated using the following formula: Where L(x) represents the Lanczos kernel function; x represents the relative offset of the sample position; a represents the size of the Lanczos kernel; f(k) represents the source pixel value in the image; L(xk) f(k) represents the Lanczos kernel; L(xk) represents the translated form of the Lanczos kernel function L(x), f(x) represents the interpolated pixel value; k represents the integer position of the pixel in the original image.
3. The multi-modal and multi-target intelligent identification method for bridge defects under complex background according to claim 1 is characterized in that: The at least two multi-target detection models include at least two of a YOLOv8 model, a YOLOv11 model, a RT-DETR model, and a Mamba-YOLO model.
4. The multi-modal and multi-target intelligent identification method for bridge defects under complex background according to claim 3 is characterized in that: The YOLOv8 model includes a C2f module, a neck network part and a detection head part, wherein the C2f module includes three convolution modules and n BottleNeck modules; the detection head part adopts a decoupled head structure; when training the YOLOv8 model, an adaptive data enhancement strategy is adopted to automatically adjust the strength of data enhancement according to the difficulty of target detection.
5. The multi-modal and multi-target intelligent identification method for bridge defects under complex background according to claim 3 is characterized in that: The YOLOv11 model includes a backbone network and a SPFF module; the backbone network includes a C3k2 module and a C2PSA module; the C2PSA module includes a position-sensitive attention model and a C2PASA block; the C2PSA module applies position-sensitive attention and a feedforward network to the input tensor, thereby enhancing feature extraction and processing capabilities, and the SPFF module aggregates multi-scale contextual information through multiple maximum pooling operations with different kernel sizes.
6. The multi-modal and multi-target intelligent identification method for bridge defects under complex background according to claim 3 is characterized in that: The RT-DETR model is based on the DETR model, in which a CORS-based backbone and an efficient hybrid encoder are introduced to achieve real-time speed. The RT-DRET model handles multi-scale features by decoupling scale interactions and cross-scale fusion.
7. The multi-modal and multi-target intelligent identification method for bridge defects under complex background according to claim 3 is characterized in that: The Mamba-YOLO model applies a state-space transition model based on the state-space model to each layer of YOLO to effectively capture global dependencies and utilizes the advantages of local convolution to improve detection accuracy and the model's ability to understand complex scenes while maintaining real-time performance.
8. The multi-modal and multi-target intelligent identification method for bridge defects under complex background according to claim 1 is characterized in that: The evaluation indicators include precision, recall, average precision and F1 score, which are calculated by the following formula: Where recall is the recall rate, that is, the proportion of positive examples predicted to be true to all samples that are actually positive examples, precision is the accuracy, TP is the positive example predicted to be true, FN is the positive example predicted to be false, FP is the negative example predicted to be true, F1-score is the weighted average of precision and recall, IOU is the intersection-overlap ratio, which is used to quantify the proximity of two bounding boxes, area of overlap is the intersection area of the bounding boxes, and area of union is the union area of the two bounding boxes.
9. A multi-modal and multi-target intelligent identification device for bridge defects under complex background, characterized in that: The device comprises: A data acquisition unit is configured to acquire an image data set and a label data set; wherein the data set includes a plurality of images of bridges with diseases, the bridge images include various textures, lighting conditions and complex backgrounds, and the label data set includes labels for identifying various types of bridge damage and key bridge components; a data annotation unit, configured to annotate the image dataset based on the label dataset to obtain an annotated dataset; A data expansion unit is configured to perform interpolation processing on the labeled data set, and perform random flipping, noise addition and grayscale processing to obtain an expanded data set; A model building unit is configured to establish at least two multi-target detection models, and train the multi-target detection models based on the expanded data set to obtain a plurality of trained multi-target detection models; The model screening unit is configured to evaluate the multiple trained multi-target detection models using evaluation indicators, screen out the trained multi-target detection model with the best performance as the intelligent multi-target detection model for defects, and realize multi-modal multi-target intelligent identification of bridge defects under complex backgrounds based on the intelligent identification of defects. 10 . A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, perform the method according to claim 1 .
Citation Information
Patent Citations
Regional bridge risk prediction method and system
CN110807562A
Bridge surface disease automatic identification method
CN114092460A
Multi-label bridge surface defect detection method and system based on deep learning
CN117408947A
Bridge crack disease detection method and device based on crack self-segmentation model
CN117809082A
Small target detection method, system and device in high-altitude overlook scene and medium
CN118968011A
Cited By
Concrete bridge efflorescence defect detection method, device and system and storage medium
CN120235884A
Concrete bridge efflorescence defect detection method, device, system, and storage medium
CN120235884B
Corrosion reinforced concrete crack detection method and device based on YOLOV11
CN120259319A
Railway bridge intelligent inspection method and system based on unmanned aerial vehicle
CN120913109A