Intelligent bridge identification method and system based on high-resolution remote sensing image
Through the GLDViT model combining the bridge recognition method of global and local features, the problems of limited feature extraction and low processing efficiency in bridge recognition are solved, and high-precision and stable bridge recognition effect are achieved.
Patent Information
- Application Number
- CN202510367892.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has limited feature extraction and low processing efficiency in bridge recognition, and cannot achieve ideal accuracy and speed in large-scale bridge detection, and traditional methods perform unstable in complex environments.
The GLDViT model is adopted, combining the advantages of global and local characteristics, and paying attention to global information and local details through the self-attention mechanism, and high-resolution remote sensing images are used to identify bridge structures and potential damage.
It realizes high accuracy and robustness of bridge recognition, can maintain stable performance in complex environments, improves the efficiency and recognition accuracy of large-scale remote sensing data processing, and reduces manual intervention.
Smart Images

Figure CN120298892A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing and intelligent recognition, and particularly relates to a method and system for intelligent recognition of bridges based on high-resolution remote sensing images. Background Art
[0002] As an indispensable part of the modern transportation system, bridges play a crucial role in connecting different regions and spanning natural or artificial obstacles. Globally, the total length of bridges for high-speed railways, highways, and urban rails has reached millions of kilometers. Especially in China, the total length of highway bridges exceeds 9.5 million kilometers, and the total length of railway bridges is approximately 40,000 kilometers. With the continuous growth of transportation demand, these data will continue to increase. With the advancement of urbanization and continuous investment in infrastructure, bridges are not only an important guarantee for traffic flow but also a key support for economic activities and social development. However, bridges face severe challenges during long-term operation, including corrosion by the natural environment, increased traffic loads, and aging problems. Many old bridges face potential risks due to structural damage or material fatigue, seriously affecting traffic safety. Therefore, carrying out regular monitoring and health assessment of bridges to ensure their structural safety and proper function is an urgent need in current traffic infrastructure maintenance.
[0003] Traditional bridge recognition methods mainly rely on manual inspections and basic image processing techniques. Although manual inspections can delve into every detail and conduct targeted checks, this method is time-consuming, has a limited coverage area, and is easily affected by human factors, making it unable to efficiently process large-scale bridge data. In addition, the inspection of bridges by manual inspections mainly relies on the experience of inspectors, and some minor damages or hidden problems may be missed, resulting in the failure to timely detect potential safety hazards. In the application of remote sensing images, traditional image processing methods usually use techniques such as edge detection, image segmentation, and morphological operations. Although the basic features of bridges can be extracted, due to the uneven quality of remote sensing images and the complex structure of bridges, traditional methods often perform unstably in complex scenarios. Especially under the interference of environmental factors such as changes in lighting, differences in viewing angles, and weather effects, the effectiveness of traditional methods is greatly reduced. In addition, these methods also face problems such as limited feature extraction and low processing efficiency, and cannot achieve ideal accuracy and speed in large-scale bridge detection.
[0004] With the development of deep learning technology, convolutional neural networks (CNNs) and vision transformers (ViTs) have made significant progress in the field of image processing. Deep learning can automatically extract high-order features from images, avoiding the complexity of manually designing features in traditional methods and significantly improving the accuracy and efficiency of image recognition. However, although CNNs perform excellently in local feature extraction, due to the limitation of the local receptive field, it is difficult to effectively capture the global context information in images. Especially when dealing with remote sensing images with complex backgrounds and long-range dependencies, its recognition ability has limitations. ViT processes global information through the self-attention mechanism, but when dealing with details and local features, it is often less accurate than CNNs and is prone to losing details or misidentifying. Summary of the Invention
[0005] To solve the problems existing in the above-mentioned prior art, the present invention provides a bridge intelligent recognition method and system based on high-resolution remote sensing images, solving the technical problems of limited feature extraction, low processing efficiency, and inability to achieve ideal accuracy and speed in large-scale bridge detection in the prior art.
[0006] In view of the deficiencies of the prior art, this patent proposes the GLDViT (Global and Local Diffusion Vision Transformer) model, which combines the advantages of global and local features. GLDViT simultaneously focuses on global information and local details through the self-attention mechanism, and can effectively identify bridge structures and potential damages in remote sensing images. Compared with traditional CNN and ViT models, GLDViT has higher accuracy and robustness when dealing with complex remote sensing images. Especially in environments such as lighting changes and perspective differences, it can still maintain stable performance. This model is suitable for processing large-scale and high-frequency remote sensing data, providing a more accurate and efficient solution for bridge intelligent recognition.
[0007] Compared with traditional manual interpretation methods, deep learning models based on computer vision can greatly improve the recognition efficiency, reduce the dependence on human resources, and save time and energy. By obtaining data through high-resolution optical and SAR images, more comprehensive bridge feature information can be provided in complex environments. Training with the GLDViT model, which combines global and local information, effectively improves the accuracy and robustness of bridge recognition. This method not only realizes the automatic recognition and extraction of bridges in transportation corridors, but also can process large-scale remote sensing data, greatly improving the monitoring efficiency and reducing manual intervention.
[0008] A bridge intelligent recognition method based on high-resolution remote sensing images includes:
[0009] Step S1: Select high-resolution remote sensing images, including optical and SAR remote sensing images, and select the area around the bridge for sample data collection. Through expert identification and visual interpretation, manually delineate the sample area of the bridge and label the non-bridge part as the background value.
[0010] Step S2: Since most bridges are large-scale buildings, the method of segmentation-prediction is used to calculate the bridge images, and the segmented images are enhanced to obtain the spatial position and distribution of the bridges in the remote sensing images.
[0011] Step S3: Set the initial parameters and training parameters of the GLDViT deep learning model, use the bridge remote sensing images and corresponding sample data obtained in Step 1, and perform backpropagation training using the Structural Perceptual Loss function.
[0012] Step S4: Randomly select remote sensing images of traffic corridors containing bridges for identification and verify the accuracy of the identification results.
[0013] Furthermore, the processing of the remote sensing images by the deep learning model GLDViT includes: using the Transformer module to obtain the global information of the images, then using the sparse attention mechanism to screen the feature maps, calculating the importance score of each local area through the diffusion matrix, and selecting the sparse local area most relevant to the task as the local feature. The sparse prompt information is as follows:
[0014]
[0015]
[0016] In the formula, i, j, h are the time step, denoising layer, and cross-attention head in the diffusion model respectively. is the diffusion model, I is the feature image, and w is the cross-attention mapping method.
[0017] In the formula, Q represents the target position, K represents the content of the area, d is the feature scaling factor, softmax is a normalization function, and Topk is a selection mechanism that only retains the local area with the highest attention score.
[0018] Combine the local features and local information through linear projection and a fixed-size feature vector, and then use Random drop to discard a part of the feature vectors.
[0019] Use the Softmax recognizer to calculate the category in the fully connected layer, and further optimize the positioning accuracy through bounding box regression to obtain a more accurate target boundary.
[0020] Further, the specific steps of step S1 are as follows:
[0021] Step S11: Based on high-resolution optical satellite remote sensing images, use the software ArcMap to outline the bridge contour in the images.
[0022] Step S12: Set the bridge code to 1 and other features and backgrounds to 0. Save as label data as a TIFF raster image with the same spatial resolution as the remote sensing image.
[0023] Step S13: Calculate the vegetation index and the backscattering coefficient of the SAR image, and stack them into the feature image of the deep learning model in the format of TIFF raster data.
[0024] Further, the specific steps of step S2 are as follows:
[0025] Step S21: To further improve the accuracy of the model, before inputting the remote sensing image into the model, first perform preprocessing and image enhancement operations on the image data. This includes image rotation, image flipping, noise removal, and multi-spectral band fusion processing on the original remote sensing image, so as to improve the clarity and detail expressiveness of the image, and at the same time enhance the sensitivity of the model to target features.
[0026] Step S22: To meet the requirements of model input, crop the image data into blocks of a fixed size, with a resolution of 512×512 pixels for each block. During the cropping process, the overlapping sliding window technique is adopted to avoid truncating or missing key target features. The cropped image data is saved as a JPEG format file, and at the same time, the corresponding label image is also cropped into blocks of the same size to ensure that the model can accurately match the input features with the labels during the training process.
[0027] Further, step S3 includes: putting the original image into the GLDViT model to extract feature information, where the loss function of Structural Perceptual Loss is:
[0028]
[0029] In the formula, φ(.) is the feature extractor of the deep learning model, Calculate the prediction result and the Euclidean distance between the true label in the feature space. ▽ represents the edge or gradient information of the image, which is used to measure the gradient consistency between the prediction result and the true value, capture the edge information and improve the detail performance.
[0030] Further, the specific steps of step S4 are as follows:
[0031] Step S41: Randomly select an image containing a bridge as the test data and input it into the model to perform intelligent identification of the bridge, and at the same time evaluate the confidence of the identification result to verify the identification performance and accuracy of the model.
[0032] Step S42: Considering the impact of different types of bridges on identification and monitoring, input the remote sensing images of different types of bridges into the model to test the generalization ability of the model under different structural and environmental changes.
[0033] A bridge intelligent identification system based on high-resolution remote sensing images includes: a data input module, an identification module, and a result output module. The data input module is used to input remote sensing images containing bridges. The identification module is equipped with a trained deep learning model GLDViT. The trained deep learning model GLDViT identifies the remote sensing images. The result output module is used to output the remote sensing images after identifying the bridges.
[0034] The beneficial effects of the present invention include:
[0035] 1. This solution adopts a method that combines high-resolution optical images and SAR images. Through the training of the GLDViT deep learning model, it can realize the automatic identification and positioning of bridges, greatly improving the efficiency of bridge identification, and at the same time ensuring the accuracy of target detection.
[0036] 2. A bridge intelligent identification method based on the fusion of global and local features is constructed. This solution screens the feature regions through a sparse attention mechanism and combines the global context information, enabling the model to effectively capture the key region features, thereby greatly improving the identification accuracy.
[0037] 3. In the image tests of different types of bridges, the model has good robustness to complex backgrounds and various environmental changes. The identification results are very stable and the accuracy differences are small, demonstrating strong generalization ability. Description of the Drawings
[0038] Figure 1 It is a flowchart of a bridge intelligent identification method according to an embodiment of the present application.
[0039] Figure 2 It is a backbone network structure diagram of the GLDViT model according to an embodiment of the present application.
[0040] Figure 3 It is a drawn bridge raster map according to an embodiment of the present application.
[0041] Figure 4 It is a bridge identification result map according to an embodiment of the present application. Detailed Embodiments
[0042] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Therefore, the detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents the selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0043] Embodiment 1
[0044] A bridge intelligent recognition method based on high-resolution remote sensing images, as Figure 1 shown, includes the following steps:
[0045] Step S1: Select high-resolution optical and SAR remote sensing images, and select an area within about 500 meters around the bridge for sample data collection. Through expert interpretation and visual interpretation, manually delineate the sample area of the bridge and label the non-bridge part as the background value.
[0046] Step S2: Since most bridges are large-scale buildings, a segmentation-prediction method is used to calculate the bridge images, and image enhancement is performed on the segmented images to obtain the spatial positions and distributions of the bridges in the remote sensing images.
[0047] Step S3: Set the initial parameters and training parameters of the GLDViT deep learning model. Use the bridge remote sensing images and corresponding sample data obtained in Step 1, and perform backpropagation training using the Structural Perceptual Loss function.
[0048] Step S4: Randomly select traffic corridor remote sensing images containing bridges for recognition, and verify the accuracy of the recognition results.
[0049] The specific steps of Step S1 are as follows:
[0050] Step S11: Based on high-resolution optical satellite remote sensing images, use the software ArcMap to outline the bridge contour in the images;
[0051] Step S12: Set the bridge code to 1 and other features and the background to 0. Save as a label data TIFF raster image with the same spatial resolution as the remote sensing image; specifically as Figure 3 shown, where the white code of 1 is the bridge and the black code of 0 is the background;
[0052] Step S13: Calculate the vegetation index and the backscattering coefficient of the SAR image, and stack them into a feature image for the deep learning model in the format of TIFF raster data.
[0053] The specific steps of step S2 are as follows:
[0054] Step S21: To further improve the accuracy of the model, before inputting the remote sensing image into the model, the image data is first preprocessed and image enhancement operations are performed. This includes image rotation, image flipping, noise removal, and multi-spectral band fusion processing on the original remote sensing image, so as to improve the clarity and detail expressiveness of the image, and at the same time enhance the sensitivity of the model to target features.
[0055] Step S22: To meet the requirements of the model input, the image data is cropped into blocks of a fixed size, and the resolution of each block is 512×512 pixels. The overlapping sliding window technique is used during the cropping process to avoid truncating or missing key target features. The cropped image data is saved as a JPEG format file, and at the same time, the corresponding label image is also cropped into blocks of the same size to ensure that the model can accurately match the input features with the labels during the training process.
[0056] The specific steps of step S3 are as follows:
[0057] Step S31: Put the original image into the GLDViT model to extract feature information, and the loss function of StructuralPerceptual Loss is:
[0058]
[0059] In the formula, φ(.) is the feature extractor of the deep learning model, Calculates the prediction result and the Euclidean distance between the true label in the feature space. ▽ represents the edge or gradient information of the image, which is used to measure the gradient consistency between the prediction result and the true value, capture edge information, and improve detail performance.
[0060] Step S32: To improve the classification accuracy of the model, a new deep learning model (GLDViT) is developed by combining global and local feature information and a diffusion model. After the feature image is input into the model, the Transformer module is used to obtain the global information of the image. Then, the sparse attention mechanism is used in it to screen the feature map, calculate the importance score of each local area through the diffusion matrix, and select the sparse local area most relevant to the task as the local feature. The sparse extraction information is as follows:
[0061]
[0062] In formula (2), i, j, and h are the time step, denoising layer, and cross-attention head in the diffusion model, respectively. is the diffusion model, I is the feature image, and w is the cross-attention mapping method.
[0063] In formula (3), Q represents the target position, K represents the content of the region, d is the feature scaling factor, softmax is a normalization function, and Topk is a selection mechanism that only retains the local region with the highest attention score.
[0064] Step S33: Combine local features and local information through linear projection and obtain a feature vector of a fixed size. Then, use Random drop to discard a part of the feature vectors.
[0065] Step S34: Use a Softmax recognizer in the fully connected layer to calculate the category, and further optimize the localization accuracy through bounding box regression to obtain a more accurate target boundary.
[0066] Specifically, as Figure 2 shown, the flowchart shows a deep learning model - GLDViT that combines global and local feature information and a diffusion model. First, the model receives training data (annotated images) and performs feature encoding on the original images. Then, the Transformer module in the attention pooling module is used to extract global information, which can capture the overall context information of the image and reduce the spatial dimension of the feature map, thereby reducing the computational amount and controlling overfitting. Then, local features are obtained through the convolutional layer, and the feature map is screened through the sparse attention mechanism to help the model focus on important regions in the image and reduce the computational amount. Next, the importance score of each local region is calculated through the diffusion matrix, and the sparse local regions most relevant to the task are selected as local features. These local features will be further optimized through the multi-scale local loss to improve the classification accuracy. In addition, the model also uses the Prompt Dropout technique to optimize the global and local information. As shown in Table 1, by randomly discarding the global and local information, the remaining information is forced to develop independently and express different semantic information, and finally, the global loss of the training data and labels is fused to ensure the consistency of the global information. Finally, the processed feature data is input into the classifier to complete the final classification task.
[0067] Table 1 Pseudo-code of Prompt Dropout
[0068]
[0069] The specific steps of step S4 are as follows:
[0070] Step S41: Randomly select an image containing a bridge as the test data and input it into the model to perform intelligent identification of the bridge. At the same time, evaluate the confidence of the identification result to verify the identification performance and accuracy of the model.
[0071] Step S42: Considering the influence of different types of bridges on identification and monitoring, input the remote sensing images of different types of bridges into the model to test the generalization ability of the model under different structural and environmental changes.
[0072] Specifically, the bridge identification results are as Figure 4 shown, including the original image, the intelligent interpretation result, and the final bridge identification result. The recall rate reaches 94.75%, and the IOU reaches 87.21%. Among them, compared with the existing DeepLabV3+ and Unet, the accuracy of the recall rate has increased by 6.81%-8.6%.
[0073] Embodiment 2
[0074] A bridge intelligent identification system based on high-resolution remote sensing images, comprising: a data input module, an identification module, and a result output module. The data input module is used to input remote sensing images containing bridges. The identification module is provided with the trained deep learning model GLDViT involved in Embodiment 1. The trained deep learning model GLDViT identifies the remote sensing images. The result output module is used to output the remote sensing images after identifying the bridges.
[0075] The above embodiments only represent the specific implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation to the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A bridge intelligent recognition method based on high-resolution remote sensing images, characterized in that Including: Step S1: Select high-resolution remote sensing images, select the bridge area for sample data collection, determine the bridge part, and label the non-bridge part as the background value as sample data; Step S2: Calculate the bridge image through the method of segmentation prediction, and perform image enhancement on the segmented image to obtain the spatial position and distribution of the bridge in the remote sensing image; Step S3: Construct a deep learning model GLDViT and set the initial parameters and training parameters, and use the bridge remote sensing image and the corresponding sample data obtained in Step 1 for backpropagation training; Step S4: Randomly select remote sensing images containing bridges, use the trained deep learning model GLDViT for recognition, and verify the accuracy of the recognition results.
2. The method for intelligent identification of bridges based on high-resolution remote sensing images according to claim 1, wherein The processing of the remote sensing image by the deep learning model GLDViT includes: using the Transformer module to obtain the global information of the image, then using the sparse attention mechanism to screen the feature map, calculating the importance score of each local area through the diffusion matrix, and selecting the sparse local area most relevant to the task as the local feature. The sparse prompt information is as follows: Wherein, i, j, and h are the time step, denoising layer, and cross-attention head in the diffusion model, respectively, is the diffusion model, I is the feature image, and w is the cross-attention mapping method; In the formula, Q represents the target position, K represents the content of the area, d is the feature scaling factor, softmax is a normalization function, and Topk is a selection mechanism that only retains the local area with the highest attention score; Combine the local feature and local information through linear projection and a fixed-size feature vector, and then use Randomdrop to discard a part of the feature vectors; Use the Softmax recognizer to calculate the category in the fully connected layer, and further optimize the positioning accuracy through bounding box regression to obtain a more accurate target boundary.
3. The bridge intelligent recognition method based on high - resolution remote sensing images according to claim 1, wherein The specific steps of the said Step S1 are: Step S11: Based on the high-resolution optical satellite remote sensing image, use the software ArcMap to outline the bridge contour in the image; Step S12: Set the bridge code to 1, and set other ground objects and the background to 0, store it as a TIFF raster image, and the spatial resolution is the same as that of the remote sensing image; Step S13: Calculate the vegetation index and the backscattering coefficient of the SAR image, and stack them as the feature image of the deep learning model, and store it as TIFF raster data.
4. The method for intelligent bridge recognition based on high-resolution remote sensing images according to claim 1, characterized in that, The specific steps of the said Step S2 are: Step S21: Before inputting the remote sensing image into the deep learning model GLDViT, first perform preprocessing and image enhancement operations on the image data, including image rotation, image flipping, noise removal, and multi-spectral band fusion processing on the original remote sensing image; Step S22: Then use the overlapping sliding window to crop the image data into fixed-size blocks, save the cropped image data as a JPEG format file, and at the same time synchronously crop the corresponding label image into blocks of the same size.
5. The bridge intelligent recognition method based on high-resolution remote sensing images according to claim 1, wherein The said Step S3 includes: putting the original image into the GLDViT model to extract the feature information, and using the loss function of StructuralPerceptual Loss for backpropagation training. The formula is as follows: Where φ(.) is the feature extractor of the deep learning model, The predicted result is calculated and the Euclidean distance between the predicted result and the ground truth in the feature space; represents the edge or gradient information of the image, which is used to measure the gradient consistency between the predicted result and the ground truth, capture edge information, and improve the detail performance.
6. The intelligent bridge recognition method based on high-resolution remote sensing images according to claim 1, characterized in that The specific steps of the said Step S4 are: Step S41: Randomly select a remote sensing image containing a bridge as the test data and input it into the model to perform intelligent identification of the bridge. At the same time, evaluate the confidence level of the identification result to verify the identification performance and accuracy of the model. Step S42: Considering the impact of different types of bridges on identification and monitoring, input the remote sensing images of different types of bridges into the model to test the generalization ability of the model under different structural and environmental changes.
7. A bridge intelligent recognition system based on high-resolution remote sensing images, characterized in that, It includes: A data input module, an identification module, and a result output module. The data input module is used to input the remote sensing image containing the bridge. The identification module is provided with the trained deep learning model GLDViT described in any one of claims 1-6. The trained deep learning model GLDViT identifies the remote sensing image. The result output module is used to output the remote sensing image after identifying the bridge.