A long-span bridge small target identification enhancement method based on super-resolution technology

CN118262232BActive Publication Date: 2026-09-25TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410328827.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2026-09-25
Estimated Expiration
2044-03-21

AI Technical Summary

Technical Problem

然而,受单目相机焦距与景深的限制,对于超大跨径桥梁而言,通常需要采用多目相机和重识别技术来实现全桥范围覆盖,因此,不可避免的存在系统建设成本高,数据传输和处理代价大和系统稳健性弱等问题

Benefits of technology

[0039]1)综合性清晰度评价,通过结合拉普拉斯算子、傅里叶变换和噪声评估,本发明提供了一个全面而准确的图像清晰度评价模型。这种方法能够捕捉图像的多种特征,包括边缘锐度、纹理细节和噪声水平,从而更准确地识别出模糊区域。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118262232B_ABST
    Figure CN118262232B_ABST
Patent Text Reader

Abstract

The application relates to a long-span bridge small target recognition enhancement method based on a super-resolution technology, which comprises the following steps: performing perspective transformation and unfolding on a bridge image collected by a monocular camera, and dividing the image into multiple regions with different sharpness according to actual distances; constructing an image sharpness evaluation model based on a Laplacian operator, a Fourier transform and noise evaluation, evaluating the sharpness of each region, and identifying a fuzzy region; designing a super-resolution image enhancement network to repair the resolution of the identified fuzzy region and improve the image quality; and using a pre-trained YOLOv8 vehicle target detection model to perform vehicle recognition and detection on the repaired image, so that vehicle load monitoring on a bridge with a large field of view is realized. Compared with the prior art, the super-resolution technology is combined with target detection in the application, and is used for vehicle load monitoring on a long-span bridge, so that the detection accuracy is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for identifying small targets on long-span bridges, and more particularly to an enhanced method for identifying small targets on long-span bridges based on super-resolution technology. Background Technology

[0002] With the rapid development of the social economy, highway traffic loads are increasing year by year. This indicates that as highway transportation further develops, bridges built according to current standards will face increasingly higher load-bearing capacity requirements, and their service life will inevitably be significantly reduced due to structural damage caused by the increasing transport loads. Therefore, constructing a vehicle load model based on measured traffic flow loads is of great significance for bridge design verification, safety maintenance, performance evaluation, and operation management.

[0003] However, vehicle movement on bridge decks is highly random, and effective monitoring methods are currently lacking. The mainstream approach is to construct vehicle load models using vehicle load data collected by a Weigh-in-Motion (WIM) system. WIM accurately records key vehicle load information for each vehicle passing through the weighing section, including speed, length, total weight, axle type, number of axles, and axle load. It can also record the lanes traversed by vehicles based on the location of the weighing sensors. With the rapid development of vehicle load monitoring systems integrating WIM and multi-view vision technologies, accurate acquisition of bridge vehicle loads over a large field of view has become crucial for bridge health monitoring. However, due to limitations in focal length and depth of field of monocular cameras, for ultra-long-span bridges, multi-view cameras and re-identification technologies are typically required to achieve full-bridge coverage. This inevitably leads to high system construction costs, high data transmission and processing costs, and weak system robustness. Summary of the Invention

[0004] In order to effectively expand the observation range of vehicle load monitoring and improve the detection range of monocular cameras by using computer vision, thereby reducing the operation and maintenance costs of long-span bridges and reducing data transmission and processing costs, this invention proposes a vehicle load monitoring method for long-span bridges based on super-resolution technology.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] This invention provides a method for enhancing the identification of small targets on long-span bridges based on super-resolution technology, comprising the following steps:

[0007] S1: Perform perspective transformation on the bridge image captured by the monocular camera and divide the image into multiple regions with different clarity according to the actual distance;

[0008] S2: Construct an image sharpness evaluation model based on the Laplacian operator, Fourier transform, and noise assessment to evaluate the sharpness of each region obtained in S1 and identify blurred regions.

[0009] S3: Design a super-resolution image enhancement network to restore the resolution of the blurred areas identified in S2 and improve image quality;

[0010] S4: Using a pre-trained YOLOv8 vehicle target detection model, vehicle recognition and detection are performed on the repaired image in S3 to achieve large-view bridge surface vehicle load monitoring.

[0011] Furthermore, in S1, the perspective transformation unfolding is to map the image captured by the camera to a new perspective through perspective transformation, so that the vehicle target in the image remains consistent in size and shape.

[0012] Furthermore, in S2, the evaluation of image sharpness specifically includes:

[0013] clarity calculation based on the Laplacian operator;

[0014] Detail and texture evaluation based on Fourier transform;

[0015] Image quality assessment based on noise evaluation.

[0016] Furthermore, the clarity evaluation based on the Laplacian operator includes:

[0017] Apply the Laplacian operator: Apply the Laplacian operator to each region of the image acquired in S1;

[0018] Variance calculation: Calculate the variance of the image after processing by the Laplacian operator. The larger the variance, the richer the edge and detail information in the image, and the clearer the corresponding image.

[0019] Furthermore, the Fourier transform-based detail and texture evaluation specifically includes:

[0020] Perform Fourier transform: Perform a Fourier transform on each image region acquired in S1, transforming it from the spatial domain to the frequency domain;

[0021] Analyzing high-frequency components: In the frequency domain, the high-frequency components of an image are analyzed. These high-frequency components are related to the details and texture of the image; that is, the more high-frequency components there are, the clearer the image is.

[0022] Furthermore, the image quality evaluation based on noise assessment specifically includes:

[0023] Image noise assessment: The noise level of each image region is assessed using a pre-selected noise assessment algorithm, which includes one of signal-to-noise ratio (SNR) and peak signal-to-noise ratio (PSNR).

[0024] Combined Sharpness Metrics: Integrating noise levels with sharpness metrics based on the Laplacian operator and Fourier transform provides a more comprehensive evaluation of image sharpness.

[0025] Furthermore, in S3, the super-resolution image enhancement network adopts a residual dense network (RDN) as its basic architecture, and uses dense feature extraction and residual learning mechanisms to restore the resolution of blurred areas;

[0026] The Residual Dense Network (RDN) includes multiple dense blocks and a dense feature fusion function. Each dense block contains multiple convolutional layers and nonlinear activation functions for extracting features from the image.

[0027] The dense feature fusion function fuses the features of different dense blocks to obtain the repaired high-resolution image.

[0028] Furthermore, in S4, the specific process includes:

[0029] Input the repaired image: Use the repaired image from S3 as input to the YOLOv8 vehicle target detection model;

[0030] Image preprocessing: Perform necessary preprocessing operations on the input image to adapt it to the input requirements of the YOLOv8 model;

[0031] Model loading and configuration: Load the pre-trained YOLOv8 vehicle target detection model and configure the corresponding detection parameters;

[0032] Object detection: The preprocessed image is input into the YOLOv8 model for forward propagation calculation to obtain the vehicle's bounding box, class probability, and confidence score;

[0033] Post-processing steps such as confidence thresholding and non-maximum suppression are applied to filter out detection results with low confidence and redundant bounding boxes with high overlap, so as to obtain the final vehicle target detection results.

[0034] Furthermore, in S4, the YOLOv8 vehicle target detection model is trained on a clear image training set, and the YOLOv8 vehicle target detection model is used to identify and detect vehicle targets in the repaired image.

[0035] S3 and S4 also include the process of integrating and cross-training the super-resolution image enhancement network and the YOLOv8 vehicle target detection model. By freezing some network parameters, the training can focus on the efficiency of super-resolution enhancement or object detection, thereby improving the overall performance.

[0036] Furthermore, in S3 and S4, the integration and cross-training processes include:

[0037] The super-resolution image enhancement network is placed as the head in the YOLOv8 framework. First, the input image is processed by super-resolution. Then, global feature fusion and YOLOv8 are applied to detect and classify vehicle targets. The overall performance is optimized by adjusting the network parameters and training strategy to achieve accurate and efficient large field-of-view bridge vehicle load monitoring.

[0038] Compared with the prior art, the present invention has the following technical advantages:

[0039] 1) Comprehensive Sharpness Evaluation: By combining the Laplacian operator, Fourier transform, and noise assessment, this invention provides a comprehensive and accurate image sharpness evaluation model. This method can capture multiple features of an image, including edge sharpness, texture detail, and noise level, thereby more accurately identifying blurred areas.

[0040] 2) Adaptive Region Segmentation and Processing: This invention can divide an image into multiple regions of varying sharpness based on the actual distance to the image, and perform independent sharpness evaluation and processing on each region. This adaptive method ensures appropriate processing of regions with different sharpness, improving the overall image quality.

[0041] 3) Efficient super-resolution restoration: By designing a super-resolution image enhancement network, this invention can efficiently restore the resolution of identified blurred areas. This not only improves the visual effect of the image but also helps in subsequent vehicle target detection and recognition.

[0042] 4) An integrated target detection model: By integrating a super-resolution image enhancement network with the YOLOv8 vehicle target detection model, this invention realizes an end-to-end bridge deck vehicle load monitoring system. This integrated approach simplifies the processing flow and improves detection efficiency and accuracy. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the overall logic of the small target recognition enhancement method for long-span bridges based on super-resolution technology in this invention.

[0044] Figure 2 A YOLOv8 model trained with sharp images for small target detection of bridge vehicles is shown on a test set.

[0045] Figure 3 A schematic diagram showing how an image captured by a monocular camera is unfolded through perspective transformation and divided into four equal segments according to actual distance.

[0046] Figure 4 This is a schematic diagram of the ensemble model, in which RDN is placed as the header within the YOLOv8 framework;

[0047] Figure 5 Example image of a video frame captured by a surveillance camera located in the middle of the crossbeam of the bridge tower of Jiangyin Bridge;

[0048] Figure 6 A schematic diagram of a specific super-resolution (SR) training dataset developed for each part;

[0049] Figure 7 This is a graph showing the peak signal-to-noise ratio (PSNR) values ​​for parts 2 through 4 acquired during training.

[0050] Figure 8 This is a comparison image of the super-resolution processed image in Part 2 and the original image;

[0051] Figure 9 This is a comparison image of the super-resolution processed image in Part 3 and the original image;

[0052] Figure 10 This is a comparison image of the super-resolution processed image in Part 4 and the original image.

[0053] Figure 11 This is a bar chart showing the recall and mAP values ​​for parts 2, 3, and 4 after super-resolution processing. Detailed Implementation

[0054] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. Component models, material names, connection structures, control methods, algorithms, and other features not explicitly described in this technical solution are considered common technical features disclosed in the prior art.

[0055] Example 1

[0056] YOLOv8 model pre-training

[0057] To achieve vehicle identification and tracking, a target detection model is required. This technique involves detecting the effectiveness of region cropping and recognition, thereby assessing the target recognition performance in blurred regions.

[0058] To subsequently test the detection performance of blurred regions and the target detection performance after image restoration, a YOLOv8 model for small target detection of bridge vehicles was first trained using clear images. The training dataset consisted of 4000 images of the Jiangyin Bridge, both nighttime and daytime. The model's performance on the test set after training is shown below. Figure 2 .

[0059] It can be seen that the trained model has achieved good results in terms of accuracy, recall, and map for the identification of various types of vehicles, which prepares the ground for subsequent vehicle identification and detection.

[0060] Cropping of blurred areas and multi-dimensional sharpness evaluation

[0061] Perspective transformation and region division

[0062] First, the image captured by a monocular camera is unfolded using perspective transformation, and then divided into four equal segments according to the actual distance, such as... Figure 3 As shown.

[0063] After the above operations, four images with different resolutions, part1 to part4, were obtained.

[0064] Multi-dimensional clarity evaluation

[0065] Clarity calculation based on Laplace operator

[0066] The Laplacian operator, a second-order derivative operator in mathematics, is commonly used in image processing to detect edges and enhance details. In image processing, the Laplacian operator can quickly react to abrupt changes and edges in an image, thereby extracting high-frequency information. By calculating the value of the Laplacian transform on an image, a scalar value representing image sharpness can be obtained. This method not only provides a clear distinction between blurred and sharp images but also serves as a reference standard for image enhancement, image restoration, and other related applications. The advantages of using a Laplacian-based sharpness calculation method include fast computation speed, no need for complex parameter settings, and good adaptability to various types of images. This approach offers new solutions to many problems in the field of image processing.

[0067] Calculating the Laplacian sharpness of an image typically involves applying a Laplacian filter to the image and calculating its transformed variance. Higher variance generally indicates that the image has more edges and higher sharpness. Below is the pseudocode and related formula for calculating Laplacian sharpness:

[0068] Calculating the Laplacian sharpness of an image typically involves applying a Laplacian filter to the image and calculating its transformed variance. Higher variance generally indicates that the image has more edges and higher sharpness. Below is the pseudocode and related formula for calculating Laplacian sharpness:

[0069] Given an image R after Laplacian transform, its sharpness S can be calculated using the following equation:

[0070]

[0071] in:

[0072] N is the total number of pixels in the image.

[0073] R i It is the pixel value at position i in the image.

[0074] It is the average pixel value of the image.

[0075]

[0076]

[0077] Detail and texture evaluation

[0078] More texture and detail can facilitate object detection and tracking. Evaluating the texture and detail of a computed image by leveraging noise is also one of the indicators for evaluating image quality.

[0079] The Fourier transform shifts an image from the spatial domain to the frequency domain, revealing details about its sharpness.

[0080] The two-dimensional discrete Fourier transform (DFT) of a digital image is defined as follows:

[0081]

[0082] Here, M and N are the dimensions of the image, and the summation is performed over all pixels. The high-frequency components in the DFT spectrum are related to image edges and details, and are indicators of sharpness. A dominance of high-frequency components indicates a sharper image, while a lack of these frequencies suggests a blurry image.

[0083] The pseudocode is as follows:

[0084]

[0085] Super-resolution enhancement network repairs blurry areas

[0086] Residual Dense Networks for Image Super-Resolution

[0087] RDN was chosen as the foundation for super-resolution networks due to its outstanding ability to reconstruct high-resolution images from low-resolution images, primarily because of its dense feature extraction and residual learning mechanisms. This architecture excels at preserving image details, which is crucial for subsequent detection processes.

[0088] The core principle of RDN is embodied in the dense feature fusion strategy, which can be formalized as follows:

[0089] F DF (x)=H DF ([F D1 ,F D2 ,...,F DD ])

[0090] Where F Di Let H represent the dense block function, where D represents the number of dense blocks, x is the input feature map, and H is the dense block function. DF This represents a dense feature fusion function.

[0091] Global feature fusion, defined as F GF This integrates the original image features with the dense features:

[0092] F GF (x)=H GF (F DF (x),F0(x))

[0093] F0 is the initial shallow features extracted from the input image x.

[0094] Data preparation:

[0095] A large amount of paired image data is collected, consisting of low-resolution images and their corresponding high-resolution images. This data is used to train the network to learn the mapping relationship from low resolution to high resolution.

[0096] The images undergo necessary preprocessing, such as cropping, scaling, and normalization, to make them conform to the network's input requirements.

[0097] The RDN construction process includes:

[0098] Network Construction: Design the structure of the super-resolution image enhancement network, which typically includes a feature extraction layer, a nonlinear mapping layer, and an image reconstruction layer. The feature extraction layer is responsible for extracting features from the low-resolution image, the nonlinear mapping layer learns to map these features to the high-resolution space, and the image reconstruction layer is responsible for generating the final high-resolution image.

[0099] The weight parameters of the network are typically initialized using random initialization or a pre-trained model.

[0100] Loss function design: Define a loss function to measure the difference between the high-resolution image output by the network and the real high-resolution image. Commonly used loss functions include mean squared error (MSE) and structural similarity index (SSIM). Custom loss functions can be designed according to specific needs to better capture the details and texture information of the image.

[0101] Algorithm selection: Choose a suitable optimization algorithm to update the network's weight parameters to minimize the loss function. Commonly used optimization algorithms include stochastic gradient descent (SGD) and Adam. Hyperparameters such as learning rate and batch size are set to control the optimization process.

[0102] Iterative optimization during training: Pairs of low-resolution and high-resolution images are input into the network, and forward propagation is performed to obtain the output high-resolution image.

[0103] Calculate the loss function value and update the network's weight parameters using the backpropagation algorithm. Repeat the above steps for multiple epochs until the network converges or reaches a preset stopping condition.

[0104] During training, a validation set can be used to monitor network performance and make adjustments as needed, such as adjusting the learning rate or applying regularization techniques.

[0105] Model Evaluation and Selection: The performance of the trained super-resolution image augmentation network is evaluated using a test set. Evaluation metrics may include Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), etc.

[0106] Based on the evaluation results, the model with the best performance is selected as the final super-resolution image enhancement network.

[0107] YOLOv8-based vehicle target detection model

[0108] YOLOv8 is used as the object detection model in this invention due to its fast and accurate detection capabilities. The model works by segmenting the image into a grid and predicting bounding boxes and class probabilities for each grid cell.

[0109] The key operations of YOLOv8 can be summarized by the following formula:

[0110] O(x) = σ(Conv(F(x)))

[0111] Where F(x) represents the feature extraction function, Conv represents the convolution operation, and σ is the Sigmoid function used for probability prediction.

[0112] ensemble model

[0113] The ensemble model strategically places RDN as a header within the YOLOv8 framework, such as... Figure 4 As shown in the diagram, this design allows the system to first perform super-resolution processing on the input image, focusing on enhancing edges and textures, which is crucial for accurate object detection.

[0114] Combinatorial operations can be represented as:

[0115] y(x)=YOLOv8(F GF (RDN(x)))

[0116] Where RDN(x) enhances the input image x, then F GF Global feature fusion is applied, while YOLOv8 performs detection and classification.

[0117] This model can be cross-trained, and the RDN or YOLOv8 parts can be frozen by adjusting the parameters. This helps to focus training on super-resolution enhancement or object detection efficiency.

[0118] Application Example 1

[0119] Enhanced Detection Verification of Small Targets on Long-Span Bridges—An Experiment Based on Data from the Jiangyin Bridge

[0120] Data and Test Area

[0121] A surveillance camera, located in the middle of the crossbeam of the Jiangyin Bridge tower, is shooting 1920x1080 pixel video at a height of 25 frames per second from 85 meters. (See attached image) Figure 5 The recorded area spanned 192 meters by 29.5 meters. After recording, the video was cropped to 295 x 1920 pixels using 12 ground control points (GCPs), with each pixel representing 10 true ground diameters. Subsequently, a vehicle benchmark dataset was created for detector training, containing 4000 video frames labeled with three vehicle types.

[0122] Vehicle detection based on YOLOv8

[0123] The dataset was divided into 80% for training, 10% for validation, and 10% for testing. Different network architectures were trained on this dataset using a batch size of 32 and an initial learning rate of 0.004. Data augmentation included random horizontal flipping and cropping. The network was initialized using weights pre-trained on the COCO dataset and trained on an NVIDIA GeForce RTX 2080Ti GPU via the Tensorflow ObjectDetection API. Model performance was evaluated using the mAP score at an IoU threshold, defined as:

[0124]

[0125] Where A and B are the areas of the real object and the detected object, respectively. Precision and recall are calculated using the formulas... The values ​​are calculated using TP, FP, and FN, where TP, FP, and FN represent true positives, false positives, and false negatives, respectively. The intersection-over-union ratio (IoU) and confidence threshold are both set to 0.5.

[0126]

[0127] Creating a pairing dataset

[0128] I. Segmented Testing

[0129] Two hundred frames of images of the Jiangyin Bridge were labeled and divided into four parts according to the description above. Each part...

[0130] Part 1 415.79 Part 2 206.43 Part 3 49.88 Part 4 14.76

[0131] The Laplacian sharpness and detail texture assessment is as follows:

[0132] Part 1 20.19 Part 2 17.22 Part 3 16.33 Part 4 14.92

[0133] Next, vehicle target detection was performed using the YOLOv8 model for each part. Results showed that detection accuracy decreased with increasing distance from the camera, which is related to changes in blurriness across parts. To address the image quality degradation caused by camera depth of field, a specific super-resolution (SR) training dataset was developed for each part, including high-resolution images similar to Part 1 and low-resolution corresponding images for Parts 2 through 4. See [link to dataset]. Figure 6 The trained SR network aims to uniformly improve image sharpness to the level of Part 1.

[0134] To accurately simulate the real-world blurring effects in the datasets from Parts 2 to 4, the sharpness assessment method described earlier was used. The downsampling process involved interpolation and various blurring techniques, with parameters finely tuned to reflect actual image conditions. This resulted in three sets of high-resolution and low-resolution image pairs for each part, effectively demonstrating the quality improvements within each part.

[0135] II. Super-resolution enhancement

[0136] Using the generated matching dataset, a Residual Dense Network (RDN) was trained on three different parts of the image. Training settings: learning rate 1×10⁻⁶. -4 The batch size was 32, the image segmentation patch size was 32, and the total number of epochs was 400. During training, the peak signal-to-noise ratio (PSNR) values ​​for parts 2 through 4 were obtained, such as... Figure 7 As shown.

[0137] PSNR (Peak Signal-to-Noise Ratio) is a key metric in image processing. It quantifies the quality of super-resolution reconstruction by comparing the quality of the reconstructed image to the original high-quality image (the ground truth). A PSNR value exceeding 30 for all three reconstructed parts indicates successful image enhancement.

[0138] III. Conclusion

[0139] After the new model was trained, it was used for image restoration and detection, and the results are as follows.

[0140] For Part 2, which originally had relatively high clarity, recall and mAP50 showed marginal improvements. However, a significant improvement was observed at mAP50-95, see [link to relevant documentation]. Figure 11 This indicates that the super-resolution reconstruction in this part leads to a higher probability of accurate recognition, mainly due to the refinement of image details; see [link to relevant documentation]. Figure 8 .

[0141] For Part 3, which has lower resolution, improvements were primarily observed in recall and mAP50. See [link to relevant documentation]. Figure 11 Post-processing using the network of this invention significantly reduces instances of false and duplicate detections. This improvement can be attributed to the super-resolution network's ability to recover vehicle textures and finer details, thereby improving overall detection accuracy. See [link to relevant documentation]. Figure 9 .

[0142] For the fourth section, which has the lowest resolution, the main improvement lies in mAP50, see [link / reference]. Figure 11 Super-resolution processing significantly improves the detection of vehicles that were previously undetectable, likely because super-resolution networks make vehicles more detectable by enhancing edges. See [link to documentation]. Figure 10 .

[0143] This invention introduces a novel method combining super-resolution technology with object detection for vehicle load monitoring on long-span bridges, significantly improving detection accuracy. Experiments show that this method is highly effective in extending the monitoring capabilities of single-camera systems. The results indicate potential for future improvements in traffic management and bridge monitoring. Future work could explore algorithms combining super-resolution with object detection and analyze continuous video data to further enhance monitoring accuracy and scope.

[0144] The above description of the embodiments is provided to enable those skilled in the art to understand and use the invention. It will be apparent to those skilled in the art that various modifications can be made to these embodiments, and the general principles described herein can be applied to other embodiments without inventive effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope of the invention should be within the protection scope of the present invention.

Claims

1. A method for enhancing the identification of small targets in long-span bridges based on super-resolution technology, characterized in that, Includes the following steps: S1: Perform perspective transformation on the bridge image captured by the monocular camera and divide the image into multiple regions with different clarity according to the actual distance; S2: Construct an image sharpness evaluation model based on the Laplacian operator, Fourier transform, and noise assessment to evaluate the sharpness of each region obtained in S1 and identify blurred regions. S3: Design a super-resolution image enhancement network to restore the resolution of the blurred areas identified in S2 and improve image quality; S4: Using a pre-trained YOLOv8 vehicle target detection model, vehicle recognition and detection are performed on the repaired image in S3 to achieve large field-of-view bridge surface vehicle load monitoring. In S2, the evaluation of image sharpness specifically includes: clarity calculation based on the Laplacian operator; Detail and texture evaluation based on Fourier transform; Image quality assessment based on noise evaluation; In S3, the super-resolution image enhancement network uses a residual dense network (RDN) as its basic architecture, and repairs the resolution of blurred areas through dense feature extraction and residual learning mechanisms. The Residual Dense Network (RDN) includes multiple dense blocks and a dense feature fusion function. Each dense block contains multiple convolutional layers and nonlinear activation functions for extracting features from the image. The dense feature fusion function fuses the features of different dense blocks to obtain the repaired high-resolution image.

2. The method for enhancing small target recognition in long-span bridges based on super-resolution technology according to claim 1, characterized in that, In S1, the perspective transformation unfolding is to map the image captured by the camera to a new perspective through perspective transformation, so that the vehicle target in the image remains consistent in size and shape.

3. The method for enhancing small target recognition in long-span bridges based on super-resolution technology according to claim 1, characterized in that, The Laplacian operator-based clarity evaluation includes: Apply the Laplacian operator: Apply the Laplacian operator to each region of the image acquired in S1; Variance calculation: Calculate the variance of the image after processing by the Laplacian operator. The larger the variance, the richer the edge and detail information in the image, and the clearer the corresponding image.

4. The method for enhancing small target recognition in long-span bridges based on super-resolution technology according to claim 1, characterized in that, The Fourier transform-based detail and texture evaluation specifically includes: Perform Fourier transform: Perform a Fourier transform on each image region acquired in S1, transforming it from the spatial domain to the frequency domain; Analyzing high-frequency components: In the frequency domain, the high-frequency components of an image are analyzed. These high-frequency components are related to the details and texture of the image; that is, the more high-frequency components there are, the clearer the image is.

5. The method for enhancing small target recognition in long-span bridges based on super-resolution technology according to claim 1, characterized in that, The image quality evaluation based on noise assessment specifically includes: Image noise assessment: The noise level of each image region is assessed using a pre-selected noise assessment algorithm, which includes one of signal-to-noise ratio (SNR) and peak signal-to-noise ratio (PSNR). Combined Sharpness Metrics: Integrating noise levels with sharpness metrics based on the Laplacian operator and Fourier transform provides a more comprehensive evaluation of image sharpness.

6. The method for enhancing small target recognition in long-span bridges based on super-resolution technology according to claim 1, characterized in that, In S4, the specific process includes: Input the repaired image: Use the repaired image from S3 as input to the YOLOv8 vehicle target detection model; Image preprocessing: Perform necessary preprocessing operations on the input image to adapt it to the input requirements of the YOLOv8 model; Model loading and configuration: Load the pre-trained YOLOv8 vehicle target detection model and configure the corresponding detection parameters; Object detection: The preprocessed image is input into the YOLOv8 model for forward propagation calculation to obtain the vehicle's bounding box, class probability, and confidence score; After applying confidence thresholding and non-maximum suppression post-processing steps, low-confidence detection results and redundant bounding boxes with high overlap are filtered out to obtain the final vehicle target detection results.

7. The method for enhancing small target recognition in long-span bridges based on super-resolution technology according to claim 1, characterized in that, In S4, the YOLOv8 vehicle target detection model is trained on a clear image training set, and the YOLOv8 vehicle target detection model is used to identify and detect vehicle targets in the repaired image. S3 and S4 also include the process of integrating and cross-training the super-resolution image enhancement network and the YOLOv8 vehicle target detection model. By freezing some network parameters, the training can focus on the efficiency of super-resolution enhancement or object detection, thereby improving the overall performance.

8. The method for enhancing small target recognition in long-span bridges based on super-resolution technology according to claim 7, characterized in that, In S3 and S4, the integration and cross-training process includes: The super-resolution image enhancement network is placed as the head in the YOLOv8 framework. First, the input image is processed by super-resolution. Then, global feature fusion and YOLOv8 are applied to detect and classify vehicle targets. The overall performance is optimized by adjusting the network parameters and training strategy to achieve accurate and efficient large field-of-view bridge vehicle load monitoring.

Citation Information

Patent Citations

  • PET-MRI image fusion method based on adaptive generative adversarial network

    CN115457359A

  • Method for detecting weak and small leakage target of sealing element based on low-resolution infrared image

    CN115994893A