YOLOv8 remote sensing image vehicle target detection method and system based on multi-scale attention mechanism

By introducing a multi-scale attention mechanism module into the YOLOv8 network, the problem of small target information loss caused by convolution operations is solved, and the vehicle target detection accuracy is improved in complex contexts, achieving more accurate identification of small targets and irregular angle targets.

CN120182579APending Publication Date: 2025-06-20GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510327572.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When processing remote sensing image vehicle object detection, the existing YOLOv8 network uses convolution operations to reduce the feature map dimensions may lead to the loss of small target information, and the impact of different features and position information on the detection performance is not fully considered, resulting in a decrease in detection accuracy in complex contexts.

Method used

A multi-scale attention mechanism (MSAM) module is introduced to enhance the feature information of the target area through channel and spatial attention mechanisms, solve the problem of target information loss caused by convolutional operations, and characterize the impact of different features and position information on model detection performance from a more comprehensive perspective.

Benefits of technology

The vehicle object detection accuracy of the model in complex contexts is significantly improved, especially the recognition ability of small targets and irregular angle targets, and the ability to retain detailed features in the image and represent feature of target areas is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182579A_ABST
    Figure CN120182579A_ABST
Patent Text Reader

Abstract

The invention discloses a YOLOv8 remote sensing image vehicle target detection method and system based on a multi-scale attention mechanism, and particularly relates to the technical field of computer vision. According to the technical scheme, the target detection network of the YOLOv8 based on the multi-scale attention mechanism is characterized in that a multi-scale attention mechanism module (Multi-Scale Attention Module, MSAM) is introduced between a Neck module and a detection head module of an original YOLOv8, the MSAM firstly extracts the importance relation between key channels through adaptive global average pooling, convolution operation and an activation function, and then the importance relation between the key channels is extracted through the adaptive global average pooling, convolution operation and the activation function, so that the target detection network of the YOLOv8 based on the multi-scale attention mechanism is obtained, and the target detection network of the YOLOv8 based on the multi-scale attention mechanism is obtained. Weighted adjustment of channel importance is achieved, and attention of the network to key channels is enhanced. Next, the obtained channel attention features are input into a space attention module, global average pooling and maximum pooling operations are carried out, space position information is extracted through convolution operation, and a space attention weight is generated; the weight is used for carrying out adaptive adjustment on the input features, so that the network can capture important spatial position information more accurately. By means of the design, accurate perception of important feature areas is enhanced, and target detection is more refined. And finally, the algorithm is designed into a remote sensing image vehicle target detection system with real-time import and accurate detection functions, so that the availability of the remote sensing image vehicle target detection system in an actual scene is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and remote sensing image processing, and particularly to a method and system for vehicle target detection in remote sensing images based on a multi-scale attention mechanism for YOLOv8. Background Art

[0002] With the rapid development of remote sensing technology, high-resolution remote sensing images have been widely used in various application fields. Especially in fields such as traffic monitoring and urban planning, vehicle target detection in remote sensing images is of great significance. Existing object detection methods, such as YOLOv3 and Faster R-CNN, etc., have been able to achieve rapid detection of image targets, but the detection accuracy in complex backgrounds still needs to be improved. The YOLOv8 model, as a new generation of object detection algorithm, has significantly improved in detection speed and accuracy compared with its predecessors.

[0003] However, existing methods for vehicle target detection in remote sensing images based on YOLOv8 usually do not fully consider the influence of different features, different scales, and different position information of targets on the detection performance of the model. This makes the model show relatively low detection accuracy when dealing with targets with different sizes, rotation angles, occlusion situations, and complex backgrounds, especially for small targets or targets with irregular angles. Specifically, existing methods may not be able to effectively enhance the detailed features of the target regions in the image, resulting in the loss of target information or background interference in complex environments.

[0004] Therefore, the present invention aims to provide a method and system for vehicle target detection in remote sensing images based on a multi-scale attention mechanism for YOLOv8 to solve the above-mentioned related problems. Summary of the Invention

[0005] The technical problem to be solved by the present invention is that when the existing YOLOv8 network processes vehicle target detection in remote sensing images, using convolution operations to reduce the dimension of the feature map may lead to the loss of information of small targets in remote sensing images, which is particularly common in remote sensing images. Small targets usually have weak features, and existing models may ignore these targets, resulting in a decrease in detection accuracy. In addition, existing methods for vehicle target detection in remote sensing images based on YOLOv8 often do not consider the influence of different features and different position information of targets on the detection performance of the model, restricting the detection accuracy of the model in different sizes, different angles, and complex backgrounds. To solve these problems, the purpose of the present invention is to enhance the performance of YOLOv8 by introducing the MSAM module and optimize the detection accuracy of vehicle targets in remote sensing images.

[0006] The present invention is achieved through the following technical solutions:

[0007] A method for detecting vehicle targets in remote sensing images based on the YOLOv8 with multi-scale attention mechanism, the method includes:

[0008] Obtain the remote sensing image to be detected and input it into a YOLOv8 object detection network based on the multi-scale attention mechanism to obtain the object detection result;

[0009] Among them, a YOLOv8 object detection network based on the multi-scale attention mechanism introduces an MSAM module between the Neck module and the detection head module to improve the detection accuracy of vehicle targets in the image. The MSAM module effectively enhances the feature information of the target area by introducing channel and spatial attention mechanisms, thus solving the problem of target information loss that may occur when the YOLOv8 network uses convolution to reduce the dimension of the feature map;

[0010] Furthermore, the MSAM module specifically includes:

[0011] Channel attention mechanism: First, perform adaptive global average pooling on the input feature map, then compress and reconstruct the output feature after adaptive global average pooling through 1x1 convolution, and normalize the convolution result to [0, 1] through the Sigmoid activation function to generate the channel attention weight, and multiply it element-wise with the original input feature map to generate the channel attention feature to strengthen the feature information of specific channels and optimize the accuracy of object detection;

[0012] Spatial attention mechanism: After performing average pooling and max pooling operations on the feature output by the channel attention in the spatial dimension, splice the features output by average pooling and max pooling respectively, then integrate the spliced features through a 7x7 convolution and reduce the number of channels to 1, and normalize the convolution result to [0, 1] through the Sigmoid activation function to generate the spatial attention weight, and multiply it element-wise with the channel attention feature to realize the combination of channel and spatial attention mechanisms, which can effectively strengthen the feature representation of the target area while retaining the detail features in the remote sensing vehicle image and improve the detection ability of vehicle targets under small targets and complex backgrounds.

[0013] Combined with a method for detecting vehicle targets in remote sensing images based on the YOLOv8 with multi-scale attention mechanism, the loss function of this method adopts bounding box loss (Box loss), classification loss (Cls loss), and distribution focal loss (Dfl loss). Among them, the formula of Box loss is as follows:

[0014] L box = 1 - CIoU(1)

[0015] Among them, the CIoU formula is as follows:

[0016]

[0017] Among them, IoU represents the intersection over union of the predicted bounding box and the ground truth bounding box, and ρ 2 (b, b gt ) is the Euclidean distance between the center of the predicted bounding box and the center of the ground truth bounding box, c is the diagonal length of the smallest bounding rectangle that can enclose the predicted bounding box and the ground truth bounding box, v is used to measure the consistency of the aspect ratio, and α is an adjustment factor.

[0018] The formula for Cls loss is as follows:

[0019] L cls = -α(1 - p t ) γ log(p t ) (3)

[0020] Among them, p t is the class probability predicted by the model, α is the balance factor, γ is the adjustment factor, which is used to reduce the influence of easy-to-classify samples and focus on difficult-to-classify samples.

[0021] The formula for Dfl loss is as follows:

[0022]

[0023] Among them, P ij is the predicted coordinate distribution probability, y ij is the discretized distribution of the ground truth coordinates, and n is the number of discretized intervals of the coordinates. Dfl loss is the loss function used by YOLOv8 for bounding box localization. By predicting the probability distribution of the coordinates and calculating the final coordinates through weighted averaging, the target localization accuracy is improved, which is especially suitable for small target and high-precision detection tasks.

[0024] The present invention also provides a remote sensing image vehicle target detection system based on a multi-scale attention mechanism for YOLOv8. The system includes:

[0025] An image upload module for uploading the remote sensing vehicle image to be detected;

[0026] A user interaction module: for providing a user-friendly interactive front-end interface;

[0027] An image parsing and preprocessing module: After the server receives the image uploaded by the user, this module first performs image format checking and preprocessing, decodes and reads the image to ensure format compatibility;

[0028] A target detection module: This module loads the target detection model of the improved YOLOv8 provided by the present invention, outputs the target category, target location and confidence, draws the target bounding box and class label, calculates detection statistical information, etc.

[0029] Model Deployment Module: This module is responsible for managing the loading and inference of the object detection model. It automatically loads the model during runtime and keeps the model in memory to improve the inference speed. This module supports GPU acceleration to improve the detection efficiency, ensure real-time performance, and allows different models to be replaced to adapt to different detection tasks.

[0030] Result Display Module: Used to display object detection statistical data, including: average detection accuracy, total detection time, map@50, map@50-95, total number of detected objects, and total number of detected categories.

[0031] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it can implement the remote sensing image vehicle target detection method based on the multi-scale attention mechanism of YOLOv8. The computer program includes an instruction set that can guide the processor to perform steps such as input, preprocessing, feature extraction, target detection, and post-processing of image data, so as to accurately detect and identify vehicle targets in remote sensing images.

[0032] A processing terminal includes a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, it can implement the remote sensing image vehicle target detection method based on the multi-scale attention mechanism of YOLOv8. The terminal can automatically complete the detection process of vehicle targets in remote sensing images through running the program, including input of images, extraction of features, identification of targets, and output of results, ensuring that the system can detect and identify vehicle targets in remote sensing images in real time and accurately, and has high processing capabilities.

[0033] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0034] By introducing the MSAM module, the present invention assigns attention weights to both the features and positions of image targets simultaneously, so as to be able to describe the impact of different features and different position information on the model detection performance from a more comprehensive perspective, significantly enhancing the model's ability to identify vehicle targets in complex backgrounds of remote sensing images, making the network more accurate and efficient in identifying key targets in complex backgrounds. Using Flask, the remote sensing image vehicle target detection method based on the multi-scale attention mechanism of YOLOv8 is integrated into a remote sensing image vehicle target detection system with real-time import and accurate detection functions, thus ensuring its high efficiency and usability in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings. In the drawings:

[0036] Figure 1 It is a flowchart of a remote sensing image vehicle target detection method based on a multi-scale attention mechanism for YOLOv8 in this embodiment;

[0037] Figure 2 It is a network structure diagram of a remote sensing image vehicle target detection method based on a multi-scale attention mechanism for YOLOv8 in this embodiment;

[0038] Figure 3 It is a structure diagram of the MSAM module in this embodiment;

[0039] Figure 4 It is a comparison diagram of the detection results between the original YOLOv8 and a YOLOv8 based on a multi-scale attention mechanism;

[0040] Figure 5 It is an overall framework diagram of a remote sensing image vehicle target detection system based on a multi-scale attention mechanism for YOLOv8 in this embodiment;

[0041] Figure 6 It is a diagram of the system result display interface; Detailed implementation manners

[0042] The following describes exemplary embodiments of the present disclosure, including various details of the embodiments of the present disclosure to facilitate understanding. They should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0043] In the present disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements does not intend to limit the positional relationship, temporal relationship or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0044] In the description of various examples in this disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in this disclosure covers any one of the listed items and all possible combinations.

[0045] As described in the background, existing remote sensing image vehicle target detection methods based on YOLOv8 usually do not fully consider the influence of different features, different scales, and different position information of targets on the detection performance of the model. This makes the model show low detection accuracy when dealing with vehicle targets with different sizes, rotation angles, occlusion situations, and complex backgrounds, especially for small targets or targets with irregular angles. Specifically, existing methods may not be able to effectively enhance the detailed features of the target area in the image, resulting in the loss of target information or background interference in complex environments.

[0046] Therefore, the present invention proposes to enhance the feature extraction ability of the model by introducing the MSAM module, especially when dealing with small vehicle targets in remote sensing images, significantly improving the target detection accuracy and robustness.

[0047] Embodiment

[0048] In one embodiment, first, 510 remote sensing vehicle image datasets in the UCAS-AOD (University of Chinese Academy of Sciences Aerial Object Dataset) are used as the basic dataset. The images in the UCAS-AOD dataset are sourced from the Google Earth software and cover multiple regions globally. This source ensures the diversity and representativeness of the data, enabling researchers to test and evaluate their algorithms under different background conditions. The calibration process of the target samples in the dataset is strictly controlled, and manual calibration is performed using self-developed software in the laboratory. After repeated inspection of each image, the accuracy of the annotation information is ensured. Then, a remote sensing image vehicle target detection method based on YOLOv8 with a multi-scale attention mechanism is used to perform target detection on the remote sensing vehicle image to be detected. See Figure 1 As shown, the method includes:

[0049] S1: Obtain the remote sensing vehicle image to be detected,

[0050] S2: Input the remote sensing vehicle image to be detected into a YOLOv8 target detection network based on a multi-scale attention mechanism to obtain the target detection result. Among them, the structure of a YOLOv8 target detection network based on a multi-scale attention mechanism is as Figure 2As shown in the figure. In this network, an MSAM module is innovatively introduced between the Neck module and the detection head module. By combining the channel attention mechanism and the spatial attention mechanism, the MSAM module enables the model to automatically focus on the key regions in the image, reduce the interference of background noise, and improve the detection performance of vehicle targets.

[0051] Specifically, as Figure 3 shown, the implementation steps of the MSAM module in this example include:

[0052] S1: First, the channel attention mechanism module performs channel attention operations on the input feature map x. The GAP module represents adaptive global average pooling, which compresses the feature map of each channel into a single value, that is, the number of output channels remains unchanged, still C. The first convolutional module below represents a 1x1 two-dimensional convolution, which reduces the number of channels from C to C / r, where r is the reduction ratio, and in this invention, r = 16 is set. The second 1x1 two-dimensional convolutional module performs channel transformation, and the Sigmoid activation function normalizes the convolution result to [0, 1]. That is, the important relationship between key channels is extracted through adaptive global average pooling and the activation function first. This enables the network to pay more attention to the feature channels with higher weights in specific detection tasks. The calculation formula of the channel attention mechanism is as follows:

[0053] ca = σ(W2(ReLU(W1(AvgPool(x))))) (1)

[0054] Where, AvgPool(x) represents the operation of performing adaptive global average pooling on the input feature map x, W1 and W2 are the weights of the two convolutional layers in the channel attention respectively, σ is the Sigmoid activation function, and ca represents the output channel attention weight.

[0055] S2: The multiplication module in the channel attention sub-module multiplies the output ca and the input feature map x element by element to obtain the feature weighted by the channel attention, so as to enhance the important channels and suppress the unimportant channels. For the convenience of representation, let the feature weighted by the channel attention be A channel .

[0056] S3: The spatial attention mechanism module takes A channelPerform spatial attention operation. First, the mean module passed through calculates the average value of each spatial position along the channel dimension. The Max module represents calculating the maximum value along the channel dimension. The concatenation module after the Max module concatenates the average value and the maximum value features to form a two-channel feature. The subsequent convolutional module represents using a 7x7 two-dimensional convolution to integrate the concatenated features and reduce the number of channels to 1. The Sigmoid activation function normalizes the convolution result to [0, 1], that is, filters and enhances the features through the convolution operation to capture important spatial position information. The calculation formula of the spatial attention mechanism is as follows:

[0057] sa = σ(Conv 7×7 (Concat(max(x), mean(x)))) (2)

[0058] Among them, max(x) and mean(x) respectively represent the maximum value and the mean value features calculated for the input x along the channel dimension. Concat represents the concatenation operation along the channel dimension. Conv 7×7 is a 7×7 two-dimensional convolution operation. σ is the Sigmoid activation function, and sa represents the output spatial attention weight.

[0059] S4: The multiplication module in the spatial attention sub-module multiplies the output sa with A channel position by position to strengthen the attention to key spatial positions, thereby realizing the fusion of channel and spatial attention and further improving the accuracy of target localization.

[0060] Among them, in this embodiment, compared with the original YOLOv8 network, a remote sensing image vehicle target detection network based on a multi-scale attention mechanism provided by the present invention is leading in terms of Recall, mAP@0.5, and mAP@0.5:.95. The comparison results are shown in the following table:

[0061]

[0062] Among them, Recall represents the recall rate, mAP@0.5 represents the average precision when the IoU threshold is 0.5, and mAP@0.5:.95 represents the average precision calculated step by step with a step size of 0.05 between the IoU thresholds from 0.5 to 0.95. Among them, the calculation formula of IoU is as follows:

[0063]

[0064] In formula (3), b t is the ground truth bounding box; b p is the predicted bounding box; A(b t ∩b p ) is the intersection of the ground truth bounding box and the predicted bounding box; A(bt ∪b p ) is the union of the ground truth bounding box and the predicted bounding box.

[0065] Recall evaluates the proportion of samples that are actually positive and are correctly predicted as positive by the model. Its specific calculation formula is as shown in (4):

[0066]

[0067] Among them, TP refers to the number of samples that the model correctly predicts as positive and are actually positive; FN refers to the number of samples that are actually positive but are incorrectly predicted as negative by the model.

[0068] For the object detection in this embodiment, after calculating the AP value for each category and taking the average, the mean average precision (mAP) is obtained. Its formula is as follows:

[0069]

[0070] Among them, N represents the number of categories in the dataset. AP is an important indicator for evaluating the performance of the model in object detection and is obtained by calculating the area under the Precision-Recall curve. In this study, N = 1 was set, and two evaluation criteria, mAP@0.5 and mAP@0.5:.95, were adopted.

[0071] A vehicle object detection method for remote sensing images based on the YOLOv8 with a multi-scale attention mechanism, based on the Recall in formula (4) and the mAP index in formula (5), obtained a Figure 4 comparison chart of the detection results between the original YOLOv8 and a vehicle object detection method for remote sensing images based on the YOLOv8 with a multi-scale attention mechanism.

[0072] In this embodiment, the loss function of the vehicle object detection method for remote sensing images based on the YOLOv8 with a multi-scale attention mechanism adopts the bounding box loss (Box loss), classification loss (Cls loss), and distribution focal loss (Dfl loss). Among them, the formula for Box loss is as follows:

[0073] L box = 1 - CIoU (6)

[0074] Among them, the CIoU formula is as follows:

[0075]

[0076] Among them, IoU represents the intersection over union of the predicted box and the ground truth box, and ρ 2 (b,b gt) is the Euclidean distance between the center of the predicted bounding box and the center of the ground truth bounding box, c is the diagonal length of the smallest bounding rectangle that can enclose the predicted bounding box and the ground truth bounding box, v is used to measure the aspect ratio consistency, and α is a scaling factor.

[0077] The formula for Cls loss is as follows:

[0078] L cls = -α(1 - p t ) γ log(p t ) (8)

[0079] where p t is the class probability predicted by the model, α is the balancing factor, and γ is the scaling factor, which is used to reduce the influence of easy-to-classify samples and focus on difficult-to-classify samples.

[0080] The formula for Dfl loss is as follows:

[0081]

[0082] where P ij is the predicted coordinate distribution probability, y ij is the discretized distribution of the ground truth coordinates, and n is the number of discretized intervals for the coordinates. Dfl loss is the loss function used by YOLOv8 for bounding box localization. By predicting the probability distribution of the coordinates and calculating the final coordinates through weighted averaging, the target localization accuracy is improved, which is especially suitable for small target and high-precision detection tasks.

[0083] In another embodiment, this embodiment also provides a remote sensing image vehicle target detection system based on a multi-scale attention mechanism for YOLOv8, as Figure 5 shown. The system includes:

[0084] Image upload module: This module allows users to select pictures locally and upload them to the server, supports multiple picture formats (such as.jpg,.png), uses Flask to handle the upload request, and stores the pictures in the corresponding directory. The front-end page design is implemented through HTML, CSS, and JavaScript, and the AJAX technology is adopted to achieve asynchronous upload. Users can complete the upload and trigger the subsequent detection process without refreshing the page.

[0085] User interaction module: This module provides a user-friendly front-end interface, including a side navigation bar, buttons for selecting files, uploading, and detecting, etc., and is responsible for managing interactive functions such as file upload and display of detection results. The page content is dynamically updated through JavaScript, and the detection results are automatically loaded without the need for users to manually refresh the page. The AJAX technology is adopted to send detection requests to the Flask server and update the interface after receiving the results.

[0086] Image parsing and preprocessing module: After the server receives the image uploaded by the user, this module first performs image format checking and preprocessing. OpenCV is used for image decoding and reading to ensure format compatibility. The preprocessing operations include size adjustment, color space conversion, and normalization. The preprocessed image will be passed as input to the object detection module. Among them, the normalization formula in the image parsing and preprocessing module is as follows:

[0087]

[0088] Among them, I raw (x, y) is the pixel value of the original image, while I normalized (x, y) is the pixel value after normalization;

[0089] Through formula (10), all pixel values are compressed to [0, 1], ensuring numerical stability. Especially when using the remote sensing image vehicle target detection model based on multi-scale attention mechanism provided by the present invention, it can accelerate the model convergence speed.

[0090] Object detection module: This module is responsible for loading and executing a remote sensing image vehicle target detection model based on multi-scale attention mechanism, and outputting the target category, location, and confidence information in the image. Specifically, it receives the preprocessed image data from the server and passes it to the object detection network for processing. In addition, the object detection module will also draw target boxes and class labels on the image to ensure target visualization for subsequent statistical analysis. This module also calculates and returns detection statistics, such as the number of targets and the confidence of each target.

[0091] Model deployment module: This module is responsible for managing the loading and inference of the object detection model to ensure the efficient and smooth progress of the entire detection process. During operation, this module will automatically load the pre-trained model and keep the model in memory to avoid the time overhead caused by frequent loading, thereby improving the inference speed. This module supports GPU acceleration to ensure real-time performance and efficiency in complex image detection tasks. The functions of the model deployment module also include dynamic model switching and inference speed optimization.

[0092] Result display module: This module uses HTML, CSS, and JavaScript for page rendering to display the object detection statistical data, including: average detection accuracy, total detection time, map@50, map@50 - 95, total number of detected objects, and total detected classes. Among them, the system result display interface diagram is as Figure 6 shown.

[0093] A computer-readable storage medium has a computer program stored thereon. It is characterized in that when the computer program is executed by a processor, it can implement a method for detecting vehicle targets in remote sensing images based on a multi-scale attention mechanism. The computer program includes an instruction set that can guide the processor to perform steps such as input, preprocessing, feature extraction, target detection, and post-processing of image data, so as to accurately detect and identify vehicle targets in remote sensing images.

[0094] A processing terminal includes a memory and a processor. A computer program that can run on the processor is stored in the memory. It is characterized in that when the processor executes the computer program, it can implement a method for detecting vehicle targets in remote sensing images based on a multi-scale attention mechanism. The terminal can automatically complete the detection process of vehicle targets in remote sensing images through running the program, including input of images, extraction of features, identification of targets, and output of results, ensuring that the system can detect and identify vehicle targets in images in real time and accurately, and has high processing capabilities.

[0095] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A remote sensing image vehicle target detection method based on YOLOv8 with multi-scale attention mechanism, characterized in that: The method includes: Obtain a remote sensing vehicle image to be detected, and input the remote sensing vehicle image to be detected into a YOLOv8 target detection network based on a multi-scale attention mechanism to obtain a target detection result; The YOLOv8 model based on the multi-scale attention mechanism includes introducing a multi-scale attention mechanism module (Multi-Scale Attention Module, MSAM) between the Neck module and the detection head module to enhance the model's detection capability of vehicle targets in remote sensing images; The MSAM module includes a channel attention module and a spatial attention module. In the channel attention mechanism, the input feature map is first subjected to adaptive global average pooling, followed by feature compression and reconstruction through 1x1 convolution, and then through the Sigmoid activation function, and finally the channel attention weight is generated, and the channel attention weight is multiplied element by element with the original input feature map to output the channel attention feature; in the spatial attention mechanism, the features weighted by the channel attention are subjected to maximum pooling and average pooling in the spatial dimension, and then the spatial features are enhanced through a 7x7 convolution layer, and the spatial attention weight is generated through the Sigmoid activation function, and the spatial attention weight is multiplied element by element with the channel attention feature to achieve the fusion of channel attention and spatial attention.

2. A remote sensing image vehicle target detection method based on YOLOv8 with multi-scale attention mechanism as claimed in claim 1, characterized in that: The implementation steps of the channel attention mechanism in the MSAM module include: S1: Obtain global features through adaptive global average pooling operation on the input feature map; S2: Use 1x1 convolution to reduce the global features to a lower dimension and perform nonlinear mapping through the ReLU activation function; S3: The channel information is restored to the original dimension through the second 1x1 convolution layer, and the convolution result is normalized to [0, 1] through the Sigmoid activation function to finally obtain the channel attention weight; S4: Multiply the channel attention weights by the original input feature map element-wise and output the channel attention features, thereby enhancing the useful information in the feature map.

3. The remote sensing image vehicle target detection method based on YOLOv8 of the multi-scale attention mechanism as claimed in claim 1, characterized in that: The implementation steps of the spatial attention mechanism in the MSAM module include: S1: Perform average pooling and maximum pooling on the channel attention features generated by channel attention in the spatial dimension to obtain two different spatial features; S2: Concatenate the two spatial features together and integrate the concatenated features through a 7x7 convolution and reduce the number of channels to 1. S3: Apply the Sigmoid activation function to the integrated features, normalize the convolution results to [0, 1], and generate spatial attention weights; S4: Multiply the spatial attention weights by the channel attention features element-wise to enhance the target region in the image.

4. The method for detecting vehicle targets in remote sensing images based on YOLOv8 with a multi-scale attention mechanism as claimed in claim 1, characterized in that: The MSAM module can improve the detection capability of vehicle targets in remote sensing images, especially the positioning and classification performance of vehicle targets in remote sensing images with complex backgrounds, by combining channel and spatial attention mechanisms.

5. The method for detecting vehicle targets in remote sensing images based on YOLOv8 with a multi-scale attention mechanism as claimed in claim 1, characterized in that: The loss function uses bounding box loss (Boxloss), classification loss (Cls loss), and distributed focal loss (Dfl loss). The formula of Box Loss is as follows: L box =1-CIoU (1) Among them, the CIoU formula is as follows: Among them, IoU represents the intersection-over-union ratio between the predicted box and the real box, ρ 2 (b,b gt ) is the Euclidean distance between the center of the predicted box and the center of the true box, c is the diagonal length of the minimum bounding rectangle that can enclose the predicted box and the true box, v is used to measure the consistency of the aspect ratio, and α is the adjustment factor. The formula of Cls loss is as follows: L cls =-α(1-p t ) γ log(p t ) (3) Among them, p t is the category probability predicted by the model, α is the balancing factor, and γ is the adjustment factor, which is used to reduce the influence of easy-to-classify samples and focus on difficult-to-classify samples. The formula for Dfl loss is as follows: Among them, P ij is the predicted coordinate distribution probability, y ij is the discretized distribution of the real coordinates, and n is the number of discretized intervals of the coordinates. Dfl loss is the loss function used by YOLOv8 for bounding box positioning. It improves the accuracy of target positioning by predicting the probability distribution of coordinates and calculating the final coordinates by weighted average. It is particularly suitable for small targets and high-precision detection tasks.

6. A remote sensing image vehicle target detection system based on YOLOv8 with multi-scale attention mechanism, characterized in that: The system includes: Image upload module, used to upload remote sensing vehicle images to be detected; User interaction module: used to provide a user-friendly interactive front-end interface; Image parsing and preprocessing module: After receiving the pictures uploaded by users on the server side, this module first performs image format check and preprocessing, and then decodes and reads the images to ensure format compatibility; Object detection module: This module is responsible for loading and executing the improved YOLOv8 object detection model of the present invention, and outputting the object category, location, and confidence information in the image. Specifically, it receives preprocessed image data from the server and passes it to the object detection network for processing. In addition, the object detection module also draws the object box and category label on the image to ensure the visualization of the object and facilitate subsequent statistical analysis. This module also calculates and returns detection statistics, such as the number of objects and the confidence of each object. Model deployment module: This module is responsible for managing the loading and reasoning of the target detection model, ensuring the efficiency and smoothness of the entire detection process. At runtime, the module automatically loads the pre-trained model and keeps the model in memory to avoid the time overhead caused by frequent loading, thereby improving the reasoning speed. This module supports GPU acceleration to ensure real-time and efficiency in complex image detection tasks. The functions of the model deployment module also include dynamic model switching and reasoning speed optimization. Result display module: used to display target detection statistics, including average detection accuracy, total detection time, map@50, map@50-95, total number of detected objects and total detection categories.

7. The functions of the model deployment module as claimed in claim 6 also include dynamic model switching and reasoning speed optimization, characterized in that: include: Dynamic model switching: allows different target detection models to be replaced as needed to adapt to different task requirements; Inference speed optimization: Through parallel computing and GPU acceleration, the inference time is reduced to ensure that the system can process and return detection results in real time, especially when the number of images is large or the image resolution is high.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it can implement the remote sensing image vehicle target detection method based on YOLOv8 of a multi-scale attention mechanism as described in any one of claims 1 to 4. The computer program includes an instruction set that can guide the processor to perform steps such as image data input, preprocessing, feature extraction, target detection and post-processing, so as to accurately detect and identify vehicle targets in remote sensing images.

9. A processing terminal comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, it can implement a remote sensing image vehicle target detection method based on a multi-scale attention mechanism YOLOv8 as described in any one of claims 1 to 4. The terminal can automatically complete the vehicle target detection process in the remote sensing image by running the program, including image input, feature extraction, target recognition and result output, ensuring that the system can detect and recognize vehicle targets in the image in real time and accurately, and has efficient processing capabilities.

Citation Information

Cited By

  • Automobile hinge profile processing data processing method and system based on deep learning

    CN120930026A

  • Target detection method based on detail enhancement convolution, storage medium and terminal

    CN121392250A