UAV-borne target detection method based on adversarial training strategy

By improving the YOLOv7-Tiny model and combining the adversarial training strategies of SPCBlock and MDCA modules, the problem of redundancy and insufficient calculation accuracy of drone target detection algorithm is solved, and efficient and real-time automated detection of highway diseases is achieved.

CN120126031BActive Publication Date: 2025-09-02GUANGXI COMPREHENSIVE TRANSPORTATION BIG DATA RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510179072.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-09-02
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing drone target detection algorithms have redundancy in highway disease detection, large parameters, large models, and insufficient accuracy of lightweight algorithms, making it difficult to meet the real-time detection needs.

Method used

The improved YOLOv7-Tiny lightweight object detection model is adopted, combined with the SPCBlock module and the MDCA module, and the model is optimized through adversarial training strategies, and the PGD algorithm is introduced to load perturbations on the first layer of the model to improve detection accuracy and efficiency.

Benefits of technology

It improves the efficiency and accuracy of highway disease detection, reduces costs, ensures driver safety, and realizes automated drone detection and real-time monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126031B_ABST
    Figure CN120126031B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for detecting targets on drones based on an adversarial training strategy. The method comprises: collecting images and video data of a highway in real time through a camera carried by a drone, inputting the collected data into a lightweight drone robust target detection model based on an improved YOLOv7‑Tiny, and visualizing the detection results returned by the model on the front end, wherein the backbone network of the target detection model comprises an SPCBlock module improved based on a CS module, and the neck network of the target detection model comprises an MDCACBlock module formed by combining an Elan module and an MDCA module; the target detection model adopts an adversarial training strategy of an improved projected gradient descent algorithm to load the disturbance generated by the model on the first layer parameters of the model. The present invention replaces the corresponding module of the traditional model with a module that is lightweight and has efficient feature processing capabilities, and uses disturbed samples to train the model, thereby realizing effective detection and monitoring of highways by drones.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting unmanned aerial vehicle (UAV) targets based on an adversarial training strategy, and belongs to the fields of intelligent transportation, artificial intelligence, image segmentation, and medical image processing. Background Art

[0002] With the rapid development of the transportation industry, highways have become critical infrastructure for modern socioeconomic activities, connecting cities, facilitating logistics, and improving economic development efficiency. As society's demand for transportation continues to grow, the highway network is expanding year by year, significantly increasing the number of vehicles and traffic loads. As traffic pressure continues to grow and the natural environment impacts, highway pavements and structures are increasingly threatened by various defects. These include pavement defects such as cracks and potholes, localized shoulder subsidence and damage, bridge deck cracks and expansion joint failures, tunnel lining cracks, water leakage and seepage, and support structure damage. Therefore, efficiently and accurately identifying and monitoring highway defects is of great practical significance.

[0003] Intelligent transportation is a key area of ​​modern urban development, aiming to improve the efficiency and safety of transportation systems through the use of advanced information technology and data analysis methods. With the rapid development of intelligent transportation systems, traditional fixed cameras are no longer able to meet the needs of large-scale, highly maneuverable traffic monitoring. In this context, drones, with their advantages of high operability, wide viewing angles, high flexibility, and low cost, are widely used in intelligent traffic management and video surveillance.

[0004] Existing technology initially relied on manual image collection of various highway pavement defects. However, to save manpower, reduce unnecessary time, and increase detection rates, drone inspections are now increasingly being used. This involves using drones to capture road conditions on target sections and store them locally. The captured videos or images are then manually inspected visually. While the high performance and maneuverability of drones significantly increase data collection rates compared to manual data collection, post-collection data analysis still requires significant labor and time, resulting in low detection efficiency.

[0005] The rapid advancement of computer vision and deep learning technologies has brought new opportunities for drone traffic monitoring. Current object detection can be divided into two stages: object detection based on traditional handcrafted feature design and object detection based on deep learning. Traditional handcrafted feature-based object detection uses a sliding window technique to acquire the target region, extract features from the target region, and then implements a classifier and regressor to achieve multi-target classification and regression. Deep learning-based object detection uses convolutional neural networks to extract target features and is categorized as either a two-stage or one-stage approach, depending on whether a candidate bounding box is present. Two-stage object detection first delineates candidate regions that may contain targets and then uses feature extraction and classification regression to determine the target region. While it offers the best performance in terms of target accuracy and recall, it cannot meet real-time detection speed requirements. One-stage object detection uses convolutional neural networks for direct classification and regression. While fast enough to meet real-time detection requirements, it is inferior to two-stage in terms of accuracy and recall. Consequently, existing object detection algorithms suffer from computational redundancy, a high number of parameters, and large model size. Furthermore, the accuracy of existing lightweight object detection algorithms needs to be improved. For example, although the YOLO algorithm is famous for its real-time processing capabilities, the target detection algorithm improved based on YOLO has many parameters and a large model, and the performance of the edge devices carried by drones is limited, making it difficult to effectively run such algorithms. Although many scholars have developed lightweight detection algorithms for drone scenarios, most lightweight algorithms generally achieve lightweightness at the expense of target accuracy. Summary of the Invention

[0006] The present invention provides a method for detecting targets onboard an unmanned aerial vehicle (UAV) based on an adversarial training strategy, aiming to solve at least one of the technical problems existing in the prior art.

[0007] The technical solution of the present invention relates to a target detection method based on a drone, and the method according to the present invention comprises the following steps:

[0008] S100, uses the camera onboard the drone to collect real-time images and video data of the highway;

[0009] S200, inputs the collected data into a lightweight UAV robust target detection model based on the improved YOLOv7-Tiny;

[0010] S300: The front end visualizes the detection results returned by the model;

[0011] The backbone network of the target detection model includes a SPCBlock module improved based on the CS module, and the neck network of the target detection model includes an MDCACBlock module formed by combining an Elan module and an MDCA module;

[0012] The target detection model adopts an adversarial training strategy of an improved projected gradient descent algorithm to load the perturbations generated by the model on the first layer parameters of the model.

[0013] Furthermore, in the target detection model,

[0014] The SPCBlock module includes two SPC modules based on the StarBlock structure combined with partial convolution.

[0015] Furthermore, the SPC module is configured to:

[0016] A partial convolution with a 3×3 convolution kernel combined with a batch normalization layer is used to process a quarter of the input features, reducing the amount of computation while maintaining the accuracy of the model.

[0017] The features after the above convolution processing are spliced ​​together to form richer semantic information; the spliced ​​features are input into the feature weighting module to extract more critical features;

[0018] The weighted features are channel compressed through 1×1 convolution, and then processed by batch normalization and SiLU nonlinear activation to further enhance the expressive power of the model.

[0019] The processed features are connected to the original features of the third branch through residual connection to alleviate the gradient vanishing problem during training and retain the original feature information.

[0020] Furthermore, the feature weighting module consists of two sub-branches, one of which consists of a 1×1 convolution and a Sigmoid function, and the other branch includes a 1×1 convolution.

[0021] Furthermore, the weighting process of the feature weighting module is expressed as follows:

[0022] y=σ(W1*x+b1)⊙(W2*x+b2)

[0023] Where y represents the output feature, x represents the input feature, W1 and W2 represent the learned weight matrices, b1 and b2 represent the bias terms, σ represents the sigmoid activation function, and ⊙ represents element-wise multiplication.

[0024] Furthermore, in the target detection model,

[0025] The MDCA module is composed of a multi-scale dilated convolution and a CA attention mechanism module;

[0026] The MDCA module is configured as follows:

[0027] It includes three parallel dilated convolution branches, each of which uses a 3×3 dilated convolution with a dilation rate of 3 to capture contextual information of different scales;

[0028] The captured features of the three dilated convolution branches are concatenated with the convolution output with a dilation rate of 3 to form a richer multi-scale feature representation;

[0029] The above fused features are input into the CA attention mechanism module to effectively capture the dependencies in the spatial and channel dimensions;

[0030] The original input features are added to the processed features through the residual connection structure.

[0031] Furthermore, the MDCA module is configured to:

[0032] The input of the module passes through a 1×1 convolutional layer to reduce the dimension and extract preliminary features;

[0033] The preliminary features are transformed nonlinearly through batch normalization and ReLU activation function to enhance the expressiveness of the model and improve the stability of training;

[0034] The features enter a 3×3 dilated convolution layer with a dilation rate of 3 to significantly expand the receptive field without adding additional parameters; wherein the dilated convolution layer includes a 3×3 convolution kernel with a dilation rate of 3;

[0035] The output features after the dilated convolution processing are concatenated with the original input features through the Concatenate operation to further enrich the feature expression and strengthen the model's ability to capture features at different levels.

[0036] The technical solution of the present invention also relates to a computer-readable storage medium having program instructions stored thereon, and the above-mentioned method is implemented when the program instructions are executed by a processor.

[0037] The technical solution of the present invention also relates to a target detection system based on a drone, wherein the system includes a computer device containing the above-mentioned computer-readable storage medium.

[0038] Furthermore, the target detection system includes: a drone camera acquisition module for collecting real-time images of the drone during flight; a Websocket communication module for interacting the drone's back-end instructions with the model algorithm; a core algorithm processing module for receiving special instructions from the back-end to perform different algorithm detections; a data storage and management module for downloading resources before detection and uploading data after detection; a push-pull stream module for pushing the detected image to a specified address; and a result visualization module for combining the push stream address with the front-end interface.

[0039] The beneficial effects of the present invention are as follows:

[0040] The present invention improves the target detection algorithm based on the lightweight yolov7-tiny, replaces the corresponding module of the traditional model with a lightweight module with efficient feature processing capabilities, and uses interfered samples to train the model. The present invention improves the backbone network and neck network of the model respectively, adopting the SPC module and the MDCA module. Both modules adopt a lightweight design with a simple and efficient structure, which ensures lightweight while improving the target detection accuracy of the model. The CA attention mechanism is introduced into the MDCA module, which effectively improves the ability to focus on key target areas through adaptive adjustment of channel and spatial information. The present invention introduces an adversarial training strategy based on the PGD algorithm. Compared with directly interfering with samples, the model is trained with interfered samples, and the original prediction results are changed by perturbing the first-layer parameters of the model, thereby achieving the effect of interfering with the model and realizing the effect of adversarial training.

[0041] The present invention improves the detection efficiency of highway pavement and surrounding environment, reduces costs, and ensures the safety of drivers. It proposes a drone-based target detection method for automated identification and monitoring of highway defects, providing a fast and convenient technology for vehicles to obtain real-time information about the surrounding environment of the highway, as well as for highway safety and maintenance inspections, enabling it to effectively serve the drone and traffic management related industries. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is a basic flow chart of the method according to the present invention.

[0043] Figure 2 is a flow chart of a target detection algorithm according to the method of the present invention.

[0044] Figure 3 2 is a schematic diagram of the structure of the Yolov7-Tiny module before improvement according to an embodiment of the present invention.

[0045] Figure 4 2 is a schematic diagram of the improved Yolov7-Tiny module structure according to an embodiment of the present invention.

[0046] Figure 5 2 is a schematic diagram of the structure of the Elan module before improvement according to an embodiment of the present invention.

[0047] Figure 6 FIG. 1 is a structural diagram of an SPCBlock module after improving the Elan module according to an embodiment of the present invention.

[0048] Figure 7 2 is a schematic structural diagram of an SPC module according to an embodiment of the present invention.

[0049] Figure 8 4 is a flow chart of Pconv processing features according to an embodiment of the present invention.

[0050] Figure 9 2 is a schematic diagram of the improved MDCABlock module structure according to an embodiment of the present invention.

[0051] Figure 10 FIG. 4 is a schematic diagram of the MDCA module structure according to an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The following will provide a clear and complete description of the concept, specific structure and technical effects of the present invention in conjunction with the embodiments and drawings to fully understand the purpose, scheme and effects of the present invention.

[0053] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. The singular forms of "," "said" and "the" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the invention. The term "and / or" used herein includes any combination of one or more of the related listed items.

[0054] Should be understood that, although the present disclosure may adopt the term first, second, third etc. to describe various elements, these elements should not be limited to these terms.These terms are only used to distinguish the elements of the same type from each other.For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element.The use of any and all examples or exemplary language ("for example", "such as" etc.) provided herein is only intended to better illustrate embodiments of the present invention, and unless otherwise required, will not impose limitations on the scope of the present invention.

[0055] Reference Figures 1 to 10 In some embodiments, the target detection method based on a drone according to the present invention includes at least the following steps:

[0056] S100, uses the camera onboard the drone to collect real-time images and video data of the highway;

[0057] S200, inputs the collected data into a lightweight UAV robust target detection model based on the improved YOLOv7-Tiny;

[0058] S300: The front end visualizes the detection results returned by the model;

[0059] The backbone network of the target detection model includes a SPCBlock module improved based on the CS module, and the neck network of the target detection model includes an MDCACBlock module formed by combining an Elan module and an MDCA module;

[0060] Among them, the target detection model adopts an adversarial training strategy of an improved projected gradient descent algorithm to load the perturbation generated by the model on the first layer parameters of the model.

[0061] In some embodiments, the drone module of the system of the present invention includes a drone wireless communication module, a high-definition camera module, a drone obstacle avoidance module, and a GPS positioning module. The drone's wireless communication module establishes a data communication link with the ground station's data transmission device, or server. The data captured by the camera and GPS positioning information are encoded, modulated, and transmitted via the link as a 5.8 GHz wireless signal. A ground-based receiving server receives, decodes, and demodulates the received data. The high-definition camera primarily captures images or videos of the scene the drone passes through, and transmits the captured data to the drone's wireless communication module. The drone's obstacle avoidance module is typically equipped with various sensors for detecting the surrounding environment, such as lidar, ultrasonic sensors, and infrared sensors. These sensors acquire information such as the location, size, and distance of obstacles, establish an environmental perception model, and interact with the drone's high-definition camera to prevent the drone from encountering obstacles that could affect its trajectory. The GPS module primarily determines the drone's current geographic location and interacts with the ground. The GPS positioning module accurately monitors the drone's flight trajectory and prevents deviations. Based on received satellite signals, the GPS module calculates the drone's current latitude and longitude coordinates, altitude, speed, and other information and transmits this information to the wireless communication module. The front-end interface then displays the drone's specific navigation position in real time.

[0062] Compared to previous drone target detection methods and systems, the present invention combines deep learning-based methods. By using deep learning's inherent machine learning approach, which mimics the human brain's ability to judge external objects, it can significantly reduce the time required for manual judgment, thereby improving time utilization and work efficiency. Furthermore, the present invention utilizes drone-based automated detection, with automation primarily reflected in the fact that the algorithmic detection process from acquisition to target detection, as well as the subsequent information storage and push-pull streaming processes, are all run as a whole, without the need for human intervention. Drones, with their flexibility and small size, offer inherent advantages for data collection. Furthermore, the drone's onboard communication module connects to a server, allowing the collected data to be transmitted to a high-performance server via wireless communication. By selecting the appropriate detection method based on the drone's mode, it can classify various highway defects, including whether they require repairs, and conduct real-time detection. The detection results are then transmitted to the central computer that initiated the inspection in real time. All detected targets and problematic targets are classified using different labeling methods, visualized on the central computer, and the results are transmitted to inspectors, facilitating subsequent updates and monitoring. Furthermore, the invention can expand upon this foundation with the addition of more algorithmic modules, enabling the realization of multifunctional automated drone inspections. Compared to manual inspections, drone inspections are less expensive and more efficient, eliminating the need for on-site personnel and minimizing potential safety hazards. Furthermore, the present invention provides cloud backup of inspected video streams or images, facilitating subsequent review and inspection of highway road conditions at a specific time.

[0063] It should be noted that the present invention adopts an automated identification, inspection and monitoring method for highway diseases based on drones, improves the functions based on deep learning and python libraries, and uses the target detection technology of the one-stage yolov5.7.0 version based on deep learning. Among them, in order to meet the purpose of real-time detection, there are certain restrictions on the rate of target detection. According to the development history of target detection, in the target detection based on deep learning, the two-stage target detection method has high accuracy but low detection efficiency, and the rate of 3-4 frames per second is not competent. Therefore, the present invention adopts a one-stage target detection method, and utilizes the fast speed of the yolo series method (45FPS rate per second) and the good advantages of yolov5 with stability and good integration, so that the present invention can meet the real-time detection needs.

[0064] In some embodiments, the automated detection system of the present invention includes a drone camera acquisition module, a Websocket communication module, a core algorithm processing module, a data storage and management module, a push-pull flow module, and a result visualization module.

[0065] In one embodiment, the camera acquisition module of the present invention is used to capture real-time footage of a drone in flight. Specifically, the front-end interface controls the drone's exit from the warehouse and vertical takeoff. At approximately 150 meters above the ground, the drone's posture is adjusted to a horizontal position. The drone's camera is positioned at an angle of approximately 75 degrees to the ground and then flies to a designated inspection location. The camera is then turned on, and the drone flies along the designated route, streaming the captured footage in real time to a public network address for subsequent algorithm-based stream detection. The captured real-time footage is then saved to the local network.

[0066] In one embodiment, the WebSocket module of the present invention is used to communicate between drone backend commands and algorithms. Specifically, a WebSockets module, implemented in a Python library, enables real-time asynchronous communication with the backend. It periodically sends heartbeat packets to the backend to prevent disconnection, and detects WebSocket anomalies to implement a communication reconnection mechanism.

[0067] In one embodiment, the core algorithm processing module of the present invention is used to implement switching between algorithms and different models trained with different algorithms. Specifically, before testing, a folder is created to store downloads from the drone or database, and a folder and output address for test results are established. The core algorithm processing module follows the WebSocket module and performs different algorithm tests upon receiving special instructions from the backend. For example, receiving an RTSP keyword from the backend enables real-time drone testing; otherwise, testing is performed on offline static images or video streams.

[0068] Furthermore, for the core algorithm processing module, in the part of real-time detection of data collected by the drone, the detection model needs to be trained with yolo in advance, the model needs to be started by the operation command line using the os module, and the stream is pulled to the address where the drone pushes the stream, and the output video after frame-by-frame detection is synthesized.

[0069] Furthermore, in the core algorithm processing module, for offline static images and video streams, the image or video stream to be tested must first be downloaded locally using a multi-process approach. This involves sending a post request, saving the corresponding resource locally by name, and using the time module to record the download time. The resource is then tested using the same real-time detection method, and the test results are output to a local folder for upload to the Minio storage space.

[0070] In one embodiment, the data storage and management module of the present invention is used to download resources before testing and transfer data to an OSS database after testing. Specifically, before testing, the main task is to create a recursive resource folder using the os.makedir function in Python. After testing, the OSS2 module stores the test result address in a bucket, thereby uploading the test results.

[0071] In one application embodiment, the push-pull stream module of the present invention mainly uses FFmpeg's streaming technology to push the detected picture to the specified address in the format of 640*480 for each frame of the video or picture in the detect.py file of yolov5.7.0 during the algorithm execution, thereby realizing the broadcast characteristics of the detection.

[0072] In one application embodiment, the result visualization module of the present invention utilizes the video tag in HTML technology or the video playback library Video.js to include the streaming address, thereby integrating the streaming address with the front-end interface, and providing a convenient screen for data monitoring personnel to conveniently check whether the highway pavement and surrounding protective equipment need to be updated or repaired.

[0073] See also Figure 1 and Figure 2 In the method for automated identification and monitoring of highway defects based on drones of the present invention, the drone's takeoff attitude and angle are first adjusted to reach the target area, and then the drone camera is turned on to collect data in the target area and broadcast live to the target address. The live video is streamed and the algorithm is executed to push the detection results to the designated address for live broadcast, and finally the live content is visualized at the front end. Furthermore, the processing process of its target detection model includes: first sending a detection instruction to the drone, performing two-way communication through Websocket, and then judging whether RTSP is 1. If RTSP is not 1, a static image and the collected video are downloaded to a local directory, and then the local image or video is taken out for detection in turn, and then the target detection model is used for detection and the result is returned; if RTSP is 1, the target address is pulled → the target detection model detects and returns the result → the result is pushed to the designated target address for live broadcast. Then, the detection result is uploaded to the OSS cloud, and finally the front end plays the detection result in real time.

[0074] See also Figures 3 to 10 The target detection method based on drones of the technical solution of the present invention aims to solve the problems of computational redundancy, large number of parameters, and large model in existing drone target detection models, as well as the problem that the accuracy of existing lightweight target detection algorithms needs to be improved. The present invention improves the lightweight yolov7-tiny target detection algorithm, replaces the corresponding module of the traditional model with a lightweight module with efficient feature processing capabilities, and changes the model training strategy.

[0075] Specifically, the present invention improves the backbone network and the neck network of the model. Specifically, the present invention designs two modules: the SPC (StarPConv) module and the MDCA (Multi-scale Dilated Convolution Coordinate Attention) module. Both modules are characterized by lightweight design, simple and efficient structure, ensuring lightweight while improving the model's object detection accuracy.

[0076] In one application embodiment, the SPC module of the present invention focuses on enhancing the efficiency of feature extraction by simplifying the calculation process in the processing of drone aerial images. The SPC module ensures the effective capture of the target area while reducing redundant calculations through partial convolution (PConv) operations. The SPC module of the present invention can streamline the network structure so that it can run efficiently on resource-constrained drone equipment, and is particularly suitable for scenarios that require rapid response. Specifically, the SPC module simultaneously performs weighted processing on features, effectively improving the recognition ability of key areas, ensuring that important features are enhanced and highlighted, and irrelevant information is minimized, so that the SPC module can quickly locate the target area in the aerial image, avoiding excessive calculation of irrelevant background, and improving the efficiency and accuracy of image processing.

[0077] In one application embodiment, the MDCA module of the present invention exhibits remarkable lightweight properties in drone aerial image processing, capable of extracting rich features of key targets while maintaining computational efficiency. Through multi-scale dilated convolution operations, the module effectively expands the receptive field, addressing the target detection requirements arising from varying altitudes and perspectives during drone flight. The MDCA module helps ensure that accurate target capture is not lost due to scale differences during processing, particularly in complex scenes, while maintaining high-resolution detail.

[0078] In one application embodiment, the MDCA module of the present invention introduces a Coordinate Attention (CA) mechanism, which effectively improves the ability to focus on key target areas by adaptively adjusting channel and spatial information. The CA attention mechanism further optimizes the extraction of important features, enhancing the model's object detection accuracy in complex backgrounds while maintaining a lightweight design, ensuring higher efficiency when processing high-resolution images.

[0079] In an application embodiment, see Figure 3 and Figure 4As shown in the structure of Yolov7-Tiny model before and after improvement, the present invention improves the backbone network (Backbone) and neck network (Neck) of yolov7-tiny. Figure 5 and Figure 6 As shown in the Elan module before and after improvement, the SPCBlock module improved based on the Elan model replaces the only two 3*3 convolutions in the Elan model with the SPC (StarPconv) module, where the SPC module structure is shown in Figure 7 As shown in FIG, the SPC module of the present invention draws on the StarBlock structure in StarNet and adds the Pconv (Partial Convlution) proposed by FasterNet.

[0080] In one embodiment, the SPC module of the present invention is a lightweight and efficient convolutional structure, which draws on the design concept of StarBlock. Figure 7 and Figure 8 ,The SPC module of the present invention consists of three main branches.

[0081] Among them, the first branch of the SPC module adopts partial convolution (PConv) combined with batch normalization layer. PConv uses 3×3 convolution kernel, which is different from traditional convolution. The PConv of the present invention only processes one quarter of the input features. Specifically, based on the observation of channel feature maps, it is found that there is a high degree of similarity and redundancy between different channels. The present invention adopts the method of processing only part of the channels to make PConv. For example, the memory access volume of ordinary convolution is h×w×2c+k 2 ×c 2 ≈h×w×2c (Formula 1), and the memory access capacity of PConv of the present invention can reach Where c is the number of input channels, c p Equal to 1 / 4C, which greatly reduces the computational requirements of PConv, where h and w are the height and width of the feature map, respectively, and k is the convolution kernel size.

[0082] Among them, the second branch of the SPC module does not directly process the features, but concatenates these features with the features processed by PConv to form richer semantic information. Subsequently, these concatenated features enter a feature weighting module to extract more critical features. Specifically, the weighting module consists of two sub-branches: one is composed of a 1×1 convolution and a Sigmoid function, and the other contains only a 1×1 convolution. In this way, the present invention not only maintains the lightweight of the module, but also introduces a simple attention mechanism to enhance the feature extraction capability.

[0083] Furthermore, the feature weighting process can be expressed by the following formula: y = σ(W1*x+b1)⊙(W2*x+b2) (Equation 3), where x is the input feature, W1 and W2 are the learned weight matrices, b1 and b2 are bias terms, σ represents the sigmoid activation function, and ⊙ represents element-wise multiplication.

[0084] The weighted features are then compressed using a 1×1 convolution, followed by batch normalization and SiLU nonlinear activation to further enhance the model's expressiveness. Finally, these processed features are connected to the original features from the third branch using a residual connection, alleviating the vanishing gradient problem during training while effectively preserving the original feature information.

[0085] In an application embodiment, in the neck network of the present invention, an MDCA module is added after the Elan module to form a new MDCACBlock module, see Figure 5 and Figure 9 The MDCA module of the present invention is composed of a combination of multi-scale dilated convolution and CA attention mechanism, see Figure 6 As shown in the figure, it should be noted that MDCA (Multi-scale Dilated Convolution Coordinate Attention) is an efficient feature extraction structure that combines multi-scale dilated convolution with the coordinate attention mechanism. The MDCA module can expand the receptive field and capture multi-level features while maintaining high computational efficiency.

[0086] Specifically, in the MDCACBlock module formed by the combination of the present invention, first, the input of the module passes through a 1×1 convolution layer, which plays the role of dimensionality reduction and preliminary feature extraction. Then, the features are transformed nonlinearly through batch normalization (BN) and ReLU activation function to enhance the expressiveness of the model and improve the stability of training. Next, the features enter a 3×3 dilated convolution layer with an expansion rate of 3. This dilated convolution is the core component of the MDCA module, which can greatly expand the receptive field without adding additional parameters. Specifically, the receptive field of the 3×3 convolution kernel with an expansion rate of 3 is equivalent to that of a 7×7 ordinary convolution kernel, but its number of parameters is only 9, which is about 18.37% of the number of parameters of the 7×7 convolution kernel. The above-mentioned network model structure of the present invention enables the model to capture a wide range of information while still maintaining a high computational efficiency. Then, the output features after the dilated convolution processing are spliced ​​with the original input features through the Concatenate operation to further enrich the feature expression and enhance the model's ability to capture features at different levels.

[0087] See also Figure 10The multi-scale nature of the MDCA model is achieved through three parallel dilated convolution branches described below. Specifically, each branch progressively extracts features using a 3×3 dilated convolution with a dilation rate of 3. These captured features are then concatenated with the output of the initial dilated convolution with a dilation rate of 3. This concatenated feature is then fed into a 1x1 fused multi-scale feature representation, forming a richer multi-scale feature representation. This progressive dilated convolution and feature concatenation strategy enables it to simultaneously capture both local and broader contextual information, enhancing the model's ability to perceive diverse scene information.

[0088] Then, after the feature fusion, the MDCA module introduces the CA attention mechanism (see Figure 10 The CA attention mechanism can effectively capture the dependencies in spatial and channel dimensions, improve the model's attention to key features, and thus enhance the model's performance in complex scenarios.

[0089] Finally, the MDCA module adopts a residual connection structure to add the original input features to the processed features, which not only helps to alleviate the gradient vanishing problem in deep networks, but also retains the original information, making it easier for the network to learn the identity mapping, thereby improving the training effect and optimization ability of the model.

[0090] In some embodiments, the present invention introduces an adversarial training strategy based on the PGD (Projected Gradient Descent) algorithm. This improves the robustness of the model, allowing it to maintain a certain level of accuracy even with perturbed images. Compared to directly interfering with samples and then using them to train the model, the present invention perturbs the first-layer parameters of the model, altering the original predictions and thus achieving the effect of interfering with the model and achieving the effect of adversarial training.

[0091] It should be noted that the original PGD iteratively generates adversarial samples for a batch of samples. Specifically, the traditional model uses forward propagation, and finally calculates the loss, generates the corresponding gradient backpropagation, and updates the model parameters. When the gradient backpropagates to the first layer, the gradient at this moment is used to generate perturbations and interfere with clean samples, and further generate adversarial samples to train the model. This cycle is repeated a certain number of times, continuously maximizing the loss to form adversarial gradients to generate perturbations, to update the model parameters, and then interferes with the next batch after the iteration is completed.

[0092] This invention improves the interference method. Specifically, although the overall idea is still based on PGD adversarial training, the perturbation generated by the model of this invention is not added to the clean samples, but to the first layer parameters of the model. This invention changes the prediction results by perturbing the parameters, which is the same as the effect of directly interfering with the samples. It is convenient and efficient, and does not require additional generation of adversarial samples. The parameter update formula of this invention is as follows:

[0093]

[0094] Specifically, the adversarial training process based on the PGD (Projected Gradient Descent) algorithm is as follows:

[0095] The model inputs are: model M, first-layer model parameters θ0, training data X, label y, step size α, perturbation range S, number of iterations K, and learning rate η. The model output is the model M with enhanced robustness.

[0096] A1、θ adv ←θ0 / / Initialize adversarial parameters.

[0097] A2、θ original ←θ adv / / Back up the original parameters.

[0098] A3. / / Calculate the initial gradient

[0099] Perform K-1 cycles.

[0100] A4, / / Update parameters, where Π represents the projection operator, which projects the parameters to θ original The sphere with radius S as the center.

[0101] After the last iteration, the network parameters are updated.

[0102] A5.

[0103]

[0104] A6. Return M / / Return the model with enhanced robustness.

[0105] A specific embodiment is used here for illustration.

[0106] First, download a public dataset. Here, we choose the UAV-PDD2023 road damage dataset. The dataset contains 2,440 three-channel JPG images and corresponding VOC-format annotation files. Six types of road damage are labeled in the images: longitudinal cracks (LC), transverse cracks (TC), alligator cracks (AC), oblique cracks (OC), patch (RP), and pothole (PH). The VOC-format labels are then converted into a YOLO txt file. The dataset is then randomly divided into training, validation, and test sets in a 7:2:1 ratio.

[0107] The UAV target detection model based on the improved yolov7-tiny of the present invention is constructed, and the model training is performed. The number of iterations of PGD adversarial training, as well as the training strategies such as batch size and epoch number are adjusted according to the implementation results. The batch size is set to 16, the epoch is set to 150, the learning rate is 0.01, and the optimizer is Adam. After the training is completed, the optimal weight model is retained, and the trained model is experimentally tested with a divided test set. It should be noted that the experimental platform of the embodiment of the present invention is based on the Linux Ubuntu22 operating system. The training platform uses Nvidia GeForce RTX 309024G GPU, Intel (R) Xeon (R) CPU E5-2697A v4 @ 2.60GHz, 64G memory, the code running framework is PyTorch, and the operating environment is Python3.7, CUDA 12.4.

[0108] It should be appreciated that the method steps in the embodiments of the present invention can be implemented or executed by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can be run on a programmed application-specific integrated circuit.

[0109] Furthermore, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that can be executed by one or more processors.

[0110] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, an RSM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted over a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.

[0111] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data that is stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.

[0112] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the aforementioned embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various modifications and variations of the technical solutions and / or implementation methods are possible.

Claims

1. The target detection method based on drone is characterized by: The method comprises the following steps: S100, uses the camera onboard the drone to collect real-time images and video data of the highway; S200, inputs the collected data into a lightweight UAV robust target detection model based on the improved YOLOv7-Tiny; S300: The front end visualizes the detection results returned by the model; The backbone network of the target detection model includes a SPCBlock module improved based on the CS module, and the neck network of the target detection model includes an MDCACBlock module formed by combining an Elan module and an MDCA module; the SPCBlock module includes two SPC modules based on a StarBlock structure combined with partial convolution; The target detection model adopts an adversarial training strategy of an improved projected gradient descent algorithm to load the perturbations generated by the model on the first layer parameters of the model; Wherein, the SPC module is configured to: A partial convolution with a 3×3 convolution kernel combined with a batch normalization layer is used to process a quarter of the input features, reducing the amount of computation while maintaining the accuracy of the model. The features after the above convolution processing are spliced ​​together to form richer semantic information; the spliced ​​features are input into the feature weighting module to extract more critical features; The weighted features are channel compressed through 1×1 convolution, and then processed by batch normalization and SiLU nonlinear activation to further enhance the expressive power of the model. The processed features are connected to the original features of the third branch through residual connection to alleviate the gradient vanishing problem during training and retain the original feature information; The MDCA module is composed of a multi-scale dilated convolution and a CA attention mechanism module; The MDCA module is configured as follows: including three parallel dilated convolution branches, each of which uses a 3×3 dilated convolution with a dilation rate of 3 to capture contextual information of different scales; the captured features of the three dilated convolution branches are spliced ​​again with the convolution output with a dilation rate of 3 to form a richer multi-scale feature representation; the spliced ​​features are fused through 1x1 convolution and dimensionality reduction is performed to reduce the amount of calculation, and the fused features are further input into the CA attention mechanism module to effectively capture the dependencies in the spatial dimension and channel dimension; the original input features are combined with the processed features through the residual connection structure. The input of the module passes through a 1×1 convolution layer to reduce the dimension and extract preliminary features. The preliminary features are transformed nonlinearly through batch normalization and ReLU activation function to enhance the expressiveness of the model and improve the stability of training. The features enter a 3×3 dilated convolution layer with a dilation rate of 3 to greatly expand the receptive field without adding additional parameters. The dilated convolution layer includes a 3×3 convolution kernel with a dilation rate of 3. The output features after the dilated convolution processing are concatenated with the original input features through the Concatenate operation to further enrich the feature expression and strengthen the model's ability to capture features at different levels.

2. The target detection method based on drone according to claim 1, characterized in that: The feature weighting module consists of two sub-branches, one of which consists of a 1×1 convolution and a Sigmoid function, and the other includes a 1×1 convolution.

3. The target detection method based on drone according to claim 2, characterized in that: The weighting process of the feature weighting module is expressed as follows: y=σ(W1*x+b1)⊙(W2*x+b2) Where y represents the output feature, x represents the input feature, W1 and W2 represent the learned weight matrices, b1 and b2 represent the bias terms, σ represents the sigmoid activation function, and ⊙ represents element-wise multiplication.

4. A computer-readable storage medium having program instructions stored thereon, wherein when the program instructions are executed by a processor, the target detection method based on a drone according to any one of claims 1 to 3 is implemented.

5. The target detection system based on drone is characterized by: include: A computer device comprising the computer-readable storage medium according to claim 4.

6. The target detection system based on drone according to claim 5, characterized in that: include: A drone camera acquisition module used to capture real-time images of the drone during flight; Websocket communication module for interaction between the drone's backend commands and model algorithms; The core algorithm processing module is used to receive special instructions from the backend to perform different algorithm detection; Data storage and management module for downloading resources before testing and uploading data after testing; Push-pull stream module for pushing detected images to a specified address; A result visualization module used to combine the streaming address with the front-end interface.

Citation Information

Patent Citations

  • Lightweight small target detection method based on improved YOLOv7

    CN116206185A

  • Ship visible light image target detection method based on improved YOLOv7-tiny

    CN118865273A