Unmanned aerial vehicle-mounted target detection method based on adversarial training strategy

By introducing SPCBlock and MDCA modules into the UAV target detection model and adopting PGD adversarial training strategy, the problems of redundancy in the existing technology, large parameters and insufficient detection accuracy are solved, and efficient, accurate and automated highway disease detection is achieved.

CN120126031AActive Publication Date: 2025-06-10GUANGXI COMPREHENSIVE TRANSPORTATION BIG DATA RES INST +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510179072.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing UAV target detection methods have problems such as redundancy in highway disease detection, large parameters, complex models, and insufficient detection accuracy and recall in highway disease detection, which is difficult to meet the real-time detection needs.

Method used

The robust target detection model of lightweight drone based on improved YOLOv7-Tiny is adopted. By introducing SPCBlock module and MDCA module into the backbone network and neck network, and using the adversarial training strategy of PGD algorithm, the detection accuracy and efficiency of the model are improved.

Benefits of technology

It improves the detection efficiency of highway road surface and surrounding environment, reduces costs, ensures the safety of drivers' lives, realizes automatic identification and monitoring of highway diseases, and provides fast and convenient technical support for traffic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126031A_ABST
    Figure CN120126031A_ABST
Patent Text Reader

Abstract

The invention relates to an unmanned aerial vehicle-mounted target detection method based on an adversarial training strategy. The method comprises the steps that image and video data of a road are collected in real time through a camera carried by an unmanned aerial vehicle, the collected data are input into a lightweight unmanned aerial vehicle robust target detection model based on improved YOLOv7-Tiny, and a front end carries out visualization processing on a detection result returned by the model, a backbone network of the target detection model comprises an improved SPCBlock module based on a CS module, and a neck network of the target detection model comprises an MDCACBlock module formed by combining an ELin module and an MDCA module; the target detection model adopts an adversarial training strategy of an improved projection gradient descent algorithm, so that disturbance generated by the model is loaded on a first-layer parameter of the model. According to the invention, a light-weight module with efficient feature processing capability is adopted to replace a corresponding module of a traditional model, and the model is trained by an interfered sample, so that effective detection and monitoring of the unmanned aerial vehicle on the road are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting unmanned aerial vehicle targets based on an adversarial training strategy, and belongs to the fields of intelligent transportation, artificial intelligence, image segmentation and medical image processing. Background Art

[0002] With the rapid development of the transportation industry, highways have become a key infrastructure for modern social and economic activities, connecting cities, promoting logistics, and improving economic development efficiency. With the continuous growth of society's demand for transportation, the highway network is also expanding year by year, and the number of vehicles passing and traffic load has increased significantly. With the continuous growth of traffic pressure and the impact of the natural environment, highway pavement and structural facilities are gradually facing the threat of various diseases, such as pavement diseases such as cracks and potholes, local sinking and damage of shoulders, bridge cracks and expansion joint failures, as well as tunnel lining cracks, water leakage and seepage, and support structure damage. Therefore, it is of great practical significance to efficiently and accurately identify and monitor highway diseases.

[0003] Among them, intelligent transportation is a key direction for the development of modern cities, aiming to improve the efficiency and safety of transportation systems by using advanced information technology and data analysis methods. With the rapid development of intelligent transportation systems, traditional fixed cameras can no longer meet the needs of large-scale, highly maneuverable traffic monitoring. In this context, drones are widely used in intelligent traffic management and video surveillance due to their advantages of strong operability, wide viewing angle, high flexibility and low cost.

[0004] In the existing technology, it is possible to collect pictures of various road surface diseases on highways manually at first. In order to save human resources, reduce unnecessary time waste, and improve the detection rate, drone inspection is currently more often used. The drone collects the road conditions of the target section and stores them locally, and finally performs manual visual inspection on the collected videos or images. Compared with manual data collection, the high performance and strong maneuverability of drones greatly improve the data collection rate, but the analysis of data after collection still requires a lot of manpower and time, and the detection efficiency is low.

[0005] At this time, the rapid progress of computer vision and deep learning technology has brought new opportunities for drone traffic monitoring. The current target detection can be divided into two stages in time: target detection based on traditional manual feature design and target detection based on deep learning. The target detection based on traditional manual features is to obtain the area to be detected by sliding window technology, extract the features of the area to be detected, and finally pass through the classifier and regressor to achieve multi-target classification and regression. The target detection based on deep learning extracts target features based on convolutional neural networks, and is divided into two-stage and one-stage according to whether there are candidate boxes. Among them, the two-stage target detection is to first divide the candidate area that may contain the target and then obtain the target area through feature extraction and classification regression. It performs best in terms of target accuracy and recall rate, but cannot meet the real-time requirements in terms of detection rate. One-stage target detection is to directly classify and regress through convolutional neural networks. The detection speed is fast and can meet the requirements of real-time detection, but it is not as good as two-stage in terms of accuracy and recall rate. Therefore, the existing target detection algorithm model has redundant calculations, a large number of parameters, and a large model. In addition, the existing lightweight target detection algorithm has a need to improve its accuracy. For example, although the YOLO algorithm is famous for its real-time processing capabilities, the target detection algorithm improved based on YOLO has many parameters and a large model, and the performance of the edge devices carried by drones is limited, making it difficult to effectively run such algorithms. Although many scholars have developed lightweight detection algorithms for drone scenarios, most lightweight algorithms generally achieve lightweightness at the expense of target accuracy. Summary of the invention

[0006] The present invention provides a method for detecting unmanned aerial vehicle targets based on an adversarial training strategy, aiming to solve at least one of the technical problems existing in the prior art.

[0007] The technical solution of the present invention relates to a target detection method based on a drone, and the method according to the present invention comprises the following steps:

[0008] S100, collects images and video data of the highway in real time through the camera carried by the drone;

[0009] S200, inputs the collected data into the lightweight UAV robust target detection model based on the improved YOLOv7-Tiny;

[0010] S300, the front end performs visualization processing on the detection results returned by the model;

[0011] The backbone network of the target detection model includes a SPCBlock module improved based on the CS module, and the neck network of the target detection model includes an MDCACBlock module formed by combining an Elan module and an MDCA module;

[0012] The target detection model adopts an adversarial training strategy of an improved projected gradient descent algorithm to load the disturbance generated by the model on the first layer parameters of the model.

[0013] Furthermore, in the target detection model,

[0014] The SPCBlock module includes two SPC modules based on the StarBlock structure combined with partial convolution.

[0015] Further, the SPC module is configured to:

[0016] A partial convolution with a 3×3 convolution kernel combined with a batch normalization layer is used to process a quarter of the input features, which reduces the amount of computation while maintaining the accuracy of the model.

[0017] The features after the above convolution processing are concatenated to form richer semantic information; the concatenated features are input into the feature weighting module to extract more critical features;

[0018] The weighted features are compressed through 1×1 convolution, and then processed by batch normalization and SiLU nonlinear activation to further enhance the expressiveness of the model.

[0019] The processed features are residually connected with the original features of the third branch to alleviate the gradient vanishing problem during training and retain the original feature information.

[0020] Furthermore, the feature weighting module consists of two sub-branches, one of which consists of a 1×1 convolution and a Sigmoid function, and the other of which includes a 1×1 convolution.

[0021] Furthermore, the weighting process of the feature weighting module is expressed as follows:

[0022] y=σ(W 1 *x+b 1 )⊙(W 2 *x+b 2 )

[0023] Where y represents the output feature, x represents the input feature, W1 and W2 represent the learned weight matrices, b1 and b2 represent the bias terms, σ represents the sigmoid activation function, and ⊙ represents element-wise multiplication.

[0024] Furthermore, in the target detection model,

[0025] The MDCA module is composed of a multi-scale dilated convolution and a CA attention mechanism module;

[0026] Among them, the MDCA module is configured to:

[0027] Include three parallel dilated convolutional branches, each of the dilated convolutional branches using 3×3 dilated convolution with a dilation rate of 3 respectively to capture context information at different scales;

[0028] The captured features of the three dilated convolutional branches are concatenated again with the convolution output with a dilation rate of 3 to form a richer multi-scale feature representation;

[0029] The above fused features are input into the CA attention mechanism module to effectively capture the dependencies in the spatial dimension and channel dimension;

[0030] The original input features and the processed features are added through a residual connection structure.

[0031] Furthermore, the MDCA module is configured to:

[0032] The input of the module passes through a 1×1 convolutional layer to reduce the dimension and extract preliminary features;

[0033] The preliminary features are non-linearly transformed through batch normalization and the ReLU activation function to enhance the expression ability of the model and improve the stability of training;

[0034] The features enter a 3×3 dilated convolutional layer with a dilation rate of 3 to greatly expand the receptive field without increasing additional parameters; among them, the dilated convolutional layer includes a 3×3 convolutional kernel with a dilation rate of 3;

[0035] The output features after dilated convolution processing and the original input features are concatenated through a Concatenate operation to further enrich the feature expression and enhance the model's ability to capture features at different levels.

[0036] The technical solution of the present invention also relates to a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the above method is implemented.

[0037] The technical solution of the present invention also relates to a target detection system based on a drone, and the system includes a computer device, and the computer device includes the above computer-readable storage medium.

[0038] Furthermore, the target detection system includes: a drone camera acquisition module for acquiring real-time images during the flight of the drone; a Websocket communication module for the interaction between the backend instructions of the drone and the model algorithm; a core algorithm processing module for receiving special instructions sent by the backend to perform different algorithm detections; a data storage and management module for resource downloading before detection and data uploading after detection; a push-pull stream module for pushing the detected images to a specified address; and a result visualization module for combining the push stream address with the front-end interface.

[0039] The beneficial effects of the present invention are as follows:

[0040] The present invention improves the lightweight yolov7-tiny target detection algorithm, replaces the corresponding modules of the traditional model with lightweight modules with efficient feature processing capabilities, and uses the disturbed samples to train the model. The present invention improves the backbone network and the neck network of the model respectively by adopting the SPC module and the MDCA module. Both modules adopt lightweight designs, with simple and efficient structures, which can improve the target detection accuracy of the model while ensuring lightweight, and introduce the CA attention mechanism into the MDCA module, which effectively improves the attention ability to key target areas through the adaptive adjustment of channel and spatial information. The present invention introduces an adversarial training strategy based on the PGD algorithm. Compared with directly disturbing the samples, it uses the disturbed samples to train the model, and by perturbing the parameters of the first layer of the model, the original prediction results are changed, so as to achieve the effect of disturbing the model and realizing the effect of adversarial training.

[0041] The present invention improves the detection efficiency of highway pavements and the surrounding environment, reduces costs, and ensures the safety of drivers. It proposes a drone-based target detection method for automatic identification and monitoring of highway diseases, providing a fast and convenient technology for vehicles to obtain real-time information about the surrounding environment of highways and for the safety and maintenance inspections of highways, enabling it to effectively serve the industries related to drones and traffic management. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is the basic flowchart of the method according to the present invention.

[0043] Figure 2 is the flowchart of the target detection algorithm of the method according to the present invention.

[0044] Figure 3 is the schematic diagram of the Yolov7-Tiny module structure before improvement according to the embodiment of the present invention.

[0045] Figure 4 is the schematic diagram of the Yolov7-Tiny module structure after improvement according to the embodiment of the present invention.

[0046] Figure 5 It is a schematic structural diagram of the Elan module before improvement according to an embodiment of the present invention.

[0047] Figure 6 It is a schematic structural diagram of the SPCBlock module after improving the Elan module according to an embodiment of the present invention.

[0048] Figure 7 It is a schematic structural diagram of the SPC module according to an embodiment of the present invention.

[0049] Figure 8 It is a Pconv processing feature flow chart according to an embodiment of the present invention.

[0050] Fig. 9 It is a schematic structural diagram of the improved MDCABlock module according to an embodiment of the present invention.

[0051] Fig.10 It is a schematic structural diagram of the MDCA module according to an embodiment of the present invention. Detailed implementation manners

[0052] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments and the drawings, so as to fully understand the purpose, solution and effects of the present invention.

[0053] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. The singular forms of "a", "the" and "said" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0054] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all example or exemplary language (such as "for example", "such as", etc.) provided herein is only intended to better illustrate the embodiments of the present invention and will not impose a limitation on the scope of the present invention unless otherwise required.

[0055] Refer to Figures 1 to 10 , in some embodiments, the drone-based target detection method according to the present invention at least includes the following steps:

[0056] S100. Real-time collect image and video data of the road through the camera carried by the drone;

[0057] S200. Input the collected data into a lightweight drone robust target detection model based on the improved YOLOv7-Tiny;

[0058] S300. The front end visually processes the detection results returned by the model;

[0059] Among them, the backbone network of the target detection model includes an SPCBlock module improved based on the CS module, and the neck network of the target detection model includes an MDCACBlock module formed by combining an Elan module and an MDCA module;

[0060] Among them, the target detection model adopts an adversarial training strategy of an improved projected gradient descent algorithm to load the perturbations generated by the model on the parameters of the first layer of the model.

[0061] In some embodiments, the drone module of the system of the present invention includes a drone wireless communication module, a high-definition camera module, a drone obstacle avoidance module, and a GPS positioning module. Among them, the wireless communication module on the drone establishes a data communication link with the data transmission device of the ground station, that is, the server, encodes and modulates the data collected by the camera and the GPS positioning information, and sends it through the link as a 5.8 GHz wireless signal. The receiving server on the ground receives the data, decodes, and demodulates the received data. The high-definition camera is mainly used to collect the scenes passed by the drone, perform image or video collection on it, and transmit the collected data to the wireless communication module of the drone. The drone obstacle avoidance module is usually equipped with various sensors for detecting the surrounding environment, such as lidar, ultrasonic sensors, infrared sensors, etc. It obtains information such as the position, size, and distance of obstacles through the sensors, establishes an environmental perception model to cooperate with the high-definition camera of the drone, and prevents the drone from contacting obstacles and thus affecting the trajectory of the drone. The GPS module is mainly used to obtain the geographical location of the current drone and interact with the ground. Through the GPS positioning module, the flight trajectory of the drone can be accurately grasped to prevent trajectory deviation. The GPS module calculates information such as the current longitude and latitude coordinates, altitude, and speed of the drone according to the received satellite signals and transmits this information to the wireless communication module. Subsequently, the specific navigation position of the drone can be observed in real time on the front-end interface.

[0062] Compared with previous drone target detection methods and systems, the present invention combines deep learning-based methods. Through the machine learning method of imitating the human brain's judgment of external things by deep learning itself, the manual judgment time can be greatly reduced, and the utilization rate of time and the work efficiency can be improved. At the same time, the present invention is an automated detection based on drones, and its automation is mainly reflected in that the algorithm detection process from data collection to target, as well as the subsequent information storage and push / pull stream processes, all operate as a whole without manual intervention. Among them, drones have the characteristics of flexibility and small shape and volume, and have natural advantages for data collection. In addition, the communication module carried by the drone is connected to the server, and the collected data can be transmitted wirelessly to a high-performance server. And according to the mode of the drone, the corresponding detection method can be selected to classify various different highway diseases, detect problems such as whether they are damaged and need to be updated in real time, and transmit the detection results to the central computer that initiates the detection in real time. The detected targets and the targets with problems are classified in different annotation methods and visualized on the central computer, and the results are transmitted to the detector, which is convenient for subsequent updates and monitoring. Moreover, more algorithm modules can be expanded on this basis to achieve multi-functional automated drone inspections. Compared with manual inspections, drone detection has lower costs and higher efficiency, without the need for personnel to enter the site, avoiding potential personal safety hazards. At the same time, the present invention performs cloud backup operations on the detected video streams or pictures, which is convenient for subsequent review and inspection of the highway pavement conditions at a certain time.

[0063] It should be noted that the present invention adopts an automated recognition, inspection and monitoring method for highway diseases based on drones, improves the functions based on deep learning and the python library, and uses the object detection technology of the one-stage yolov5.7.0 version based on deep learning. Among them, due to the need to meet the purpose of real-time detection, there are certain limitations on the rate of object detection. According to the development process of object detection, in object detection based on deep learning, the two-stage object detection method has high accuracy but low detection efficiency, and the rate of 3-4 frames per second cannot meet the requirements. Therefore, the present invention adopts the one-stage object detection method and utilizes the advantages of the fast speed of the yolo series methods (45 FPS per second) and the good stability and good integration of yolov5, so that the present invention can meet the real-time detection requirements.

[0064] In some embodiments, the automated detection system of the present invention includes a drone camera acquisition module, a Websocket communication module, a core algorithm processing module, a data storage and management module, a push / pull stream module, and a result visualization module.

[0065] In an application embodiment, the camera acquisition module of the present invention is used to acquire real-time images during the flight of the drone. Specifically, the front-end interface controls the drone to leave the warehouse and take off vertically. When it is about 150 meters above the ground, the drone adjusts its posture to be horizontal, and the camera of the drone flies to the designated inspection location at an angle of approximately 75 degrees with the ground. Then the camera is turned on, and the drone flies along the designated route and pushes the acquired images to a public network address in real time for subsequent algorithm pulling and detection, and saves the acquired real-time images to the local network.

[0066] In an application embodiment, the Websocket module of the present invention is used to interact the backend instructions of the drone with the algorithm. Specifically, the Websocket is built using the Websockets module of the python library, which can communicate with the backend in real time asynchronously and send heartbeat packets to the backend regularly to prevent the disconnection between the Websocket module and the backend, and perform anomaly detection on the Websocket to implement the disconnection reconnection mechanism for communication.

[0067] In an application embodiment, the core algorithm processing module of the present invention is used to realize the mutual switching between algorithms and different models trained by different algorithms. Specifically, before detection, a download folder for storing data downloaded from the drone or the database is constructed, and a folder and output address for outputting the detection results are set up. The core algorithm processing module is after the Websocket module and performs different algorithm detections when receiving special instructions sent from the backend. For example: when receiving the RTSP keyword sent from the backend, the real-time detection of the drone is started, otherwise the offline static pictures or video streams are detected.

[0068] Furthermore, for the core algorithm processing module, in the part of real-time detection of the data collected by the drone, the detection model needs to be trained in advance using yolo. The os module is used to start the model through the command line operation, and the pulling operation is performed on the address where the drone pushes the stream. The video is detected frame by frame and the output video after synthesis detection is output.

[0069] Furthermore, for the core algorithm processing module, in the part of offline static pictures and video streams, first, the pictures or video streams to be detected are locally downloaded using the multi-process method for the video or pictures, that is, the request is sent using post, and the corresponding resources are saved locally by name, and the time required for downloading is recorded using the time module. Subsequently, the resources are also subjected to model detection in the same way as real-time detection, and the detection results are output to the local folder waiting to be uploaded to the minio storage space.

[0070] In an application embodiment, the data storage and management module of the present invention is used for resource download before detection and data transmission to the oss database after detection. Specifically, before detection, mainly a resource recursive folder is constructed through the os.makedir function in Python. After detection, the detection result address is put into the Bucket through the oss2 module, thereby realizing the upload of the detection result.

[0071] In an application embodiment, the push-pull stream module of the present invention mainly uses the push stream technology of FFmpeg for each frame of the video or picture in the detect.py file of yolov5.7.0 during the algorithm execution, and pushes the detected picture in the format of 640*480 to the specified address, realizing the broadcast feature of the detection.

[0072] In an application embodiment, the result visualization module of the present invention includes the push stream address by using the video tag in HTML technology or the video playback library Video.js, realizing the integration of the push stream address and the front-end interface, and providing a convenient picture for data monitoring personnel to facilitate whether the highway pavement and surrounding protection equipment need to be updated or repaired.

[0073] See Figure 1 and Figure 2 In the method for automatic identification and monitoring of highway diseases based on drones of the present invention, first, adjust the take-off attitude and angle of the drone and reach the target area, then turn on the drone camera, collect the data of the target area and live broadcast it to the target address, pull the live video stream and execute the algorithm to push the detection result to the specified address for live broadcast, and finally visualize the live content at the front end. Further, the processing process of its target detection model includes: first, send a detection instruction to the drone, conduct two-way communication through Websocket, and then judge whether RTSP is 1. If RTSP is not 1, then construct and download static pictures and the collected videos to the local directory, then take out the local pictures or videos in turn for detection, and then the target detection model detects and returns the result; if RTSP is 1, then pull the stream from the target address → the target detection model detects and returns the result → push the result to the specified target address for live broadcast. Then, upload the detection result to the oss cloud, and finally the front end plays the detection result in real time.

[0074] See Figures 3 to 10 In the drone-based target detection method of the technical solution of the present invention, aiming at the problems of redundant calculation, large number of parameters, and large model size in the existing drone target detection model, as well as the problem that the accuracy of the existing lightweight target detection algorithm needs to be improved, the present invention improves the lightweight yolov7-tiny target detection algorithm, replaces the corresponding modules of the traditional model with lightweight and highly efficient feature processing modules, and changes the model training strategy.

[0075] Specifically, the present invention improves the backbone network and the neck network of the model respectively. Specifically, the present invention designs two modules respectively: the SPC (StarPConv) module and the MDCA (Multi-scale Dilated Convolution Coordinate Attention) module. The characteristics of both modules are lightweight design, with simple and efficient structures, which can improve the object detection accuracy of the model while ensuring lightweight.

[0076] In an application embodiment, the SPC module of the present invention focuses on enhancing the efficiency of feature extraction by simplifying the calculation process in the processing of UAV aerial photography images. Through partial convolution (PConv) operations, the SPC module reduces redundant calculations while ensuring effective capture of the target area. The SPC module of the present invention can streamline the network structure, enabling it to operate efficiently on resource-constrained UAV devices, especially suitable for scenarios that require quick response. Specifically, the SPC module weights the features simultaneously, effectively improving the recognition ability of key areas, ensuring that important features are enhanced and highlighted, while irrelevant information is minimized, so that the SPC module can quickly locate the target area in aerial photography images, avoiding excessive calculation of irrelevant backgrounds, and improving the efficiency and accuracy of image processing.

[0077] In an application embodiment, the MDCA module of the present invention has significant lightweight characteristics in the processing of UAV aerial photography images, and can extract rich features of key targets while maintaining computational efficiency. Through multi-scale dilated convolution operations, the module effectively expands the receptive field to meet the object detection requirements brought by different heights and perspective changes during UAV flight. The MDCA module helps to ensure that the precise capture of the target will not be lost due to scale differences during the processing, especially maintaining high-resolution detail performance in complex scenarios.

[0078] In an application embodiment, the MDCA module of the present invention introduces the CA attention mechanism (Coordinate Attention), and effectively improves the attention ability to key target areas through adaptive adjustment of channel and spatial information. The CA attention mechanism can further optimize the extraction of important features, enhance the object detection accuracy of the model in complex backgrounds, and maintain a lightweight design at the same time, ensuring that it can operate with higher efficiency when processing high-resolution images.

[0079] In an application embodiment, refer to Figure 3 and Figure 4As shown in the improved Yolov7-Tiny model structure before and after, the present invention has made improvements in the backbone network and neck network of yolov7-tiny. Further, referring to Figure 5 and Figure 6 As shown in the Elan module before and after improvement, the SPCBlock module improved based on the Elan model replaces the only two 3*3 convolutions in the Elan model with the SPC (StarPconv) module, where the structure of the SPC module is shown in Figure 7 The SPC module of the present invention draws on the StarBlock structure in StarNet and incorporates the Pconv (Partial Convlution) proposed by FasterNet.

[0080] In an application embodiment, the SPC module of the present invention is a lightweight and efficient convolutional structure, drawing on the design concept of StarBlock. Referring to Figure 7 and Figure 8 The SPC module of the present invention consists of three main branches.

[0081] Among them, the first branch of the SPC module uses partial convolution (PConv) combined with a batch normalization layer. PConv uses a 3×3 convolution kernel, which is different from traditional convolution. The PConv of the present invention only processes one-fourth of the input features. Specifically, based on the observation of the channel feature map, it is found that there is a high degree of similarity and redundancy between different channels. The present invention adopts the method of only processing some channels, making PConv. For example, the memory access amount of ordinary convolution is h×w×2c + k 2 ×c 2 ≈h×w×2c (Equation 1), while the memory access amount of the PConv of the present invention can reach where c is the number of input channels, c p is equal to 1 / 4C, thus greatly reducing the computational requirements of PConv. Among them, h and w are the height and width of the feature map respectively, and k is the convolution kernel size.

[0082] Among them, the second branch of the SPC module does not directly process the features, but concatenates these features with the features processed by PConv, thus forming richer semantic information. Subsequently, these concatenated features enter a feature weighting module to extract more critical features. Specifically, the weighting module consists of two sub-branches: one is a branch composed of a 1×1 convolution and a Sigmoid function, and the other only contains a 1×1 convolution. Thus, the present invention not only maintains the lightweight of the module, but also introduces a simple attention mechanism to enhance the feature extraction ability.

[0083] Further, the feature weighting process can be expressed by the following formula: y = σ(W 1 *x + b 1 ) ⊙ (W 2 *x + b 2 )(Equation 3), where x is the input feature, W1 and W2 are the learned weight matrices, b1 and b2 are the bias terms, σ represents the sigmoid activation function, and ⊙ represents element-wise multiplication.

[0084] The weighted features are compressed in channels through 1×1 convolution, and then through batch normalization and SiLU non-linear activation processing to further enhance the model's expressive ability. Finally, these processed features are connected to the original features of the third branch by residual connection, which can alleviate the problem of gradient disappearance during training and effectively retain the original feature information.

[0085] In an application embodiment, in the neck network of the present invention, an MDCA module is added after its Elan module to form a new MDCACBlock module, as shown in Figure 5 and Fig. 9 shown. Among them, the MDCA module of the present invention is composed of a multi-scale dilated convolution and a CA attention mechanism, as shown in Figure 6 shown. It should be noted that MDCA (Multi-scale DilatedConvolution Coordinate Attention) is an efficient feature extraction structure that combines multi-scale dilated convolution with coordinate attention mechanism. The MDCA module can expand the receptive field and capture multi-level features while maintaining high computational efficiency.

[0086] Specifically, in the MDCACBlock module formed by the combination of the present invention, first, the input of the module passes through a 1×1 convolutional layer, which plays the role of dimensionality reduction and preliminary feature extraction. Then, the features pass through batch normalization (BN, BatchNormalization) and the ReLU activation function for non-linear transformation to enhance the model's expressive ability and improve the stability of training. Next, the features enter a 3×3 dilated convolutional layer with a dilation rate of 3. This dilated convolution is the core component of the MDCA module, which can greatly expand the receptive field without adding extra parameters. Specifically, the receptive field of a 3×3 convolutional kernel with a dilation rate of 3 is equivalent to that of a 7×7 ordinary convolutional kernel, but its number of parameters is only 9, which is about 18.37% of the number of parameters of a 7×7 convolutional kernel. The above network model structure of the present invention enables the model to capture large-range information while still maintaining high computational efficiency. Then, the output features after dilated convolution processing and the original input features are concatenated through the Concatenate operation to further enrich the feature expression and enhance the model's ability to capture different-level features.

[0087] See Fig.10 The multi-scale characteristics of the MDCA model are realized by the following three parallel dilated convolutional branches. Specifically, each branch uses a 3×3 dilated convolution with a dilation rate of 3 to progressively extract features. These captured features are then concatenated again with the output of the initial dilated convolution with a dilation rate of 3, and then the concatenated features are input into a 1x1 fusion of multi-scale features to form a richer multi-scale feature representation. Such a progressive dilated convolution and feature concatenation strategy enables it to capture both local and broader context information simultaneously, enhancing the model's perception ability of diverse information in the scene.

[0088] Next, after feature fusion in the MDCA module, the MDCA module introduces the CA attention mechanism (see Fig.10 CAteention). The CA attention mechanism can effectively capture the dependencies in the spatial and channel dimensions, improve the model's attention to key features, and thus enhance the model's performance in complex scenarios.

[0089] Finally, the MDCA module adopts a residual connection structure, adding the original input features to the processed features, which not only helps to alleviate the vanishing gradient problem in deep networks but also retains the original information, facilitating the network to learn the identity mapping, thereby improving the model's training effect and optimization ability.

[0090] In some embodiments, the present invention introduces an adversarial training strategy based on the PGD (Projected Gradient Descent) algorithm. By introducing adversarial training, the robustness of the model can be improved, such that the model can still maintain a certain accuracy even for images with interference. Compared with directly interfering with samples, training the model with the interfered samples, the present invention perturbs the parameters of the first layer of the model to change the original prediction result, thereby achieving the effect of interfering with the model and realizing the effect of adversarial training.

[0091] It should be noted that the original PGD iteratively generates adversarial samples for a batch of samples. Specifically, after the traditional model performs forward propagation and finally calculates the loss, the corresponding gradients are backpropagated to update the model parameters. When the gradients are backpropagated to the first layer, the gradients at this moment are used to generate perturbations and interfere with the clean samples, further generating adversarial samples to train the model. This process is repeated a certain number of times, continuously maximizing the loss to form adversarial gradients to generate perturbations and update the model parameters. After the iteration is completed, the next batch is interfered with.

[0092] The present invention improves the interference method. Specifically, although the overall idea is still based on adversarial training of PGD, the perturbation generated by the model of the present invention is not added to the clean sample, but to the parameters of the first layer of the model. The present invention changes the prediction result by perturbing the parameters, which has the same effect as directly interfering with the sample, and is convenient with high computational efficiency without the need to additionally generate adversarial samples. Among them, the parameter update of the present invention is as shown in the following formula:

[0093]

[0094] Specifically, the process of adversarial training based on the PGD (Projected Gradient Descent) algorithm is as follows:

[0095] Among them, model input: model M, parameters θ of the first layer of the model 0 , training data X, label y, step size α, perturbation range S, number of iterations K, learning rate η. Model output: model M with enhanced robustness.

[0096] A1. θ adv ←θ 0 / / Initialize adversarial parameters.

[0097] A2. θ original ←θ adv / / Backup the original parameters.

[0098] A3. / / Calculate the initial gradient

[0099] Execute K - 1 times of loop.

[0100] A4. / / Update the parameters, where Π represents the projection operator, and project the parameters into a sphere centered at θ original with a radius of S.

[0101] After the last iteration, the network parameters are updated.

[0102] A5.

[0103]

[0104] A6. Return M / / Return the model with enhanced robustness.

[0105] Here, a specific embodiment is used for illustration.

[0106] First, download the public dataset. Here, the road damage dataset from the perspective of drones (UAV-PDD2023) is selected. The dataset contains a total of 2,440 three-channel JPG format images and corresponding VOC format annotation files. Six types of road damages are marked in the images, namely longitudinal cracks (LC), transverse cracks (TC), crocodile cracks (AC), oblique cracks (OC), repairs (RP), and potholes (PH). Then, convert the VOC format labels of the dataset into yolo's txt format files, and randomly divide the dataset into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0107] Construct the drone target detection model based on the improved yolov7-tiny of the present invention and perform model training. Adjust the iteration times of PGD adversarial training, as well as training strategies such as the batchsize size and the number of epochs according to the implementation results. The batchsize is selected to be set to 16, the number of epochs is set to 150, the learning rate is 0.01, and the optimizer is Adam. After training is completed, retain the optimal weight model and conduct experimental tests on the trained model using the divided test set. It should be noted that the experimental platform of the embodiment of the present invention runs based on the Linux Ubuntu22 operating system. The training platform uses an Nvidia GeForce RTX 3090 24G GPU, an Intel(R) Xeon(R) CPU E5-2697A v4 @ 2.60GHz, with 64G of memory. The code running framework is PyTorch, and the running environment is Python3.7 and CUDA 12.4.

[0108] It should be recognized that the method steps in the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit for programming.

[0109] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed commonly on one or more processors, by hardware, or by a combination thereof. The computer program includes multiple instructions executable by one or more processors.

[0110] Further, the method can be implemented in any type of computing platform operably connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RSM, ROM, etc., such that it can be read by a programmable computer and can be used to configure and operate the computer to execute the processes described herein when the storage medium or device is read by the computer. In addition, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. When such media includes instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the inventions described herein include these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.

[0111] The computer program can be applied to the input data to perform the functions described herein, thereby transforming the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.

[0112] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. As long as the same means are used to achieve the technical effects of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners can have various different modifications and changes.

Claims

1. The target detection method based on drone is characterized by: The method comprises the following steps: S100, collects images and video data of the highway in real time through the camera carried by the drone; S200, inputs the collected data into the lightweight UAV robust target detection model based on the improved YOLOv7-Tiny; S300, the front end performs visualization processing on the detection results returned by the model; The backbone network of the target detection model includes a SPCBlock module improved based on the CS module, and the neck network of the target detection model includes an MDCACBlock module formed by combining an Elan module and an MDCA module; The target detection model adopts an adversarial training strategy of an improved projected gradient descent algorithm to load the disturbance generated by the model on the first layer parameters of the model.

2. The method according to claim 1, characterized in that In the target detection model, The SPCBlock module includes two SPC modules based on the StarBlock structure combined with partial convolution.

3. The method according to claim 2, characterized in that The SPC module is configured to: A partial convolution with a 3×3 convolution kernel combined with a batch normalization layer is used to process a quarter of the input features, which reduces the amount of computation while maintaining the accuracy of the model. The features after the above convolution processing are concatenated to form richer semantic information; the concatenated features are input into the feature weighting module to extract more critical features; The weighted features are compressed through 1×1 convolution, and then processed by batch normalization and SiLU nonlinear activation to further enhance the expressiveness of the model. The processed features are residually connected with the original features of the third branch to alleviate the gradient vanishing problem during training and retain the original feature information.

4. The method according to claim 3, characterized in that The feature weighting module consists of two sub-branches, one of which consists of a 1×1 convolution and a Sigmoid function, and the other includes a 1×1 convolution.

5. The method according to claim 5, characterized in that: The weighting process of the feature weighting module is expressed as follows: y=σ(W1*x+b1)⊙(W2*x+b2) Where y represents the output feature, x represents the input feature, W1 and W2 represent the learned weight matrices, b1 and b2 represent the bias terms, σ represents the sigmoid activation function, and ⊙ represents element-wise multiplication.

6. The method according to claim 5, characterized in that In the target detection model, The MDCA module is composed of a multi-scale dilated convolution and a CA attention mechanism module; Wherein, the MDCA module is configured as follows: It includes three parallel dilated convolution branches, each of which uses a 3×3 dilated convolution with a dilation rate of 3 to capture contextual information of different scales; The captured features of the three dilated convolution branches are concatenated again with the convolution output with a dilation rate of 3 to form a richer multi-scale feature representation; The above concatenated features are fused through 1x1 convolution and dimensionality reduction is performed to reduce the amount of calculation. The fused features are further input into the CA attention mechanism module to effectively capture the dependencies in the spatial and channel dimensions. The original input features are added to the processed features through a residual connection structure.

7. The method according to claim 6, characterized in that The MDCA module is configured to: The input of the module passes through a 1×1 convolutional layer to reduce the dimension and extract preliminary features; The preliminary features are transformed nonlinearly through batch normalization and ReLU activation function to enhance the expressiveness of the model and improve the stability of training; The features enter a 3×3 dilated convolution layer with a dilation rate of 3 to significantly expand the receptive field without adding additional parameters; wherein the dilated convolution layer includes a 3×3 convolution kernel with a dilation rate of 3; The output features after the dilated convolution process are concatenated with the original input features through the Concatenate operation to further enrich the feature expression and strengthen the model's ability to capture features at different levels. 8 . A computer-readable storage medium having program instructions stored thereon, wherein the program instructions, when executed by a processor, implement the method according to any one of claims 1 to 7.

9. The target detection system based on drone is characterized by: include: A computer device comprising a computer readable storage medium according to claim 9.

10. The target detection system based on drone according to claim 9, characterized in that: include: A drone camera acquisition module used to collect real-time images of drones in flight; Websocket communication module for interaction between the drone's backend commands and model algorithms; A core algorithm processing module for receiving special instructions from the backend to perform different algorithm tests; Data storage and management module for downloading resources before testing and uploading data after testing; A push-pull stream module used to push the detected image to the specified address; A result visualization module used to combine the streaming address with the front-end interface.

Citation Information

Patent Citations

  • Lightweight small target detection method based on improved YOLOv7

    CN116206185A

  • Ship visible light image target detection method based on improved YOLOv7-tiny

    CN118865273A

  • Light regulation method, system, and apparatus for growth environment of leafy vegetables

    US11978210B1