Deep Learning-Based UAV Target Intelligent Recognition and Perception System and Method
By adopting improved YOLOv5 model and key point detection technology in the drone system, combined with trajectory matching and counters, the repetition and omission problems of drones when identifying wildlife populations are solved, achieving more accurate statistics and better real-time performance.
Patent Information
- Application Number
- CN202510175096.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-18
AI Technical Summary
When drones identify wildlife numbers, due to the small appearance differences between wildlife individuals and complex activities, count statistics are prone to duplication or omission.
An intelligent drone target recognition and perception system based on deep learning is proposed, and the improved YOLOv5 model is used for target detection, combining key point detection and trajectory matching, the visibility score and confidence of key points are calculated, feature descriptors are constructed, and accurate statistics of target counts are ensured through counters.
It improves the accuracy of wildlife counting in the target area, reduces duplicate counting and omissions, enhances the recognition and tracking performance, and has better real-time and environmental adaptability.
Smart Images

Figure CN119649254B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an intelligent drone target recognition and perception system and method based on deep learning. Background Art
[0002] Intelligent drone target recognition and perception is an important research field in the drone system. It mainly involves using technologies such as computer vision, image processing, and pattern recognition to achieve functions such as automatic detection, recognition, and tracking of targets by drones.
[0003] The Chinese patent with the authorization announcement number CN116310894B discloses an intelligent recognition method for small-sample and small-target Tibetan antelopes based on drone remote sensing. It includes collecting orthophotos of the drone by the drone; constructing an auxiliary set and a small-sample library of drone Tibetan antelopes; constructing a small-sample deep learning model with context-aware fusion and contrast analysis; determining the model loss function; and training the small-sample deep learning model with context-aware fusion and contrast analysis using the auxiliary set and the small-sample library of drone Tibetan antelopes.
[0004] When counting wild animals by drones to count their numbers, due to the small appearance differences between individual wild animals and the complex animal activities, there are prone to repetitions or omissions in the counting and statistics of wild animals. Summary of the Invention
[0005] This application aims to solve at least one of the technical problems in the related technologies to some extent. For this reason, one objective of this application is to propose an intelligent drone target recognition and perception system and method based on deep learning to improve the accuracy of counting the number of wild animals in the target area.
[0006] One aspect of this application provides an intelligent drone target recognition and perception method based on deep learning, including:
[0007] Step S100: Divide the target area according to the geographical scope of the wildlife reserve, extract the important features of each target area and calculate the regional importance score, generate a target area monitoring plan, and the drone obtains continuous frame images based on the target area monitoring plan;
[0008] Step S200: Use an improved YOLOv5 model for target detection, input the continuous frame images, output the bounding boxes of suspected target wild animals in each frame of image and map them to the actual geographical coordinate system to obtain the actual position coordinates of the suspected target wild animals in each frame of image;
[0009] Step S300: Detect key points of the t-th frame image, extract the coordinates of the key points of the suspected target wild animals, calculate the visibility scores of the key points and the confidence of each key point, and the K key points with the highest confidence constitute the feature descriptor;
[0010] Step S400: Perform object detection and key point detection on the (t + 1)-th frame image to obtain the actual position coordinates and feature descriptors of the candidate objects, perform trajectory matching between the suspected target wild animals and the candidate objects. If they match, update the actual position coordinates and feature descriptors of the suspected target wild animals. If they do not match, define the candidate object as a new target animal;
[0011] Step S500: According to the trajectory matching results between the suspected target wild animals and the candidate objects, count the identified suspected target wild animals and new target animals through a counter to obtain the total number of target wild animals in the target area.
[0012] Specifically, the method for dividing the target area according to the geographical scope of the wildlife reserve, extracting the importance features of each target area and calculating the regional importance score, and generating the target area monitoring plan, and the UAV obtaining continuous frame images based on the target area monitoring plan is as follows:
[0013] Step S110: Divide the geographical scope of the wildlife reserve into a target areas and assign unique numbers, record the basic information and historical monitoring data of each target area, and establish a target area database. The basic information includes the food richness, concealment conditions, water source availability, and area of the target area. The historical monitoring data includes the occurrence frequency of target wild animals, the population density of target wild animals, the frequency of each behavior pattern of target wild animals being observed, the time of each observation of the activities of target wild animals, the nearest neighbor distance between target wild animals and adjacent wild animals, and the number of target wild animals;
[0014] Step S120: Extract the importance features of the target area from the target area database. The importance features include: the occurrence frequency of target wild animals 、the population density of target wild animals and the behavior pattern diversity index 、the regional habitat quality index 、the time distribution regularity index 、the spatial distribution regularity index ;
[0015] Step S130: Calculate the target area importance score based on the importance features ;
[0016] Step S140: Normalize the importance scores of the target regions, and map the importance scores of each target region to each monitoring parameter according to the mapping functions of the preset monitoring parameters; the monitoring parameters include: monitoring frequency MF, single monitoring duration MD, and number of drones DN.
[0017] Step S150: Generate a monitoring plan for the target regions according to the monitoring parameters of each target region, and the drones acquire consecutive frame images according to the monitoring plan for the target regions.
[0018] Specifically, the method for performing target detection using the improved YOLOv5 model, inputting consecutive frame images, outputting the bounding boxes of suspected target wild animals in each frame of the image and mapping them to the actual geographic coordinate system, and obtaining the actual position coordinates of the suspected target wild animals in each frame of the image is as follows:
[0019] Step S210: Perform wild animal target detection using the improved YOLOv5 model, and set the confidence threshold , take the consecutive frame images as the input, perform feature scaling on each input frame image to obtain a feature pyramid, input the feature pyramid into the improved YOLOv5 model, at each size, extract the target confidence prediction value, bounding box coordinate prediction value, and class probability prediction value corresponding to the prediction bounding box with a confidence greater than the confidence threshold, remove redundant prediction bounding boxes, and take the target confidence prediction values , bounding box coordinate prediction values and class probability prediction values of the remaining prediction bounding boxes and output them. Among them, represents the horizontal offset of the center of the bounding box relative to the upper left corner of the prediction bounding box, represents the vertical offset of the center of the bounding box relative to the upper left corner of the prediction bounding box, is the center point of the bounding box, is the width and height of the bounding box, represents the probability that the prediction bounding box belongs to the nth target wild animal category, where n is the total number of target wild animal categories, and the remaining prediction bounding boxes are used as the bounding boxes of suspected target wild animals;
[0020] The specific construction method of the improved YOLOv5 model is as follows:
[0021] Step S211: Construct an improved YOLOv5 model, use CSPDarknet53 as the backbone network, and add an attention mechanism to each stage of the backbone network. The attention mechanism consists of two fully connected layers and a Sigmoid activation function. Let the input feature map be X and the number of channels be C.
[0022] Step S212: The calculation process of the attention mechanism is as follows: perform global average pooling on each channel of X to obtain a feature vector Z with a size of 1×1×C; pass Z through the first fully connected layer FC1 to obtain a feature vector U with a size of 1×1× where r is a scaling factor used to control the number of parameters of the attention mechanism; pass U through the second fully connected layer FC2 to obtain a feature vector V with a size of 1×1×C; activate each element in V through the Sigmoid activation function to obtain an attention weight vector A; multiply the input feature map X by the attention weight vector A to obtain a modulated feature map ;
[0023] Step S213: Use the FPN structure to fuse feature maps of different sizes, and construct a feature pyramid through upsampling and concatenation operations;
[0024] Step S214: Set detection heads at each size of the feature pyramid. The detection head consists of a convolutional layer and a fully connected layer, and the detection head outputs the object confidence, bounding box coordinates, and class probabilities;
[0025] Step S215: Use the binary cross-entropy loss function to calculate the confidence loss between the predicted value and the actual value of the object confidence, use the GIOU loss function to calculate the bounding box loss between the predicted value and the actual value of the bounding box coordinates, use the multi-class cross-entropy loss function to calculate the class loss between the predicted value and the actual value of the class probabilities, and weight and sum the confidence loss, bounding box loss, and class loss to obtain the total loss function of the improved YOLOv5 model;
[0026] Step S216: Obtain wildlife images at historical times, annotate the class labels and bounding box coordinates of the target wildlife in each wildlife image, perform data augmentation on the wildlife images to obtain training images, use the annotated training images as input data, based on the improved YOLOv5 model, output the corresponding predicted values of object confidence, bounding box coordinates, and class probabilities, use minimizing the value of the total loss function as the training objective. When the total loss function converges, the training is completed, and the trained improved YOLOv5 model is obtained.
[0027] Step S220: According to the predicted class probabilities of the suspected target wildlife, select the target wildlife class corresponding to the largest predicted class probability as the preliminary recognition result;
[0028] Step S230: According to the size of the feature map calculate the target center coordinates on the feature map map the target center coordinates on the feature map to the continuous frame images, and calculate the normalized coordinates of the target center coordinates in the continuous frame images ;
[0029] Step S240: Based on the normalized coordinates and the current GPS coordinates of the drone , and the image resolution of the drone , map the bounding box of the suspected target wild animal to the actual geographic coordinate system to obtain the actual position coordinates of the suspected target wild animal in each frame of the continuous frame image .
[0030] Specifically, the specific method for performing key point detection on the t-th frame image, extracting the key point coordinates of the suspected target wild animal, calculating the visibility score of the key points and the confidence of each key point, and forming a feature descriptor with the K key points with the highest confidence is as follows:
[0031] Step S310: Based on the key point detection technology, extract the key point coordinates of the suspected target wild animal in the t-th frame image , and calculate the visibility score of the key points based on the conditional probability of key point visibility ;
[0032] Step S320: Set the visibility score threshold, and calculate the confidence of each key point based on the key point coordinates and the visibility score , when the visibility score is greater than or equal to the visibility score threshold, the confidence of this key point is equal to its visibility score, otherwise the confidence of this key point is equal to 0;
[0033] Step S330: Select the K key points with the highest confidence as the key point features of the suspected target wild animal to form the feature descriptor of the suspected target wild animal in the t-th frame image .
[0034] Specifically, the specific method for performing target detection and key point detection on the (t + 1)-th frame image to obtain the actual position coordinates and its feature descriptor of the candidate target, and performing trajectory matching between the suspected target wild animal and the candidate target. If they match, update the actual position coordinates and the feature descriptor of the suspected target wild animal. If they do not match, define the candidate target as a new target animal is as follows:
[0035] Step S410: Perform target detection on the (t + 1)-th frame image based on the improved YOLOv5 model, output the predicted value of the target confidence, the predicted value of the bounding box coordinates, and the predicted value of the class probability of the candidate target in the (t + 1)-th frame image, and calculate the actual position coordinates of the candidate target;
[0036] Step S420: Perform key point detection on the (t + 1)-th frame image to obtain the feature descriptor of the candidate target in the (t + 1)-th frame image ;
[0037] Step S430: Calculate the similarity between the candidate target and the suspected target wild animal based on the feature descriptors of the candidate target and the suspected target wild animal ;
[0038] Step S440: Preset a similarity threshold , if the similarity between the candidate target and the suspected target wild animal is greater than the similarity threshold, it is considered that the candidate target and the suspected target wild animal are successfully matched. Update the actual position coordinates of the suspected target wild animal to the actual position coordinates of the candidate target, and update the feature descriptor of the suspected target wild animal to the feature descriptor of the candidate target;
[0039] Step S450: If the similarity between the candidate target and the suspected target wild animal is less than or equal to the similarity threshold, it is considered that the suspected target wild animal is lost in the (t + 1)-th frame image, and define the candidate target in the (t + 1)-th frame image as a new target animal.
[0040] Specifically, the specific method for counting the identified suspected target wild animals and new target animals through a counter according to the trajectory matching result between the suspected target wild animal and the candidate target to obtain the total number of target wild animals in the target area is as follows:
[0041] Step S510: Set a counter count with an initial value of 0 to count the total number of target wild animals in the target area;
[0042] Step S520: Traverse all the suspected target wild animals that are successfully matched, and determine whether the suspected target wild animal has been counted. If the suspected target wild animal has not been counted, add 1 to the value of the counter count; otherwise, the value of the counter count remains unchanged;
[0043] Step S530: Traverse all the new target animals, and determine whether the new target animal has been counted. If the new target animal has not been counted, add 1 to the value of the counter count; otherwise, the value of the counter count remains unchanged;
[0044] Step S540: Take the final value of the counter count as the total number of target wild animals in the target area.
[0045] One aspect of the present application provides an intelligent recognition and perception system for UAV targets based on deep learning, including:
[0046] An image monitoring and acquisition module, configured to divide a target area according to the geographical scope of a wildlife reserve, extract important features of each target area and calculate a regional importance score, generate a target area monitoring plan, and the UAV obtains continuous frame images based on the target area monitoring plan;
[0047] An image target detection module, which is used to perform target detection using an improved YOLOv5 model, input consecutive frame images, output the bounding boxes of suspected target wild animals in each frame of image and map them to the actual geographic coordinate system, and obtain the actual position coordinates of suspected target wild animals in each frame of image;
[0048] An image key point detection module, which is used to perform key point detection on the t-th frame of image, extract the key point coordinates of suspected target wild animals, calculate the visibility score of the key points and the confidence of each key point, and the K key points with the highest confidence form a feature descriptor;
[0049] An animal trajectory matching module, which is used to perform target detection and key point detection on the (t + 1)-th frame of image to obtain the actual position coordinates and feature descriptors of candidate targets, perform trajectory matching on suspected target wild animals and candidate targets, if they match, update the actual position coordinates and feature descriptors of suspected target wild animals, if they do not match, define the candidate target as a new target animal;
[0050] An animal quantity statistics module, which is used to count the identified suspected target wild animals and new target animals through a counter according to the trajectory matching results of suspected target wild animals and candidate targets, and obtain the total number of target wild animals in the target area.
[0051] One aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it realizes the steps in the method for intelligent recognition and perception of drone targets based on deep learning.
[0052] One aspect of the present application provides a readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the method for intelligent recognition and perception of drone targets based on deep learning.
[0053] The system and method for intelligent recognition and perception of drone targets based on deep learning proposed in the present application have the following advantages compared with the prior art:
[0054] The present application proposes a number of novel feature indicators such as the behavior pattern diversity index, the regional habitat quality index, and the spatio-temporal distribution law index, designs a complex regional importance scoring formula, non-linearly synthesizes multiple indicators, comprehensively reflects the regional ecological importance, and establishes a mapping function between the score and the monitoring parameters to make the scoring result more operable.
[0055] By targeting the characteristics of wild animal targets, this application introduces an attention mechanism and multi-scale feature fusion to improve detection accuracy, establishes the mapping relationship between image coordinates and geographical coordinates, can calculate the actual position of the target in real time, realizes real-time positioning of the target, overcomes the limitation of the traditional method that requires prior modeling, and has better real-time performance and environmental adaptability.
[0056] By calculating the visibility scores of key points, this application quantifies the quality of key points, screens high-quality key points according to the confidence level to construct feature descriptors, obtains representative target feature descriptions, and improves the feature robustness of target matching and recognition.
[0057] By calculating the similarity based on the key point feature descriptors and performing trajectory matching, this application realizes cross-frame tracking of wild animal targets, effectively improving the performance of recognition and tracking.
[0058] By introducing a counter to avoid duplicate counting and based on the results of target detection, key point detection, and trajectory matching, this application gives a clear method for counting the number of targets, improving the accuracy of counting the number of wild animals in the target area. Description of the Drawings
[0059] Figure 1 is the method flow chart of the UAV target intelligent recognition and perception method based on deep learning provided by this application;
[0060] Figure 2 is the method flow chart of the target detection method provided by this application;
[0061] Figure 3 is the method flow chart of the key point detection method provided by this application;
[0062] Figure 4 is the functional module diagram of the UAV target intelligent recognition and perception system based on deep learning provided by this application. Detailed Embodiments
[0063] To better understand this application, more detailed descriptions of various aspects of this application will be made with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of the exemplary embodiments of this application and do not limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.
[0064] In the accompanying drawings, for the sake of clarity, the sizes, dimensions and shapes of the elements have been slightly adjusted. The accompanying drawings are only examples and are not drawn to an exact scale. As used herein, terms such as "substantially", "about" and similar terms are used as terms indicating approximation, rather than terms indicating degree, and are intended to account for the inherent deviations in measured or calculated values that would be recognized by a person of ordinary skill in the art. Additionally, in this application, the order in which the steps of each process are described does not necessarily represent the order in which these processes occur in actual operation, unless otherwise clearly specified or derivable from the context.
[0065] It should also be understood that expressions such as "comprising", "including", "having", "containing" and / or "including" are open-ended rather than closed-ended expressions in this specification, which mean that the stated features, elements and / or components exist, but do not exclude the existence of one or more other features, elements, components and / or their combinations. In addition, when an expression such as "at least one of..." appears after a list of listed features, it modifies the entire list of features, rather than just an individual element in the list. In addition, when describing the embodiments of this application, the use of "may" means "one or more embodiments of this application". And the term "exemplary" is intended to refer to an example or illustration.
[0066] Unless otherwise defined, all terms used herein (including engineering terms and scientific and technical terms) have the same meaning as commonly understood by a person of ordinary skill in the art to which this application pertains. It should also be understood that, unless clearly stated otherwise in this application, words defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense.
[0067] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will detail this application with reference to the accompanying drawings and in combination with the embodiments.
[0068] Example 1
[0069] As Figure 1 shown, the intelligent target recognition and perception method for drones based on deep learning provided by this application includes:
[0070] Step S100: Divide the target area according to the geographical scope of the wildlife reserve, extract the important features of each target area and calculate the regional importance score, generate a target area monitoring plan, and the drone obtains continuous frame images based on the target area monitoring plan;
[0071] The method for dividing the target area according to the geographical scope of the wildlife reserve, extracting the importance features of each target area, calculating the area importance score, and generating the target area monitoring plan, and the method for the drone to obtain continuous frame images based on the target area monitoring plan is as follows:
[0072] Step S110: Divide the geographical scope of the wildlife reserve into a target areas and assign a unique number to each. Record the basic information and historical monitoring data of each target area, and establish a target area database. The basic information includes the food richness, concealment conditions, water source availability, and target area area of the target area. The historical monitoring data includes the occurrence frequency of the target wild animal, the population density of the target wild animal, the frequency of each behavior pattern of the target wild animal being observed, the time of each observation of the target wild animal's activities, the nearest neighbor distance between the target wild animal and adjacent wild animals, and the number of target wild animals;
[0073] Step S120: Extract the importance features of the target area from the target area database. The importance features include: the occurrence frequency of the target wild animal , the population density of the target wild animal and the behavior pattern diversity index , the regional habitat quality index , the time distribution law index , and the spatial distribution law index ;
[0074] The behavior pattern diversity index reflects the richness of the behavior patterns exhibited by the target wild animal in the target area;
[0075] The calculation formula of the behavior pattern diversity index is: , where is the frequency of the jth behavior pattern being observed;
[0076] The regional habitat quality index reflects the habitat factors such as the food richness, concealment conditions, and water source availability of the target area;
[0077] The calculation formula of the regional habitat quality index is: , where represents the food richness, represents the comfort level of the concealment conditions, represents the water source availability, respectively represent the weight coefficients of the food richness, concealment condition comfort level, and water source availability;
[0078] The weight coefficients of food abundance, comfort level of hiding conditions, and water source availability are set by those skilled in the art according to experience. Among them, food abundance, comfort level of hiding conditions, and water source availability are professionally evaluated and determined by those skilled in the art based on the food distribution, hiding conditions, and water source distribution in the target area;
[0079] The time distribution regularity index reflects the time regularity and stability of the activities of the target wild animals;
[0080] The calculation formula of the time distribution regularity index is: , where NI is the number of drone observations, ni is the serial number of the drone observation times, and ni ∈ {1, 2,..., NI}, is the time angle when the target wild animals are observed for the ni-th time;
[0081] The time angle when the target wild animals are observed is represented by 0 to 360 degrees for 24 hours a day;
[0082] The spatial distribution regularity index reflects the spatial distribution pattern and aggregation degree of the target wild animals in the area. The value range of NNI is (0, 2.15). A value close to 1 indicates a random distribution, a value less than 1 indicates an aggregated distribution, and a value greater than 1 indicates a uniform distribution;
[0083] The calculation formula of the spatial distribution regularity index is: , where, is the nearest neighbor distance between the na-th target wild animal and the adjacent wild animals, refers to the arithmetic mean of the nearest neighbor distances, refers to the area of the target area, refers to the number of target wild animals;
[0084] Step S130: Calculate the importance score of the target area based on the important features ;
[0085] The calculation formula of the importance score of the target area is: , where, , , , , , are respectively the occurrence frequency of the target wild animals, the population density of the target wild animals, and the behavior pattern diversity index , the regional habitat quality index , the time distribution regularity index , and the spatial distribution regularity index Weight coefficient;
[0086] The occurrence frequency of the target wild animal , the population density of the target wild animal and the behavior pattern diversity index , the regional habitat quality index , the time distribution law index , the space distribution law index The weight coefficients are set by those skilled in the art according to experience.
[0087] Step S140: Normalize the importance scores of the target areas, and map the importance scores of each target area to each monitoring parameter according to the preset mapping functions of each monitoring parameter; the monitoring parameters include: monitoring frequency MF, single monitoring duration MD, and number of drones DN;
[0088] The mapping function of the monitoring frequency is: , where is the maximum value of the monitoring frequency, represents rounding to the nearest integer;
[0089] Preferably, is 30 times per month;
[0090] The mapping function of the single monitoring duration is: , where is the maximum value of the single monitoring duration;
[0091] Preferably, is 2 hours;
[0092] The mapping function of the number of drones is: , where is the maximum value of the number of drones, represents rounding up;
[0093] Step S150: Generate a monitoring plan for the target areas according to the monitoring parameters of each target area, and the drones obtain consecutive frame images according to the monitoring plan for the target areas;
[0094] The above steps design a complex regional importance scoring formula, comprehensively considering the non-linear contributions and interactive effects of multiple characteristic indicators. This scoring method can more comprehensively and accurately reflect the ecological importance of each target area, providing more targeted guidance for drone monitoring.
[0095] Step S200: Perform object detection using an improved YOLOv5 model. Input consecutive frame images, output the bounding boxes of suspected target wild animals in each frame image and map them to the actual geographic coordinate system to obtain the actual position coordinates of the suspected target wild animals in each frame image;
[0096] The specific method of performing object detection using an improved YOLOv5 model, inputting consecutive frame images, outputting the bounding boxes of suspected target wild animals in each frame image and mapping them to the actual geographic coordinate system to obtain the actual position coordinates of the suspected target wild animals in each frame image is as follows:
[0097] Step S210: Perform wild animal object detection using an improved YOLOv5 model and set a confidence threshold , take consecutive frame images as input, perform feature scaling on each input frame image to obtain a feature pyramid, input the feature pyramid into the improved YOLOv5 model. At each size, extract the target confidence prediction value, bounding box coordinate prediction value, and class probability prediction value corresponding to the prediction bounding boxes with a confidence greater than the confidence threshold, remove redundant prediction bounding boxes, and output the target confidence prediction value , bounding box coordinate prediction value and class probability prediction value of the remaining prediction bounding boxes. Among them, represents the horizontal offset of the bounding box center relative to the upper left corner of the prediction bounding box, represents the vertical offset of the bounding box center relative to the upper left corner of the prediction bounding box, is the center point of the bounding box, is the width and height of the bounding box, represents the probability that the prediction bounding box belongs to the nth target wild animal category, where n is the total number of target wild animal categories. The remaining prediction bounding boxes are used as the bounding boxes of suspected target wild animals;
[0098] The bounding box coordinate prediction value includes the center point of the bounding box and the width and height of the bounding box, with the cell of each prediction bounding box as the coordinate system;
[0099] The class probability prediction value refers to the probability of belonging to each target wild animal category;
[0100] The value of the confidence threshold is set by those skilled in the art according to experience.
[0101] The specific construction method of the improved YOLOv5 model is as follows:
[0102] Step S211: Construct an improved YOLOv5 model. Use CSPDarknet53 as the backbone network and add an attention mechanism to each stage of the backbone network. The attention mechanism consists of two fully connected layers and a Sigmoid activation function. Let the input feature map be X with the number of channels C.
[0103] Step S212: The calculation process of the attention mechanism is as follows: Perform global average pooling on each channel of X to obtain a feature vector Z with a size of 1×1×C; Pass Z through the first fully connected layer FC1 to obtain a feature vector U with a size of 1×1× , where r is a scaling factor used to control the number of parameters of the attention mechanism; Pass U through the second fully connected layer FC2 to obtain a feature vector V with a size of 1×1×C; Activate each element in V through the Sigmoid activation function to obtain an attention weight vector A; Multiply the input feature map X by the attention weight vector A to obtain a modulated feature map ;
[0104] The calculation formula for the calculation process of the attention mechanism is: , , , , , where represents global average pooling, represents passing Z through the first fully connected layer FC1, represents passing U through the second fully connected layer, , are the weight matrices of the first fully connected layer and the second fully connected layer respectively, , are the bias terms of the first fully connected layer and the second fully connected layer respectively, is the Sigmoid activation function, is channel-wise multiplication;
[0105] Step S213: Use the FPN structure to fuse feature maps of different sizes and construct a feature pyramid through upsampling and concatenation operations;
[0106] Step S214: Set detection heads at each size of the feature pyramid. The detection head consists of a convolutional layer and a fully connected layer. The detection head outputs object confidence, bounding box coordinates, and class probabilities;
[0107] Step S215: Calculate the confidence loss between the predicted value of the target confidence and the actual value of the target confidence using the binary cross-entropy loss function, calculate the bounding box loss between the predicted value of the bounding box coordinates and the actual value of the bounding box coordinates using the GIOU loss function, calculate the class loss between the predicted value of the class probability and the actual value of the class probability using the multi-class cross-entropy loss function, and sum the confidence loss, bounding box loss, and class loss with weights to obtain the total loss function of the improved YOLOv5 model;
[0108] Step S216: Obtain wildlife images at historical times, annotate the class labels and bounding box coordinates of the target wildlife in each wildlife image, perform data augmentation on the wildlife images to obtain training images, use the annotated training images as input data, and based on the improved YOLOv5 model, output the corresponding predicted values of the target confidence, predicted values of the bounding box coordinates, and predicted values of the class probability. Take minimizing the value of the total loss function as the training objective. When the total loss function converges, the training is completed, and the trained improved YOLOv5 model is obtained.
[0109] Step S220: According to the predicted value of the class probability of the suspected target wildlife, select the target wildlife class corresponding to the maximum predicted value of the class probability as the preliminary recognition result;
[0110] Step S230: According to the size of the feature map , calculate the target center coordinates on the feature map , map the target center coordinates on the feature map to the consecutive frame images, and calculate the normalized coordinates of the target center coordinates in the consecutive frame images ;
[0111] The calculation formula for the target center coordinates on the feature map is: , ;
[0112] The mapping formula for mapping the target center coordinates on the feature map to the consecutive frame images is: , , where is the target center coordinates mapped to the consecutive frame images, is the width of the consecutive frame images, is the height of the consecutive frame images;
[0113] The calculation formula for the normalized coordinates of the target center coordinates in the consecutive frame images is: , ;
[0114] The target center coordinates are based on the global coordinate system of the feature map;
[0115] Step S240: Based on the normalized coordinates and the current GPS coordinates of the drone , the drone image resolution , map the bounding box of the suspected target wild animal to the actual geographic coordinate system to obtain the actual position coordinates of the suspected target wild animal in each frame of the continuous frame image ;
[0116] The mapping formula for the actual position coordinates of the suspected target wild animal is: , ;
[0117] The above steps optimize and improve the YOLOv5 algorithm according to the characteristics of wild animal targets, introduce the attention mechanism and the multi-scale feature fusion strategy, and improve the detection accuracy of wild animals. The above steps realize the real-time calculation of the target position by establishing the mapping relationship between the image coordinate system and the geographic coordinate system, overcome the limitation of the traditional method that requires a complex environment model to be established in advance, and have better real-time performance and adaptability.
[0118] Step S300: Perform key point detection on the t-th frame image, extract the key point coordinates of the suspected target wild animal, calculate the visibility score of the key points and the confidence of each key point, and the K key points with the highest confidence form a feature descriptor;
[0119] The specific method for performing key point detection on the t-th frame image, extracting the key point coordinates of the suspected target wild animal, calculating the visibility score of the key points and the confidence of each key point, and the K key points with the highest confidence form a feature descriptor is as follows:
[0120] Step S310: Based on the key point detection technology, extract the key point coordinates of the suspected target wild animal in the t-th frame image , and calculate the visibility score of the key points based on the conditional probability of key point visibility ;
[0121] The calculation formula for the visibility score of the key points is: , where is the body part label to which the i-th key point belongs, is the category of the suspected target wild animal;
[0122] Step S320: Set the visibility score threshold, and calculate the confidence of each key point based on the key point coordinates and the visibility score , when the visibility score is greater than or equal to the visibility score threshold, the confidence of this key point is equal to its visibility score, otherwise the confidence of this key point is equal to 0;
[0123] The calculation formula for the confidence of the key points is: , where is the visibility score threshold;
[0124] The visibility score threshold is used to filter low-quality key points and is set by those skilled in the art according to experience.
[0125] Step S330: Select the K key points with the highest confidence as the key point features of the suspected target wild animal, and form the feature descriptor of the suspected target wild animal in the t-th frame image ;
[0126] The feature descriptor is expressed as: , where is the coordinate of the K-th key point, is the confidence of the K-th key point;
[0127] Figure 2 is the flow chart of the object detection method provided by this application;
[0128] Through the above steps, the key point coordinates and visibility scores of the target wild animal are obtained through key point detection, and the feature descriptor of the key point is constructed according to the confidence. This feature descriptor can be used for subsequent object matching and recognition tasks.
[0129] Step S400: Perform object detection and key point detection on the (t + 1)-th frame image to obtain the actual position coordinates and feature descriptor of the candidate target, perform trajectory matching on the suspected target wild animal and the candidate target. If they match, update the actual position coordinates and feature descriptor of the suspected target wild animal. If they do not match, define the candidate target as a new target animal;
[0130] The specific method of performing object detection and key point detection on the (t + 1)-th frame image to obtain the actual position coordinates and feature descriptor of the candidate target, performing trajectory matching on the suspected target wild animal and the candidate target. If they match, update the actual position coordinates and feature descriptor of the suspected target wild animal. If they do not match, define the candidate target as a new target animal is as follows:
[0131] Step S410: Perform object detection on the (t + 1)-th frame image based on the improved YOLOv5 model, output the object confidence prediction value, bounding box coordinate prediction value and class probability prediction value of the candidate target in the (t + 1)-th frame image, and calculate the actual position coordinates of the candidate target;
[0132] The calculation method of the actual position coordinates of the candidate target is the same as that of the actual position coordinates of the suspected target wild animal;
[0133] Step S420: Perform key point detection on the (t + 1)-th frame image to obtain the feature descriptor of the candidate target in the (t + 1)-th frame image ;
[0134] The calculation method of the feature descriptor of the candidate target is the same as that of the suspected target wild animal;
[0135] Step S430: Calculate the similarity between the candidate target and the suspected target wild animal based on the feature descriptor of the candidate target and the feature descriptor of the suspected target wild animal ;
[0136] The calculation formula for the similarity between the candidate target and the suspected target wild animal is: , where and are the feature descriptors of the k-th key point of the suspected target wild animal f and the candidate target g respectively, represents the L2 norm, represents the exponential function with base e;
[0137] Step S440: Preset a similarity threshold , if the similarity between the candidate target and the suspected target wild animal is greater than the similarity threshold, it is considered that the candidate target and the suspected target wild animal are successfully matched, update the actual position coordinates of the suspected target wild animal to the actual position coordinates of the candidate target, and update the feature descriptor of the suspected target wild animal to the feature descriptor of the candidate target;
[0138] The similarity threshold is set by those skilled in the art according to experience.
[0139] Step S450: If the similarity between the candidate target and the suspected target wild animal is less than or equal to the similarity threshold, it is considered that the suspected target wild animal is lost in the (t + 1)-th frame image, and the candidate target in the (t + 1)-th frame image is defined as a new target animal;
[0140] Figure 3 is the flow chart of the key point detection method provided by this application;
[0141] The above steps realize the cross-frame tracking of wild animals through key point detection and matching, effectively improving the performance of automatic identification and tracking of wild animals, and providing a basis for the subsequent step of counting the total number of target wild animals in the target area.
[0142] Step S500: According to the trajectory matching result of the suspected target wild animal and the candidate target, count the identified suspected target wild animals and new target animals through a counter to obtain the total number of target wild animals in the target area;
[0143] According to the trajectory matching results of the suspected target wild animals and the candidate targets, the specific method for counting the identified suspected target wild animals and new target animals through a counter to obtain the total number of target wild animals in the target area is as follows:
[0144] Step S510: Set a counter count with an initial value of 0 to count the total number of target wild animals in the target area;
[0145] Step S520: Traverse all the successfully matched suspected target wild animals, and determine whether the suspected target wild animal has been counted. If the suspected target wild animal has not been counted, add 1 to the value of the counter count; otherwise, the value of the counter count remains unchanged;
[0146] Step S530: Traverse all the new target animals, and determine whether the new target animal has been counted. If the new target animal has not been counted, add 1 to the value of the counter count; otherwise, the value of the counter count remains unchanged;
[0147] When the value of the counter count is incremented by 1, the flag bit of the suspected target wild animal or the new target animal is set to true, indicating that it has been counted;
[0148] Step S540: Take the final value of the counter count as the total number of target wild animals in the target area.
[0149] When the technical solution provided by this application is actually applied, by setting multiple drones in the target area to simultaneously obtain consecutive frame images, and then performing target detection, key point detection, and trajectory matching on these consecutive frame images, it is ensured that there is no duplicate counting, and the accuracy of counting target wild animals is maximally ensured.
[0150] The above steps are based on the target detection, key point detection, and trajectory matching in steps S200~S400. By setting a counter, the problem of calculating the number of target wild animals in the total target area is clearly solved.
[0151] Embodiment 2
[0152] As Figure 4 shown, the drone target intelligent recognition and perception system based on deep learning provided by this application includes:
[0153] An image monitoring and acquisition module, which is used to divide the target area according to the geographical scope of the wildlife reserve, extract the important feature of each target area and calculate the regional importance score, generate a target area monitoring plan, and the drone obtains consecutive frame images based on the target area monitoring plan;
[0154] An image target detection module, which is used to perform target detection using an improved YOLOv5 model, input consecutive frame images, output the bounding boxes of suspected target wild animals in each frame image and map them to the actual geographic coordinate system, and obtain the actual position coordinates of the suspected target wild animals in each frame image;
[0155] An image key point detection module, which is used to perform key point detection on the t-th frame image, extract the key point coordinates of the suspected target wild animals, calculate the visibility scores of the key points and the confidence of each key point, and the K key points with the highest confidence form a feature descriptor;
[0156] An animal trajectory matching module, which is used to perform target detection and key point detection on the (t + 1)-th frame image to obtain the actual position coordinates and feature descriptors of the candidate targets, perform trajectory matching on the suspected target wild animals and the candidate targets. If they match, update the actual position coordinates and feature descriptors of the suspected target wild animals. If they do not match, define the candidate target as a new target animal;
[0157] An animal quantity statistics module, which is used to count the identified suspected target wild animals and new target animals through a counter according to the trajectory matching results of the suspected target wild animals and the candidate targets, and obtain the total number of target wild animals in the target area.
[0158] Embodiment 3
[0159] An embodiment of the present application provides an electronic device for an intelligent recognition and perception method of drone targets based on deep learning. The electronic device may include one or more processors and one or more memories. Among them, computer-readable code is stored in the memory, and when the computer-readable code is run by one or more processors, it can execute the intelligent recognition and perception method of drone targets based on deep learning as described above.
[0160] The electronic device may include a bus, one or more CPUs, a read-only memory (ROM), a random access memory (RAM), a communication port connected to a network, input / output components, a hard disk, etc. The storage device in the electronic device, such as ROM or hard disk, may store the deep learning-based intelligent recognition and perception method for drone targets provided in this application. The deep learning-based intelligent recognition and perception method for drone targets may, for example, include: dividing target areas according to the geographical scope of a wildlife reserve, extracting the important features of each target area and calculating the regional importance score, generating a monitoring plan for the target areas, and the drone obtaining consecutive frame images based on the monitoring plan for the target areas; performing target detection using an improved YOLOv5 model, inputting the consecutive frame images, outputting the bounding boxes of suspected target wild animals in each frame image and mapping them to the actual geographical coordinate system to obtain the actual position coordinates of the suspected target wild animals in each frame image; performing key point detection on the t-th frame image, extracting the key point coordinates of the suspected target wild animals, calculating the visibility score of the key points and the confidence level of each key point, and the K key points with the highest confidence levels form a feature descriptor; performing target detection and key point detection on the (t + 1)-th frame image to obtain the actual position coordinates and its feature descriptor of the candidate target, performing trajectory matching between the suspected target wild animals and the candidate target, if they match, updating the actual position coordinates and the feature descriptor of the suspected target wild animals, if they do not match, defining the candidate target as a new target animal; according to the trajectory matching results of the suspected target wild animals and the candidate target, counting the recognized suspected target wild animals and new target animals through a counter to obtain the total number of target wild animals in the target area. Further, the electronic device may also include a user interface. Of course, the architecture of the above electronic device is only exemplary, and when implementing different devices, one or more components in the above electronic device may be omitted according to actual needs.
[0161] Embodiment 4
[0162] An embodiment of this application provides a computer-readable storage medium for the deep learning-based intelligent recognition and perception method for drone targets. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are run by a processor, the deep learning-based intelligent recognition and perception method according to the embodiments of this application can be executed. The storage medium includes but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may, for example, include random access memory (RAM) and cache memory, etc. Non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.
[0163] In addition, according to an embodiment of the present application, the processes described above can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be run by a processor to execute instructions corresponding to the method steps provided by the present application. For example: dividing a target area according to the geographical range of a wildlife reserve, extracting the importance features of each target area and calculating the area importance score, generating a target area monitoring plan, and a drone obtaining consecutive frame images based on the target area monitoring plan; performing object detection using an improved YOLOv5 model, inputting the consecutive frame images, outputting the bounding boxes of suspected target wild animals in each frame of the image and mapping them to the actual geographic coordinate system to obtain the actual position coordinates of the suspected target wild animals in each frame of the image; performing key point detection on the t-th frame of the image, extracting the key point coordinates of the suspected target wild animals, calculating the visibility score of the key points and the confidence of each key point, and the K key points with the highest confidence form a feature descriptor; performing object detection and key point detection on the (t + 1)-th frame of the image to obtain the actual position coordinates and its feature descriptor of the candidate target, performing trajectory matching on the suspected target wild animals and the candidate target, if they match, updating the actual position coordinates and the feature descriptor of the suspected target wild animals, if they do not match, defining the candidate target as a new target animal; according to the trajectory matching results of the suspected target wild animals and the candidate target, counting the identified suspected target wild animals and new target animals through a counter to obtain the total number of target wild animals in the target area. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.
[0164] The methods, apparatuses, and devices of the present application can be implemented in many ways. For example, the methods, apparatuses, and devices of the present application can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the method is only for illustration, and the steps of the method of the present application are not limited to the specific order described above unless otherwise specifically stated. In addition, in some embodiments, the present application can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers a recording medium storing a program for executing the method according to the present application.
[0165] In addition, parts of the above technical solutions provided in the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.
[0166] The specific embodiments described above further elaborate in detail the objectives, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for intelligent recognition and perception of UAV targets based on deep learning, characterized in that: include: Divide the target area according to the geographical scope of the wildlife reserve, extract the importance features of each target area and calculate the regional importance score, generate a target area monitoring plan, and the drone acquires continuous frame images based on the target area monitoring plan; The improved YOLOv5 model is used for target detection. Continuous frame images are input, and the bounding box of the suspected target wild animal in each frame is output and mapped to the actual geographic coordinate system to obtain the actual position coordinates of the suspected target wild animal in each frame. Perform key point detection on the t-th frame image, extract the key point coordinates of the suspected target wild animal, calculate the visibility score of the key point and the confidence of each key point, and the K key points with the highest confidence constitute the feature descriptor; Perform target detection and key point detection on the t+1th frame image to obtain the actual position coordinates and feature descriptors of the candidate target, and match the trajectory of the suspected target wild animal with the candidate target. If they match, update the actual position coordinates and feature descriptors of the suspected target wild animal. If they do not match, define the candidate target as a new target animal. According to the track matching results of the suspected target wild animals and the candidate targets, the identified suspected target wild animals and new target animals are counted by a counter to obtain the total number of target wild animals in the target area; The specific method of performing key point detection on the t-th frame image, extracting the key point coordinates of the suspected target wild animal, calculating the visibility score of the key point and the confidence of each key point, and forming the feature descriptor with the K key points with the highest confidence is as follows: Based on key point detection technology, the key point coordinates of the suspected target wild animal in the t-th frame image are extracted , the visibility score of the key point is calculated based on the conditional probability of the key point visibility ; Set the visibility score threshold and calculate the confidence of each key point based on the key point coordinates and visibility score , when the visibility score is greater than or equal to the visibility score threshold, the confidence of the key point is equal to its visibility score, otherwise the confidence of the key point is equal to 0; Select the K key points with the highest confidence as the key point features of the suspected target wild animal, and form the feature descriptor of the suspected target wild animal in the t-th frame image. ; The specific method of performing target detection and key point detection on the t+1th frame image to obtain the actual position coordinates and feature descriptors of the candidate target, performing trajectory matching on the suspected target wild animal and the candidate target, and updating the actual position coordinates and feature descriptors of the suspected target wild animal if they match, and defining the candidate target as a new target animal if they do not match includes: Based on the feature descriptors of the candidate targets and the feature descriptors of the suspected target wild animals, the similarity between the candidate targets and the suspected target wild animals is calculated. ; Preset similarity threshold If the similarity between the candidate target and the suspected target wild animal is greater than the similarity threshold, the candidate target and the suspected target wild animal are considered to be matched successfully, and the actual position coordinates of the suspected target wild animal are updated to the actual position coordinates of the candidate target, and the feature descriptor of the suspected target wild animal is updated to the feature descriptor of the candidate target; If the similarity between the candidate target and the suspected target wild animal is less than or equal to the similarity threshold, the suspected target wild animal is considered to be lost in the t+1th frame image, and the candidate target in the t+1th frame image is defined as a new target animal.
2. The method for intelligent recognition and perception of unmanned aerial vehicle targets based on deep learning according to claim 1, characterized in that: The target area is divided according to the geographical scope of the wildlife reserve, the importance features of each target area are extracted and the regional importance score is calculated, and a target area monitoring plan is generated. The specific method for the drone to obtain continuous frame images based on the target area monitoring plan is as follows: Divide the geographical scope of the wildlife reserve into a target areas and assign unique numbers, record the basic information and historical monitoring data of each target area, and establish a target area database. The basic information includes the food abundance, shelter conditions and water availability of the target area, and the area of the target area. The historical monitoring data includes the frequency of occurrence of target wildlife, the population density of target wildlife, the frequency of each behavior pattern of target wildlife being observed, the time of each observation of target wildlife activity, the nearest neighbor distance between the target wildlife and its adjacent wildlife, and the number of target wildlife; Extract the importance features of the target area from the target area database, the importance features include: the frequency of occurrence of target wild animals , Population density of target wildlife and behavioral pattern diversity index , Regional Habitat Quality Index , Time Distribution Index , spatial distribution law index ; Calculate the importance score of the target area based on the importance features ; The importance scores of the target areas are normalized, and the importance scores of the target areas are mapped to various monitoring parameters according to the preset mapping functions of various monitoring parameters; the monitoring parameters include: monitoring frequency MF, single monitoring duration MD, and number of drones DN; A target area monitoring plan is generated according to the monitoring parameters of each target area, and the UAV obtains continuous frame images according to the target area monitoring plan.
3. The method for intelligent recognition and perception of unmanned aerial vehicle targets based on deep learning according to claim 2, characterized in that: The improved YOLOv5 model is used for target detection, continuous frame images are input, the bounding box of the suspected target wild animal in each frame image is output and mapped to the actual geographic coordinate system, and the specific method for obtaining the actual position coordinates of the suspected target wild animal in each frame image is: Use the improved YOLOv5 model for wildlife target detection and set the confidence threshold , take continuous frame images as input, perform feature scaling on each input frame image to obtain a feature pyramid, input the feature pyramid into the improved YOLOv5 model, extract the target confidence prediction value, bounding box coordinate prediction value and category probability prediction value corresponding to the prediction bounding box with a confidence greater than the confidence threshold at each size, remove redundant prediction bounding boxes, and convert the target confidence prediction value of the remaining prediction bounding boxes into , bounding box coordinate prediction value and class probability prediction value Output, where Indicates the horizontal offset of the center of the bounding box relative to the upper left corner of the predicted bounding box. Indicates the vertical offset of the center of the bounding box relative to the upper left corner of the predicted bounding box. is the center point of the bounding box, are the bounding box width and the bounding box height, represents the probability that the predicted bounding box belongs to the nth target wildlife category, where n is the total number of target wildlife categories, and the remaining predicted bounding boxes are used as bounding boxes of suspected target wildlife; According to the category probability prediction value of the suspected target wild animal, the target wild animal category corresponding to the largest category probability prediction value is selected as the preliminary identification result; According to the size of the feature map , calculate the target center coordinates on the feature map , map the target center coordinates on the feature map to the continuous frame images, and calculate the normalized coordinates of the target center coordinates in the continuous frame images ; Based on the normalized coordinates and the current GPS coordinates of the drone , UAV image resolution , map the bounding box of the suspected target wild animal to the actual geographic coordinate system, and obtain the actual position coordinates of the suspected target wild animal in each frame of the continuous frame image .
4. The method for intelligent recognition and perception of unmanned aerial vehicle targets based on deep learning as claimed in claim 3, characterized in that: The specific construction method of the improved YOLOv5 model is: An improved YOLOv5 model is constructed, using CSPDarknet53 as the backbone network. An attention mechanism is added to each stage of the backbone network. The attention mechanism consists of two fully connected layers and a Sigmoid activation function. The input feature map is X and the number of channels is C. The calculation process of the attention mechanism is as follows: perform global average pooling on each channel of X to obtain a feature vector Z with a size of 1×1×C; pass Z through the first fully connected layer FC1 to obtain a feature vector with a size of 1×1× The feature vector U of , where r is the scaling factor, is used to control the parameter amount of the attention mechanism; U is passed through the second fully connected layer FC2 to obtain a feature vector V of size 1×1×C; each element in V is activated by the Sigmoid activation function to obtain the attention weight vector A; the input feature map X is multiplied by the attention weight vector A to obtain the modulated feature map ; The FPN structure is used to fuse feature maps of different sizes, and a feature pyramid is constructed through upsampling and splicing operations; A detection head is set at each dimension of the feature pyramid, wherein the detection head consists of a convolutional layer and a fully connected layer, and the detection head outputs target confidence, bounding box coordinates, and category probability; The confidence loss between the target confidence prediction value and the actual target confidence value is calculated using the binary cross entropy loss function, the bounding box loss between the bounding box coordinate prediction value and the actual bounding box coordinate value is calculated using the GIOU loss function, and the category loss between the category probability prediction value and the actual category probability value is calculated using the multivariate cross entropy loss function. The confidence loss, bounding box loss, and category loss are weighted and summed to obtain the total loss function of the improved YOLOv5 model. Obtain wildlife images from historical times, annotate each wildlife image with the category label and bounding box coordinates of the target wildlife, perform data enhancement on the wildlife images to obtain training images, use the annotated training images as input data, output the corresponding target confidence prediction value, bounding box coordinate prediction value, and category probability prediction value based on the improved YOLOv5 model, and use the value that minimizes the total loss function as the training goal. When the total loss function converges, the training is completed, and a trained improved YOLOv5 model is obtained.
5. The method for intelligent recognition and perception of unmanned aerial vehicle targets based on deep learning according to claim 4, characterized in that: The specific method of performing target detection and key point detection on the t+1th frame image to obtain the actual position coordinates and feature descriptors of the candidate target, performing trajectory matching on the suspected target wild animal and the candidate target, and updating the actual position coordinates and feature descriptors of the suspected target wild animal if they match, and defining the candidate target as a new target animal if they do not match also includes: Perform target detection on the t+1th frame image based on the improved YOLOv5 model, output the target confidence prediction value, bounding box coordinate prediction value and category probability prediction value of the candidate target in the t+1th frame image, and calculate the actual position coordinates of the candidate target; Perform key point detection on the t+1th frame image to obtain the feature descriptor of the candidate target in the t+1th frame image .
6. The method for intelligent recognition and perception of unmanned aerial vehicle targets based on deep learning according to claim 5, characterized in that: The specific method of counting the identified suspected target wild animals and new target animals through a counter according to the track matching results of the suspected target wild animals and the candidate targets to obtain the total number of target wild animals in the target area is: Set the counter count, with an initial value of 0, to count the total number of target wild animals in the target area; Traverse all the suspected target wild animals that have been matched successfully, and determine whether the suspected target wild animals have been counted. If the suspected target wild animals have not been counted, add 1 to the value of the counter count; otherwise, the value of the counter count remains unchanged; Traverse all new target animals and determine whether the new target animal has been counted. If the new target animal has not been counted, add 1 to the value of the counter count; Otherwise, the value of the counter count remains unchanged; The final value of the counter count is taken as the total number of target wild animals in the target area.
7. A UAV target intelligent recognition and perception system based on deep learning, which is used to implement the UAV target intelligent recognition and perception method based on deep learning as described in any one of claims 1 to 6, characterized in that: include: The image monitoring acquisition module is used to divide the target area according to the geographical scope of the wildlife reserve, extract the importance features of each target area and calculate the regional importance score, generate a target area monitoring plan, and the drone acquires continuous frame images based on the target area monitoring plan; The image target detection module is used to perform target detection using the improved YOLOv5 model, input continuous frame images, output the bounding box of the suspected target wild animal in each frame image and map it to the actual geographic coordinate system to obtain the actual position coordinates of the suspected target wild animal in each frame image; The image key point detection module is used to detect key points of the t-th frame image, extract the key point coordinates of the suspected target wild animal, calculate the visibility score of the key point and the confidence of each key point, and the K key points with the highest confidence constitute the feature descriptor; The animal track matching module is used to perform target detection and key point detection on the t+1 frame image to obtain the actual position coordinates and feature descriptors of the candidate target, and to match the track of the suspected target wild animal with the candidate target. If they match, the actual position coordinates and feature descriptors of the suspected target wild animal are updated; if they do not match, the candidate target is defined as a new target animal; The animal number counting module is used to count the suspected target wild animals and new target animals identified through a counter based on the trajectory matching results of the suspected target wild animals and the candidate targets, so as to obtain the total number of target wild animals in the target area.
8. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for intelligent recognition and perception of unmanned aerial vehicle targets based on deep learning as described in any one of claims 1 to 6 are implemented.
9. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which is suitable for being loaded by a processor to execute the steps in the deep learning-based drone target intelligent recognition and perception method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
An intelligent identification method for small sample and small target Tibetan antelope based on UAV remote sensing
CN116310894B
Method and system for identifying fine impurities in wine bottle
CN119379592A
Occlusion-aware multi-object tracking
US20220391621A1