A method, device, equipment and medium for collaborative target tracking of UAV formation based on monocular vision

Through the pilot-following algorithm of monocular vision and improved edge detection, template matching, and kernel-related filtering algorithm, high-precision target tracking and formation adjustment of drone formation in complex environments is achieved, and the problems of large computing resource consumption and inconvenient formation adjustment in the prior art are solved.

CN119200651BActive Publication Date: 2025-07-04NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411707765.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-07-04
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

The existing drone formations are difficult to track targets in real time and with high accuracy in complex environments, and the computing resources are consumed, making formation adjustments inconvenient.

Method used

The pilot-following algorithm based on monocular vision is adopted, combined with improved adaptive edge detection, template matching and kernel-related filtering algorithms, and the target is identified and tracked through the pilot drone, and the formation is adjusted through the position solution algorithm.

Benefits of technology

Improve the target tracking accuracy of drone formations in complex environments, save computing resources, and flexibly adjust the formation to complete tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119200651B_ABST
    Figure CN119200651B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and medium for collaborative target tracking of UAV formation based on monocular vision, which relates to the technical fields of unmanned control and visual target recognition and tracking. It detects and identifies static or dynamic targets and marks the target positions, processes the targets using the kernel correlation filtering algorithm, realizes the tracking of the target by the leader UAV and improves the tracking accuracy. The formation members share the formation information and the position information of the leader UAV, and calculate the expected positions of each following UAV at the next moment through the position calculation algorithm, so as to form a formation to achieve collaborative target tracking. After the static or dynamic target is recognized by the monocular vision sensor, the UAV formation realizes collaborative tracking of the target while maintaining a predetermined formation or changing the formation. The present invention combines an improved formation control method with a target recognition and tracking method based on monocular vision to improve the system performance and ensure that the formation completes the task safely and efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of multi-agent formation control and visual target recognition and tracking, and particularly to a method, device, equipment and medium for collaborative target tracking of an unmanned aerial vehicle (UAV) formation based on monocular vision. Background Art

[0002] With the development of sensor and electronic technologies, UAV formations have been widely used in military, civilian and other fields. The target recognition and tracking of UAVs refers to the technology of real-time detecting and recognizing ground or aerial targets through on-board vision, radar and other sensors, and real-time tracking the target position and motion trajectory through the tracking of continuous image frames. This technology involves computer vision, machine learning and the aviation field, and presents significant application prospects in today's security and monitoring, border patrol, search and rescue and other fields.

[0003] Currently, the technology of collaborative target tracking of UAV formations based on vision is one of the research hotspots in the UAV field. In the prior art, when UAVs identify targets through vision sensors, they will inevitably be affected by complex environmental interference and algorithm complexity. When performing real-time tracking of targets, how to improve the algorithm accuracy and save computing resources is also a major challenge. In addition, since UAV formations often need to maintain or adjust the formation shape according to actual needs when tracking targets, and the motion direction and speed of the targets are uncertain, how to calculate and control the positions and speeds of formation members in real time has also become a technical problem to be solved urgently in this field.

[0004] Based on the above, the present invention proposes a technical solution for collaborative target tracking of UAV formations based on monocular vision, which can maintain a predetermined formation shape or adjust the formation shape according to actual needs while realizing the tracking of targets by UAV formations based on monocular vision. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, device, equipment and medium for collaborative target tracking of UAV formations based on monocular vision,

[0006] and its specific solutions are as follows:

[0007] In a first aspect, according to the current initialization stage of the system, the leader and follower agents are selected, and the communication relationship between each agent is established. Among them, a UAV is selected as the leader agent, and a monocular camera is mounted on the UAV.

[0008] The leading agent obtains image information through a monocular camera, inputs the image information into the first step, then inputs the output result of the first step into the second step, and determines the target recognized by the system according to the output result of the second step; the first step is edge detection algorithm calculation, and the second step is template matching algorithm calculation;

[0009] The leading agent inputs the target position information into the third step, and decides the movement mode of the leading agent according to the output result of the third step; the third step is single-object tracking calculation;

[0010] Solve the relative positions between the follower agents and with the leading agent, and broadcast the movement decision of the leading agent to the follower agents.

[0011] Furthermore, the leading agent obtains image information through a monocular camera, inputs the image information into the first step, then inputs the output result of the first step into the second step, and determines the target recognized by the system according to the output result of the second step. Specifically,

[0012] Input the edge line template image of the target into the system, obtain images in real time through a monocular camera, execute the edge detection algorithm in the first step, perform filtering processing and gradient magnitude and direction calculation on the original image containing the target, and obtain the edge feature information in the image by using the adaptive threshold method to obtain a binary image containing the target edge line;

[0013] Execute the template matching algorithm in the second step, generate multi-scale and multi-angle template images from the input target template image, match and judge the similarity with the extracted binary image containing the target edge line, and obtain the bounding box area and center point coordinates after successful matching.

[0014] Furthermore, the third step includes,

[0015] Execute the improved kernel correlation filtering algorithm, use the target bounding box and center coordinates output in the second step as the two parameters of the target initial size and position respectively to initialize the tracker, generate a kernel correlation filter for tracking and train it, search for the target position in each subsequent frame, thereby controlling the movement mode of the leading agent, and display the tracking status in real time on the ground station.

[0016] Furthermore, the follower agents calculate their expected positions at the next moment and control their own movement trajectories. When changing the formation according to the actual task requirements, the follower agents calculate their expected positions in the formation through the position calculation algorithm according to the change amount of their own formation positions, and the change amount is the expected angle and distance from the leading agent, so as to realize the formation transformation when tracking the target.

[0017] Further, perform an edge detection algorithm to filter the original image containing the target and calculate the gradient magnitude and direction. Obtain the edge feature information in the image by using an adaptive threshold method. Specifically,

[0018] Perform Gaussian filtering on the input image. Wherein, the value of each pixel represents the brightness or color of the image. When applying the Gaussian filter, each pixel in the image is convolved with a two-dimensional Gaussian function;

[0019] For the image after the filtering process, use the Sobel operator to calculate the gradient and intensity of each pixel point in the image. Among them, the convolution kernel template used for calculation is extended to , , 45°, and 135° in four directions, and obtain the gradient magnitudes corresponding to the four directions through convolution operations with the gradient templates in the four directions;

[0020] Divide the image into background and foreground, and maximize the inter-class variance so that the difference between the background and foreground after threshold segmentation is maximized;

[0021] Divide the gray levels according to the gray information of the image, divide the image into a first part image and a second part image, obtain the first probability and the second probability of any pixel in the image in the first part image and the second part image, calculate the first pixel average gray value and the second pixel average gray value according to the first probability and the second probability, calculate the image average gray value and the inter-class variance according to the first pixel average gray value and the second pixel average gray value, obtain the current optimal threshold when the inter-class variance is the largest, use the optimal threshold to distinguish strong and weak edge lines, and thus connect the strong and weak edge lines, and finally process the image into a binary image that only retains the edge line features.

[0022] Further, perform the second-stage template matching algorithm. Generate multi-scale and multi-angle template images from the input target template image, and match and judge the similarity with the binary image containing the target edge line extracted. After successful matching, mark the target with a red square frame as the bounding box of the target. At this time, the system will output the area size of the bounding box and the coordinates of the center point of the bounding box, which respectively represent the size and position of the target in the system's field of view at the current moment. Specifically,

[0023] For the target template image of the input system, the odd rows and columns of the target image are sampled through a Gaussian convolution kernel to gradually reduce the image resolution layer by layer. The image resolution after each sampling is 1 / 4 of the previous layer's image, thereby generating an image pyramid composed of a set of target template images with different resolutions; different scaling factors (such as 0.25, 0.5, 1.5, etc.) are set for the template images in different levels, and different rotation angles (such as 45°, 90°, 180°, etc.) are set for the template images in each layer.

[0024] Perform a normalized cross-correlation operation on the binary image containing the target edge line output in the first step and a template image of the same size to determine the matching similarity between the current target image and the template image, and judge whether the matching is successful. If the matching is successful, it means that the target is successfully recognized, output the area of the target bounding box and the coordinates of the center point of the bounding box, and set the current moment as the starting frame of the tracking process.

[0025] Furthermore, execute the improved kernel correlation filtering algorithm, use the target bounding box size and center coordinates output in the second step as the two parameters of the target initial size and position respectively to initialize the tracker, generate a correlation filter for tracking and train it. Then search for the target position in each subsequent frame, thereby controlling the movement mode of the leading intelligent agent and displaying the tracking status in real time on the ground station. Specifically,

[0026] Use the target bounding box output in the second step as the initial tracking area of the kernel correlation filtering algorithm, and the center coordinates as the target initial position. Extract features to generate a correlation classifier for tracking, and use the target and non-target regions as input samples for training. Map the samples to a high-dimensional space through a mapping function. After the training classifier step is completed, input the image features of the next frame into the classifier to judge the target position. The maximum value of the response is the predicted current target position. After determining the new target position, use the current target area as a training sample to update the classifier to obtain a new detection model, thereby correcting the tracking box in real time;

[0027] Set a threshold through the maximum response value and the average peak correlation energy to judge whether the target is lost, and judge in real time whether the target is lost. When the target is lost, stop the update, and at the same time display the tracking status in real time on the ground end;

[0028] The output result determines the movement mode of the leading intelligent agent.

[0029] In a second aspect, the present application discloses a monocular vision-based UAV formation cooperative target tracking device, including:

[0030] An initialization module, configured to select leading and following agents according to the current initialization stage of the system, establish a communication relationship between each agent, where a drone is selected as the leading agent, and a monocular camera is mounted on the drone;

[0031] A matching module, configured to enable the leading agent to obtain image information through the monocular camera, input the image information into a first stage, then input the output result of the first stage into a second stage, and determine the target recognized by the system according to the output result of the second stage; the first stage is edge detection calculation, and the second stage is template matching calculation;

[0032] A decision-making module, configured to enable the leading agent to input the target position information into a third stage, and make a decision on the movement mode of the leading agent according to the output result of the third stage; the third stage is single-target tracking calculation;

[0033] A solution module, configured to calculate the relative positions between the following agents and between the following agents and the leading agent, and broadcast the movement decision of the leading agent to the following agents.

[0034] In a third aspect, the present application discloses an electronic device, including:

[0035] A memory, configured to store a computer program;

[0036] A processor, configured to execute the computer program to implement the steps of the aforementioned method for collaborative target tracking of an unmanned aerial vehicle formation based on monocular vision.

[0037] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned method for collaborative target tracking of an unmanned aerial vehicle formation based on monocular vision are implemented.

[0038] As can be seen from the above, the present invention discloses a method, device, equipment and medium for collaborative target tracking of UAV formation based on monocular vision, which relates to the fields of multi-rotor UAV formation control and visual target recognition and tracking technologies. The leader-follower algorithm and the position calculation algorithm are adopted to optimize the formation control effect. For the leader UAV equipped with a monocular camera, the static or dynamic target is detected and recognized by the adaptive edge detection algorithm and the adaptive template matching algorithm, and the target position is marked. The kernel correlation filtering algorithm is used to process the target, so as to realize the tracking of the target by the leader UAV and improve the tracking accuracy. The formation shape information and the position information of the leader UAV are shared among the formation members, and the expected positions of each follower UAV at the next moment are calculated by the position calculation algorithm, and the formation shape is formed to realize the collaborative tracking of the target by the formation. After the static or dynamic target is recognized by the monocular vision sensor, the UAV formation realizes the collaborative tracking of the target while maintaining the predetermined formation shape or changing the formation shape. The present invention combines the improved formation control method with the target recognition and tracking method based on monocular vision to improve the system performance and ensure that the formation completes the task safely and efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0040] Figure 1 Schematic diagram of a quadrotor UAV system provided by an embodiment of the present invention;

[0041] Figure 2 Schematic diagram of a target recognition and tracking algorithm based on monocular vision provided by an embodiment of the present invention;

[0042] Figure 3 Schematic diagram of an image pyramid strategy provided by an embodiment of the present invention.

[0043] Figure 4 Schematic diagram of two-dimensional plane for calculating the expected position of the formation provided by an embodiment of the present invention;

[0044] Figure 5 Schematic diagram of a device for collaborative target tracking of UAV formation based on monocular vision provided by an embodiment of the present invention;

[0045] Figure 6 Schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] The embodiment of the present invention provides a method for cooperative target tracking of an unmanned aerial vehicle (UAV) formation based on monocular vision, which solves the problem that it is difficult for a multi-agent formation to identify and track a target under environmental interference. Through an improved target recognition and tracking method and a formation member position calculation algorithm, vision-based formation target recognition and tracking are realized. For problems such as light change and field of view angle change in the actual task environment, the leader UAV obtains image information through a monocular vision sensor, implements an improved adaptive Canny edge detection algorithm and an adaptive template matching algorithm to complete target recognition and mark the target position. The leader agent runs an improved Kernel Correlation Filter (KCF) algorithm based on the target recognition result to enter the tracking state of the target, and broadcasts the result to each follower agent. The system combines the formation information and controls the formation through the position calculation algorithm, so as to realize the cooperative tracking of the target by the formation. As Figure 1 shown, multiple UAVs and the ground terminal form a system communication network. Among them, the UAVs adopt quadrotor UAVs, and the wheelbase and model are not fixed. The agents establish communication through the equipped ZigBee communication module to transmit data such as the positions and angles of each agent; the agents and the ground terminal establish communication through the Wi-Fi module to detect image information and target recognition results in real time from the ground terminal.

[0048] The present invention proposes a method for cooperative target tracking of an unmanned aerial vehicle (UAV) formation based on monocular vision, as Figure 2 shown, the method includes:

[0049] Step S100, according to the current initialization stage of the system, select the leader and follower agents, and establish the communication relationship between each agent. Among them, select a UAV as the leader agent, and install a monocular camera on the UAV.

[0050] In this embodiment, the current initialization stage of the system includes selecting the leader agent, the follower agents and setting the initial formation parameters; establishing the communication relationship between each agent and with the ground station. Among them, the ground station saves the IPv4 address of the leader agent in the Wi-Fi network, and saves the MAC address of each agent in the Zigbee network as the initial formation parameters; the leader agent turns on the monocular camera and transmits the image back to the ground station in real time.

[0051] In this embodiment, a drone is selected as the leader agent, and a monocular camera is mounted on the drone. Specifically, the leading drone is equipped with a monocular camera and an on-board computer for real-time acquisition and processing of image data. In the system initialization stage, the leader-follower method is adopted to set the initial formation of the system according to the actual situation and mission requirements. A drone is selected as the leader agent, and the positions of the remaining follower agents and their positional relationships with the leader agent are determined. The agents establish a communication relationship using a Zigbee-based communication method to achieve real-time information interaction. A communication relationship is established between the agents and the ground terminal using a Wi-Fi-based communication method, and the image data information is displayed in real time on the ground terminal.

[0052] Step S200: The leader agent obtains image information through the monocular camera, inputs the image information into the first stage, inputs the output result of the first stage into the second stage, and determines the target recognized by the system according to the output result of the second stage; the first stage is edge detection calculation, and the second stage is template matching calculation;

[0053] In this embodiment, an improved adaptive Canny edge detection algorithm is implemented on the leader agent to obtain a binary image containing the target edge line features, and an improved adaptive template matching algorithm is implemented on the input template image. The obtained result is used as the result of target recognition, providing reliable perception data for the leader agent and the formation to perform tasks, and the recognition result is displayed in real time on the ground terminal.

[0054] In this embodiment, in order to enhance the anti-interference ability of the system and achieve performance optimization, the target recognition process of the leader agent is divided into two stages. After real-time acquisition of image information through the mounted monocular camera and on-board computer, first, an improved adaptive Canny edge detection algorithm is implemented to perform filtering processing, gradient magnitude and direction calculation on the original image that may contain the target, and an adaptive threshold method is used to obtain the edge feature information in the image, and a binary image containing the target edge line features is output. Then, an improved adaptive template matching algorithm is adopted to match the target image with the generated multi-scale and multi-angle template images. After successful matching, the target is marked with a red box as the bounding box of the target, and the size and center point coordinates of the target bounding box are output, representing the size and position of the target in the system's field of view respectively. This method can effectively cope with environmental interference to achieve performance optimization of the system.

[0055] Step S300: The leader agent inputs the position information of the target into the third stage, and determines the movement mode of the leader agent according to the output result of the third stage; the third stage is single-target tracking calculation;

[0056] Among them, the leader agent in the airborne computer implements an improved KCF algorithm to track the motion trajectory of the target in the image frame in real time. This algorithm is used to predict the motion trajectory of the target, convert the result into a control quantity, control the motion mode of the leader agent, and display the tracking status in real time on the ground end.

[0057] In this example, to save the computing resources of the system, the target recognition result in step S200 is used as a reliable input, and the moment when the target recognition is successful is used as the starting frame. The leader agent implements an improved KCF algorithm, uses the target bounding box as the initial tracking area of the kernel correlation filtering algorithm, the center coordinates as the initial target position, extracts features to generate a kernel correlation filter for tracking and trains it, and is used to search for the target position in each subsequent frame. The maximum response value and the average peak correlation energy are added to the algorithm as a method to confirm the target confidence. When it is judged that the target is lost during the tracking process, the detection update step in the algorithm is stopped, so as to save the computing resources of the system while tracking the target in real time.

[0058] Step S400, calculate the relative positions between each follower agent and the leader agent, and broadcast the motion decision of the leader agent to the follower agents.

[0059] Among them, the motion mode of the leader agent depends on the motion mode of the target. Each follower agent calculates its expected position at the next moment in real time through a position calculation algorithm according to the motion trajectory of the leader agent. In this embodiment, when the leader agent enters the target tracking state, its own position information is broadcast to each follower agent through Zigbee. Each follower agent calculates its expected position at the next moment according to the information of the leader agent and the position calculation algorithm, and controls its own motion trajectory. When changing the formation shape according to the actual task requirements, each follower agent combines its own formation position change amount with the position calculation algorithm to realize the formation transformation when tracking the target in formation.

[0060] In this embodiment, formation control is achieved by using the leader-follower method. Any member of the formation is selected as the leader agent (leader), and the rest are follower agents (followers). Communication relationships are established among the agents through Zigbee to interact position, angle and other information in real time. A communication relationship is established between the agents and the ground terminal through Wi-Fi, and the motion state of the formation is displayed on the ground terminal in real time. The leader agent obtains image information in real time through the equipped monocular camera and on-board computer, and implements object recognition algorithms and object tracking algorithms based on monocular vision to control its own real-time tracking of the target's motion trajectory. In this embodiment, the leader agent runs an improved adaptive Canny edge detection algorithm. After performing Gaussian filtering on the image, it detects and extracts the edge feature information of the image, and implements an improved adaptive template matching algorithm. The object recognition result is obtained from the matching result, and the object size and position are output. On this basis, an improved KCF algorithm is run to achieve stable tracking of the target, and the tracking state is displayed on the ground terminal in real time. According to the actual task requirements, the leader agent sends its own position information to each follower agent, and calculates the expected positions of each follower agent at the next moment through the position calculation algorithm and the formation requirements. Thus, the motion state information of the leader agent is shared with each follower agent, enabling the entire formation system to track the target, effectively improving the autonomy and performance of the system.

[0061] In this embodiment, after the leader agent obtains image information through the monocular camera, in view of the fact that the process of obtaining images by using the camera during the flight of the unmanned aerial vehicle will be affected by factors such as illumination changes and interference from other objects, the Canny edge detection algorithm is improved to improve the edge detection efficiency and accuracy of the target. The improved adaptive Canny edge detection algorithm can adaptively adjust the threshold according to the actual environment to obtain more accurate target edge feature information, and effectively reduce the influence of illumination changes and interference from other objects on the edge feature extraction, providing more reliable data for subsequent object recognition. Specifically, in step S200, the first processing link includes: extracting the edge features of the image through the improved adaptive Canny edge detection algorithm, and the specific operations are as follows:

[0062] First, Gaussian filtering is performed on the input image. In the image, the value of each pixel represents the brightness or color of the image. When applying the Gaussian filter, each pixel in the image is convolved with a two-dimensional Gaussian function. The form of the two-dimensional Gaussian function is as shown in Equation (1):

[0063] (1)

[0064] In the formula represents the weight of the Gaussian function, is the standard deviation of the Gaussian function, and Indicates the distance from a pixel to the central pixel. The Gaussian filter can reduce the noise in the image, blur the image to reduce details, and facilitate subsequent image processing steps. Its smoothing effect depends on the standard deviation of the Gaussian function. The larger the standard deviation, the more obvious the smoothing effect of the filter.

[0065] For the image after the filtering process, the Sobel operator is used to calculate the gradient and intensity of each pixel point in the image. Among them, the convolution kernel template used for calculation is extended to , , 45°, and 135°. The gradient magnitude is defined as obtained by performing a convolution operation with the gradient templates in the four directions , , and . The overall gradient magnitude and the gradient direction are calculated by the formulas shown in Equations (2) and (3):

[0066] (2)

[0067] (3)

[0068] After obtaining the gradient magnitude and direction from Equations (2) and (3), non-maximum suppression is performed along the gradient direction, which can effectively remove non-edge points, and then a threshold is selected to connect the edge lines.

[0069] Based on the Otsu method, the image is divided into two categories (background and foreground), maximizing the between-class variance, that is, maximizing the difference between the two categories after threshold segmentation. According to the gray information of the image, an image of size is divided into gray levels. Let be the number of pixels with gray level , then the probability corresponding to the pixels with gray level is . Let the selected threshold be , , and the image is divided into and two parts, where is composed of all pixels with gray values within , and is composed of all pixels with gray values within . Then the probabilities and of any pixel in the image in and are

[0070] (4)

[0071] From equation (4), we can obtain and the average gray value of the pixels in and is

[0072] (5)

[0073] (6)

[0074] From equations (5) and (6), the average gray value of the image with size can be obtained as which can be expressed as

[0075] (7)

[0076] By calculating through equation (7) and substituting into equation (4), the between-class variance is

[0077] (8)

[0078] Thus, when the between-class variance is maximized, the value is the optimal threshold, denoted as . If the existing is not unique, the average value of all values is taken as the optimal threshold. On this basis, using the result of threshold determination, the weak edge pixels and strong edge pixels are connected to form a complete edge, and finally a binary image containing the target edge line feature is obtained.

[0079] From the above operations, the edge feature information of the target can be obtained, which overcomes problems such as illumination changes to a certain extent and provides more efficient and reliable data for subsequent target recognition.

[0080] In this embodiment, aiming at the deformation of the target in scale and angle caused by the change of flight altitude and camera viewing angle during the flight of the UAV, an adaptive template matching algorithm is adopted. By generating templates for multi-scale and multi-angle matching and matching with the target, the accuracy of target recognition is improved. Specifically, in step S200, the second processing link includes: implementing target recognition through an improved adaptive template matching algorithm and outputting the target position and size. For example:

[0081] For the target template image input to the system, the odd rows and columns of the template image are sampled through a Gaussian convolution kernel to gradually reduce the image resolution. The image resolution after each sampling is 1 / 4 of the previous layer's image, as Figure 3As shown, an image pyramid composed of a set of target template images with different resolutions is generated, where the template images in different levels are scaled with different scaling factors (such as 0.25, 0.5, 1.5, etc.), and different rotation angles (such as 45°, 90°, 180°, etc.) are set for the template images in each level.

[0082] For target images and with sizes and template images , a normalized cross-correlation operation is performed to measure the similarity of the images. The calculation formula is as shown in Equation (9):

[0083] (9)

[0084] In the formula, and are the pixel values in the target image and the template image respectively, and are the means of the target image and the template image respectively, and represents the matching degree starting from the coordinate in the image. According to the calculation result of the normalized cross-correlation, the matching degree score between the current target and the template image is determined to judge whether the matching is successful. If the matching is successful, it means that the target is successfully recognized.

[0085] Through the above operations, target recognition can be accurately completed. At the same time, the target is marked with a red box as the bounding box of the target, and the size and center point coordinates of the target bounding box are output, representing the size and position of the target in the system's field of view respectively, as the input for the next stage.

[0086] In this embodiment, aiming at the problem of target loss that occurs when the drone tracks a dynamic target, the KCF algorithm is improved. During the tracking process, it is determined whether the target is lost by judging the target confidence, which effectively improves the tracking accuracy and saves computing resources. Specifically, in step S300, the third processing link includes: the leader agent implements the improved KCF algorithm to track the motion trajectory of the target in real time. For example:

[0087] Using the target recognition result in S200, with the target and non-target regions as input samples, a ridge regression model is used to train the kernel correlation filter, that is, to train the classifier ( is the classifier weight coefficient, is the basic sample), to minimize the average error between the training sample and the regression label :

[0088] (10)

[0089] In the formula represents the regularization parameter introduced to prevent overfitting. The filter weights that can match the target are obtained through training, which are used to match the image patch with the target template and output the matching degree score.

[0090] Since target tracking is a non-linear problem, mapping the samples through the mapping function to a high-dimensional space can make the non-linear problem linearly separable, and the weight coefficients of the classifier can be expressed as a linear combination of the mapped samples as shown in Equation (11):

[0091] (11)

[0092] In the formula is the linear combination coefficient. Substituting it into the classifier and Equation (10), the vector composed of has the expression as shown in Equation (12):

[0093] (12)

[0094] In the formula is the kernel function, is the regression label The column vector composed of. Perform Fourier transform on Equation (12):

[0095] (13)

[0096] In the formula is in Fourier form; is the Fourier form of the first row of the kernel function When the training of the classifier is completed, input the features of the next frame of the image into the classifier to determine the target position. Let be the current training sample, be the test sample of the next frame, and use to represent and The kernel function of, and solve for the response in the frequency domain:

[0097] (14)

[0098] In the formula represents the convolution operator. The maximum value of the response is the predicted current target position. After determining the new target position, use the current target area as a training sample to update the classifier to obtain a new detection model, so as to correct the tracking box in real time.

[0099] When problems such as partial occlusion and target loss occur in the previous frame of the tracking process, resulting in tracking failure in the next frame, the maximum response value and the average peak correlation energy are used as methods to confirm the target confidence, which are used to determine whether the target is lost. The calculation formula is as follows:

[0100] (15)

[0101] In the formula, represents the horizontal translation of the target , and the vertical translation of the detection response value after that; and represent the maximum and minimum responses of the current frame; represents the arithmetic mean operation; represents the th row and th column pixel response value of the current frame. By setting an appropriate threshold, it can be determined whether the target is lost. When the target is lost, the update is stopped.

[0102] Through the above operations, the leader agent can track the target with high precision, judge in real time whether the target is lost, and display the tracking status on the ground end in real time.

[0103] In this embodiment, in order to meet the real-time performance of formation control, a classic leader-follower control algorithm is adopted. According to the desired formation, the follower agents are set to follow the leader agent at a specific distance and angle. After the follower agents learn the position information of the leader agent, they calculate their own desired positions and yaw angles through a position calculation algorithm, so as to maintain the preset formation. Specifically, in S400, it includes:

[0104] As Figure 4 described, during the execution of tasks by multi-agent formations, when the leader agent enters the target tracking state, each follower agent in the formation needs to calculate its own desired position. After learning the desired position point, it controls itself to reach the desired position to maintain the formation. Let be the world coordinate system, be the body coordinate system of the leader UAV. The follower follows the leader at a distance , and an angle . , is the component of the following distance in the body coordinate system of the leader. , are the positions of the leader and the follower in the world coordinate system respectively, is the angle between the body coordinate system and the world coordinate system. The desired following distance , It can be expressed as:

[0105] (16)

[0106] Through coordinate system transformation, the actual following distance between the follower and the leader , It can be expressed as:

[0107] (17)

[0108] From equation (17), the expected position coordinates of the UAV in the world coordinate system can be deduced , as:

[0109] (18)

[0110] In the formation system, the follower UAVs can calculate their own expected positions from the above formulas. On this basis, by controlling the movement trajectories of each follower UAV, the formation shape can be maintained when the system executes the target tracking task.

[0111] It should be noted that when changing the formation shape according to the actual task requirements, the formation change information is sent from the ground end to each intelligent agent. The follower intelligent agents combine their own formation position change amounts with the position calculation algorithm, and the formation change during target tracking can be realized.

[0112] The embodiment of the present invention discloses a method for cooperative target tracking of UAV formation based on monocular vision, which relates to the field of multi-rotor UAV formation control and visual target recognition and tracking, and can realize formation cooperative tracking after identifying static or dynamic targets. The present invention includes: adopting a leader-follower algorithm and an improved position calculation algorithm to optimize the formation control effect; considering interference problems such as changes in light and field of view angle, for the leader UAV equipped with a monocular camera, an improved adaptive Canny edge detection algorithm and an adaptive template matching algorithm are implemented. The method of using an adaptive threshold is used to extract the edge feature information of the target in the field of view and match it with the generated multi-scale and multi-angle template images to improve the accuracy and real-time performance of target recognition in a dynamic environment; aiming at the problem of target tracking loss in a dynamic scene, after target recognition is completed, the leader UAV runs an improved KCF algorithm to perform real-time tracking on the detected stationary or moving target, and sends information such as position and movement speed to the other members in the formation in real time. Each follower UAV combines the formation shape information and the position information of the leader UAV, and calculates its own expected position at the next moment through the position calculation algorithm to realize formation cooperative tracking of the target. The present invention combines an improved formation control method with a target recognition and tracking method based on monocular vision to improve the system performance and ensure that the formation completes the task safely and efficiently.

[0113] An embodiment of the present invention provides a method for collaborative target tracking of an unmanned aerial vehicle (UAV) formation based on monocular vision. The method is used in a system, and the system includes: multiple agents and a ground terminal. The multiple agents include at least two UAVs. As Figure 1 shown, multiple UAVs and the ground terminal form a system communication network. Among them, the UAVs adopt quadrotor UAVs, and the wheelbase and model are not fixed. The agents establish communication through the equipped ZigBee communication module to transmit data such as the positions and angles of each agent; the agents and the ground terminal establish communication through the Wi-Fi module to detect image information and target recognition results in real time from the ground terminal.

[0114] The present invention also proposes a device for collaborative target tracking of an unmanned aerial vehicle formation based on monocular vision, as shown in the appendix Figure 5 shown. The device 10 includes an initialization module, a matching module, a decision-making module, and a solution module.

[0115] The initialization module is used to select a leader and follower agents according to the current initialization stage of the system, and establish a communication relationship between each agent. Among them, a UAV is selected as the leader agent, and a monocular camera is carried on the UAV.

[0116] The matching module is used for the leader agent to obtain image information through the monocular camera, input the image information into the first link, and then input the output result of the first link into the second link, and determine the target recognized by the system according to the output result of the second link; the first link is edge detection calculation, and the second link is template matching calculation.

[0117] The decision-making module is used for the leader agent to input the target position information into the third link, and make a decision on the movement mode of the leader agent according to the output result of the third link; the third link is single-target tracking calculation.

[0118] The solution module is used to calculate the relative positions between the follower agents and between the follower agents and the leader agent, and broadcast the movement decision of the leader agent to the follower agents.

[0119] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the method for collaborative target tracking of an unmanned aerial vehicle formation based on monocular vision executed by the electronic device disclosed in any of the foregoing embodiments.

[0120] In this embodiment, the power supply 23 is used to provide operating voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed thereon herein; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type thereof can be selected according to specific application requirements, and no specific limitation is imposed thereon herein.

[0121] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0122] In addition, the memory 22, as a carrier for resource storage, may be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc., and the resources stored thereon include an operating system 221, a computer program 222, data 223, etc., and the storage method may be transient storage or permanent storage.

[0123] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, so as to implement the operation and processing of the massive data 223 in the memory 22 by the processor 21. It can be Windows, Unix, Linux, etc. In addition to the computer program that can be used to complete the monocular vision-based UAV formation cooperative target tracking method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks. The data 223 may include not only the data transmitted by external devices received by the electronic device, but also the data collected by its own input / output interface 25, etc.

[0124] Furthermore, the embodiment of the present application also discloses a computer-readable storage medium, in which a computer program is stored. When the computer program is loaded and executed by a processor, the steps of the monocular vision-based UAV formation cooperative target tracking method disclosed in any of the foregoing embodiments are implemented.

[0125] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0126] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0127] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0128] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0129] The above has introduced in detail a method, device, equipment and storage medium for cooperative target tracking of UAV formation based on monocular vision provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for collaborative target tracking of UAV formation based on monocular vision, characterized in that, Including: According to the current initialization stage of the system, select the leader and follower agents, and establish the communication relationship between each agent. Among them, select a drone as the leader agent and install a monocular camera on the drone. The leader agent obtains image information through the monocular camera, inputs the image information into the first step, then inputs the output result of the first step into the second step, and determines the target recognized by the system according to the output result of the second step. The first step is the calculation of the edge detection algorithm, and the second step is the calculation of the template matching algorithm. The leader agent inputs the target position information into the third step and decides the movement mode of the leader agent according to the output result of the third step. The third step is the calculation of single-target tracking. Calculate the relative positions between the follower agents and between the follower agents and the leader agent, and broadcast the movement decision of the leader agent to the follower agents. The leader agent obtains image information through the monocular camera, inputs the image information into the first step, then inputs the output result of the first step into the second step, and determines the target recognized by the system according to the output result of the second step. Specifically, Input the edge line template image of the target into the system, obtain images in real time through the monocular camera, execute the edge detection algorithm in the first step, perform filtering processing, gradient magnitude and direction calculation on the original image that may contain the target, and obtain the edge feature information in the image by using the adaptive threshold method to obtain a binary image containing the target edge line. Execute the template matching algorithm in the second step, generate multi-scale and multi-angle template images from the input target template image, and perform matching to judge the similarity with the extracted binary image containing the target edge line. After successful matching, output the area of the target bounding box and the coordinates of the center point. The execution of the template matching algorithm in the second step, generate multi-scale and multi-angle template images from the input target template image, and perform matching to judge the similarity with the extracted binary image containing the target edge line. After successful matching, output the current size and position of the target. The third step includes executing an improved kernel correlation filtering algorithm, using the area of the target bounding box and the coordinates of the center point output in the second step as the size and position of the target at the current moment respectively, initializing the tracker, generating a classifier for tracking and training, searching for the target position in each frame, thereby controlling the movement mode of the leader agent, and displaying the tracking status in real time on the ground station.

2. The method for collaborative target tracking of UAV formation based on monocular vision according to claim 1, wherein, The follower agent calculates its expected position at the next moment and controls its own movement trajectory. When changing the formation according to the actual task requirements, the follower agent calculates its expected position in the formation through the position calculation algorithm according to the change amount of its own formation position, and the change amount is the expected angle and distance from the leader agent, so as to realize the formation transformation when tracking the target.

3. The method for collaborative target tracking of UAV formation based on monocular vision according to claim 2, wherein Execute the edge detection algorithm, perform filtering processing, gradient magnitude and direction calculation on the original image that may contain the target, and obtain the edge feature information in the image by using the adaptive threshold method. Specifically, Perform Gaussian filtering on the input image; wherein, the value of each pixel represents the brightness or color of the image. When applying the Gaussian filter, each pixel in the image is convolved with a two-dimensional Gaussian function; For the image after the filtering process, use the Sobel operator to calculate the gradient and intensity of each pixel in the image; among them, the convolution kernel template used for calculation is extended to , , and four directions of 45° and 135°. The gradient amplitudes corresponding to the four directions are obtained through convolution operations with the gradient templates in the four directions. Divide the image into background and foreground, and maximize the between-class variance so as to maximize the difference between the background and foreground after threshold segmentation; Divide the gray levels according to the gray information of the image, divide the image into a first part image and a second part image, obtain the first probability and the second probability of any pixel in the image in the first part image and the second part image, calculate the first pixel average gray value and the second pixel average gray value according to the first probability and the second probability, calculate the image average gray value and the between-class variance according to the first pixel average gray value and the second pixel average gray value, obtain the current optimal threshold when the between-class variance is the largest, use the optimal threshold to distinguish strong and weak edge lines, and thus connect the strong and weak edge lines, and finally process the image into a binary image that only retains the edge line features.

4. The method for collaborative target tracking of UAV formation based on monocular vision according to claim 2, wherein Execute the second-stage template matching algorithm. Generate multi-scale and multi-angle template images from the input target template image, and match and judge the similarity with the binary image containing the target edge line extracted. After successful matching, output the current size and position of the target. Specifically, For the target template image input to the system, sample the odd rows and columns of the template image through a Gaussian convolution kernel to gradually reduce the image resolution. The image resolution after each sampling is 1 / 4 of the previous layer's image, so as to generate a set of target template images with different resolutions to form an image pyramid; Set different scaling factors for the template images in different levels for scaling, and set different rotation angles for the template images in each layer; Perform normalized cross-correlation operation on the binary image containing the target edge line output in the first stage and the template image of the same size to determine the matching similarity between the current target image and the template image, and judge whether the matching is successful. If the matching is successful, it means that the target is successfully recognized, output the area of the target bounding box and the coordinates of the center point of the bounding box, and set the current moment as the starting frame of the tracking process.

5. The method for collaborative target tracking of UAV formation based on monocular vision according to claim 3, characterized in that Execute the improved kernel correlation filtering algorithm. Use the target bounding box and the center coordinates output in the second stage as the two parameters of the target initial size and position respectively to initialize the tracker, generate a correlation filter for tracking and train it. Then search for the target position in each subsequent frame, thereby controlling the movement mode of the leading intelligent agent, and real-time display the tracking status on the ground station. Specifically, Use the target bounding box output in the second stage as the initial tracking area of the kernel correlation filtering algorithm, and the center coordinates as the target initial position. Extract features to generate a kernel correlation filter for tracking, and use the target and non-target areas as input samples for training. Map the samples to a high-dimensional space through a mapping function. After the training classifier step is completed, input the features of the next frame image into the classifier to judge the target position. The maximum value of the response is the predicted current target position. After determining the new target position, use the current target area as a training sample to update the classifier to obtain a new detection model, so as to correct the tracking box in real time; Set a threshold to determine whether the target is lost based on the maximum response value and the average peak correlation energy, and determine in real time whether the target is lost. When the target is lost, stop the update, and at the same time, display the tracking status on the ground end in real time; Output the result to decide the movement mode of the leading agent.

6. An unmanned aerial vehicle formation cooperative target tracking device based on monocular vision, characterized in that It includes: An initialization module, which is used to select the leading and following agents according to the current initialization stage of the system, and establish the communication relationship between each agent. Among them, a drone is selected as the leading agent, and a monocular camera is carried on the drone; A matching module, which is used for the leading agent to obtain image information through the monocular camera, input the image information into the first link, and then input the output result of the first link into the second link, and determine the target recognized by the system according to the output result of the second link; the first link is edge detection calculation, and the second link is template matching calculation; A decision-making module, which is used for the leading agent to input the target position information into the third link, and decide the movement mode of the leading agent according to the output result of the third link; the third link is single-target tracking calculation; A solution module, which is used to calculate the relative positions between the following agents and the leading agent, and broadcast the movement decision of the leading agent to the following agents; The leading agent obtains image information through the monocular camera, inputs the image information into the first link, and then inputs the output result of the first link into the second link, and determines the target recognized by the system according to the output result of the second link. Specifically, Input the edge line template image of the target into the system, obtain images in real time through the monocular camera, execute the edge detection algorithm in the first link, perform filtering processing and gradient magnitude and direction calculation on the original image that may contain the target, and obtain the edge feature information in the image by using the adaptive threshold method to obtain a binary image containing the target edge line; Execute the template matching algorithm in the second link, generate multi-scale and multi-angle template images from the input target template image, and match and judge the similarity with the extracted binary image containing the target edge line. After successful matching, output the target bounding box area and the center point coordinates; the execution of the template matching algorithm in the second link, generate multi-scale and multi-angle template images from the input target template image, and match and judge the similarity with the extracted binary image containing the target edge line. After successful matching, output the current size and position of the target; The third link includes executing an improved kernel correlation filtering algorithm, using the target bounding box area and the center point coordinates output in the second link as the size and position of the target at the current moment respectively, initializing the tracker, generating a classifier for tracking and training, searching for the target position in each frame, thereby controlling the movement mode of the leading agent, and displaying the tracking status on the ground station in real time.

7. An electronic device, characterized in that, It includes: A memory, which is used to save computer programs; A processor, which is used to execute the computer program to implement the steps of the method for collaborative target tracking of an unmanned aerial vehicle formation based on monocular vision according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the monocular vision-based UAV formation cooperative target tracking method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Fixed-wing unmanned aerial vehicle cluster affine formation control method based on pilot following mode

    CN113741518A

  • Multi-unmanned aerial vehicle intelligent identification and relative positioning method based on monocular vision

    CN114581516A