Trajectory tracking method and system for multi-rotor unmanned aerial vehicle
Through video image recognition technology and deep learning network, the problem of position misdetection in multi-rotor drone flight missions is solved, and more accurate trajectory tracking is achieved.
Patent Information
- Application Number
- CN202210445628.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-04-24
AI Technical Summary
In the prior art, when performing flight missions, multi-rotor drones are prone to positional misdetection problems through radar, sound detection and radio frequency detection, resulting in missed detection and misdetection.
Video images are used for object recognition detection, combined with the first convolutional neural network, UAV-FPN neural network and YOLOHead prediction network, the location and category of multi-rotor drones are identified, and appearance feature extraction and motion trajectory matching are performed through the second convolutional neural network and DeepSort algorithm.
Vision recognition technology significantly reduces the probability of position error detection of multi-rotor drones and improves the tracking accuracy of the motion trajectory of multi-rotor drones.
Smart Images

Figure CN114757974B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle applications, and in particular to a trajectory tracking method and system for a multi-rotor unmanned aerial vehicle. Background Art
[0002] Traditional drone aerial monitoring systems usually rely on radar detection. However, multi-rotor drones have a small radar detection area, and when using radar to detect the three-dimensional position of the drone to capture its flight trajectory, missed detections often occur. Later, it was proposed to use sound detection and radio frequency detection to detect the position of the drone, but both detection methods require sharing frequency bands, which increases the probability of false detection. Therefore, how to reduce the probability of false detection of the positions of several multi-rotor drones when performing flight missions and achieve efficient trajectory tracking is a technical problem that the present invention needs to solve. Summary of the invention
[0003] The present invention provides a trajectory tracking method and system for a multi-rotor unmanned aerial vehicle to solve one or more technical problems existing in the prior art and at least provide a beneficial option or create conditions.
[0004] An embodiment of the present invention provides a trajectory tracking method for a multi-rotor unmanned aerial vehicle, the method comprising:
[0005] Obtain video images of several multi-rotor drones performing flight missions;
[0006] Performing target recognition detection on each frame of the video image to obtain position detection results and the category of the multi-rotor drones contained in each frame of the video image;
[0007] Extracting the appearance features of each frame of image according to the position detection results of all multi-rotor drones contained in each frame of image, and obtaining the appearance feature detection results of all multi-rotor drones contained in each frame of image;
[0008] The appearance feature detection results of all multi-rotor drones contained in each frame image are applied to the DeepSort algorithm, and the motion trajectories of the multi-rotor drones are matched frame by frame according to the position detection results of all multi-rotor drones contained in each frame image.
[0009] Furthermore, the target recognition detection is performed on each frame of the video image to obtain the position detection results and the category of the drones of all multi-rotor drones contained in each frame of the video image, including:
[0010] Using a pre-built first convolutional neural network to extract backbone features from each frame of the video image, and obtaining low-level feature data and high-level feature data contained in each frame of the video image;
[0011] The UAV-FPN neural network is used to fuse the low-level feature data and high-level feature data contained in each frame of the image to obtain the fused feature data contained in each frame of the image;
[0012] The YOLOHead prediction network is used to perform feature conversion on the fused feature data contained in each frame of the image to obtain the position detection results and the category of the drone to which all multi-rotor drones contained in each frame of the image belong.
[0013] Furthermore, the first convolutional neural network includes an input processing module, a low-level feature extraction module and a high-level feature extraction module connected in sequence; wherein the input processing module is used to extract downsampled feature data from any frame image, the low-level feature extraction module is used to extract low-level feature data from the downsampled feature data, and the high-level feature extraction module is used to further extract high-level feature data from the low-level feature data.
[0014] Furthermore, the input processing module includes an input layer and a Focus structure layer connected in sequence, the low-level feature extraction module includes a first BaseBlock, a second BaseBlock, a first residual convolution layer, a third BaseBlock, a second residual convolution layer, a fourth BaseBlock, a first spatial pyramid pooling layer and a third residual convolution layer connected in sequence, and the high-level feature extraction module includes a fifth BaseBlock, a second spatial pyramid pooling layer and a fourth residual convolution layer connected in sequence.
[0015] Furthermore, any BaseBlock includes an input layer, a two-dimensional convolutional layer, a maximum pooling layer, and an activation layer connected in sequence.
[0016] Furthermore, the appearance feature extraction is performed on each frame of image according to the position detection results of all multi-rotor drones contained in each frame of image, and the appearance feature detection results of all multi-rotor drones contained in each frame of image are obtained, including:
[0017] According to the position detection results of all multi-rotor drones contained in each frame of image, an image of the area where each multi-rotor drone is located is intercepted from each frame of image;
[0018] The pre-built second convolutional neural network is used to extract features from the image of the area where each multi-rotor drone is located, and the appearance feature vector of each multi-rotor drone is obtained.
[0019] Furthermore, the second convolutional neural network includes an input layer, a convolution layer, an average pooling layer and a normalization layer connected in sequence; wherein the convolution layer is used to extract the global appearance feature data of the multi-rotor drone from the image of the area where each multi-rotor drone is located, the average pooling layer is used to perform vector modulus adjustment on the global appearance feature data, and the normalization layer is used to convert the adjusted global appearance feature data into an appearance feature vector.
[0020] In addition, an embodiment of the present invention further provides a trajectory tracking system for a multi-rotor UAV, the system comprising:
[0021] at least one processor;
[0022] at least one memory for storing at least one program;
[0023] When the at least one program is executed by the at least one processor, the at least one processor implements the trajectory tracking method of the multi-rotor drone described in any one of the above items.
[0024] The present invention has at least the following beneficial effects: by combining the first convolutional neural network, the UAV-FPN neural network and the YOLOHead prediction network, it is possible to more accurately identify and detect all multi-rotor drones from each frame of the video image, thereby solving the problem of multi-rotor drones being easily missed when performing flight missions in the prior art from the perspective of vision. By combining the second convolutional neural network and the DeepSort algorithm, the appearance features of multi-rotor drones are fully considered, and the tracking effect of the motion trajectories of several multi-rotor drones can be effectively improved, thereby solving the problem of multi-rotor drones being easily missed when performing flight missions in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation on the technical solution of the present invention.
[0026] Figure 1 1 is a flow chart of a trajectory tracking method for a multi-rotor UAV in an embodiment of the present invention;
[0027] Figure 2 Schematic diagram of the structure of the first convolutional neural network in an embodiment of the present invention;
[0028] Figure 3 Schematic diagram of the structure of the UAV-FPN neural network in an embodiment of the present invention;
[0029] Figure 4It is a schematic diagram of the structural composition of the second convolutional neural network in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0031] It should be noted that, although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0032] Please refer to Figure 1 , Figure 1 1 is a flow chart of a trajectory tracking method for a multi-rotor UAV provided by an embodiment of the present invention, the method comprising the following steps:
[0033] S101. Acquire video images of a plurality of multi-rotor drones performing flight missions.
[0034] In an embodiment of the present invention, the video images of the plurality of multi-rotor drones performing flight missions are collected and acquired by a camera device built on the ground, and the type information of the plurality of multi-rotor drones is different from each other.
[0035] S102: Perform target recognition detection on each frame of the video image to obtain position detection results and the category of the multi-rotor drones contained in each frame of the video image.
[0036] In an embodiment of the present invention, a pre-built first convolutional neural network is first used to perform backbone feature extraction on each frame of the video image to obtain low-level feature data and high-level feature data contained in each frame; secondly, a UAV-FPN neural network is used to perform feature fusion on the low-level feature data and high-level feature data contained in each frame to obtain fused feature data contained in each frame; then, a YOLOHead prediction network is used to perform feature conversion on the fused feature data contained in each frame to obtain position detection results of all multi-rotor drones contained in each frame and the drone category to which they belong.
[0037] In an embodiment of the present invention, the first convolutional neural network includes an input processing module, a low-level feature extraction module and a high-level feature extraction module connected in sequence. Figure 2As shown; wherein, the input processing module is used to extract down-sampled feature data from any frame image with an input scale of 640×640×3, the low-level feature extraction module is used to extract low-level feature data with a scale of 40×40×256 from the down-sampled feature data, and the high-level feature extraction module is used to further extract high-level feature data with a scale of 20×20×512 from the low-level feature data.
[0038] Furthermore, the input processing module includes an input layer and a Focus structure layer connected in sequence, and the Focus structure layer can slice and then stack each frame of image provided by the input layer according to the pixel value to ensure that the feature information of each frame of image is extracted to the maximum extent; the low-level feature extraction module includes a first BaseBlock, a second BaseBlock, a first residual convolution layer, a third BaseBlock, a second residual convolution layer, a fourth BaseBlock, a first spatial pyramid pooling layer and a third residual convolution layer connected in sequence; the high-level feature extraction module includes a fifth BaseBlock, a second spatial pyramid pooling layer and a fourth residual convolution layer connected in sequence; wherein any BaseBlock includes an input layer, a two-dimensional convolution layer, a maximum pooling layer and an activation layer connected in sequence, and the activation layer is provided with a SiLu (Sigmoid Weighted Liner Unit) activation function.
[0039] Due to the small size of the multi-rotor UAV, when it performs a high-altitude flight mission, the proportion of its body in each frame image contained in the video image may be small and the flight distance may change at any time. In order to effectively extract more feature data of the multi-rotor UAV in each frame image, an embodiment of the present invention proposes to extract low-level feature data and high-level feature data of a fixed scale from each frame image, that is, using the first spatial pyramid pooling layer to extract low-level feature data with a scale of 40×40×256, and then using the second spatial pyramid pooling layer to extract high-level feature data with a scale of 20×20×512.
[0040] Since the activation function existing in the first convolutional neural network will cause inevitable feature information extraction loss, and as the number of network layers increases, it is easy to cause network degradation, the embodiment of the present invention proposes to use a residual convolution layer to ensure the network's feature learning ability and prevent network degradation, wherein any residual convolution layer includes an input layer, a channel stacking layer, three BaseBlocks and two two-dimensional convolutional layers, and the channel stacking layer is used to combine the low-level features processed by a single BaseBlock with the high-level features processed by two BaseBlocks and a single two-dimensional convolutional layer.
[0041] In the embodiment of the present invention, the UAV-FPN neural network (wherein UAV is the abbreviation of Unmanned Aerial Vehicle, translated as unmanned aircraft; FPN is the abbreviation of Feature Pyramid Network, translated as feature pyramid network) is mainly used to perform feature fusion on the high-level feature data output by the high-level feature extraction module and the low-level feature data output by the low-level feature extraction module. The UAV-FPN neural network includes BaseBlock, upsampling layer, residual convolution layer, two-dimensional convolution layer and two channel stacking layers, such as Figure 3 As shown, the implementation process is as follows: first, the high-level feature data is processed by the BaseBlock and upsampling layers, and the scale of the high-level feature data is enlarged to be consistent with the scale of the low-level feature data using the nearest neighbor interpolation method; secondly, a single channel stacking layer is used to fuse the enlarged high-level feature data with the low-level feature data to obtain preliminary fused feature data; then, the preliminary fused feature data is processed by the residual convolution layer and the two-dimensional convolution layer, and the scale of the preliminary fused feature data is reduced to be consistent with the scale of the high-level feature data; finally, a single channel stacking layer is used to fuse the reduced-scale preliminary fused feature data with the high-level feature data processed by the BaseBlock to obtain the final fused feature data.
[0042] In an embodiment of the present invention, the YOLOHead prediction network is mainly used to perform feature conversion on the fused feature data, and its network structure includes an Obj branch structure, a Cls branch structure and a Reg branch structure, wherein the Obj branch structure is used to determine whether the fused feature data contains a multi-rotor drone, the Cls branch structure is used to identify the type information of all multi-rotor drones contained in the fused feature data, and the Reg branch structure is used to extract the location frame information of all multi-rotor drones contained in the fused feature data.
[0043] S103, extracting appearance features of each frame of image according to the position detection results of all multi-rotor drones contained in each frame of image, to obtain appearance feature detection results of all multi-rotor drones contained in each frame of image.
[0044] In an embodiment of the present invention, first, based on the position detection results of all multi-rotor drones contained in each frame of image, an image of the area where each multi-rotor drone is located is cut out from each frame of image; secondly, a pre-built second convolutional neural network is used to perform feature extraction on the image of the area where each multi-rotor drone is located to obtain an appearance feature vector of each multi-rotor drone.
[0045] More specifically, the embodiment of the present invention targets any frame image with a scale of 640×640×3 contained in the video image, takes the position detection results of all multi-rotor drones contained in the frame image as the reference center, and extracts the image of the area where each multi-rotor drone is located from the frame image. During the entire interception process, the scale output of the image of the area where each multi-rotor drone is located is adjusted to 64×128 by calling the existing Reshape function.
[0046] In an embodiment of the present invention, the second convolutional neural network includes an input layer, a convolution layer, an average pooling layer and a normalization layer connected in sequence, such as Figure 4 As shown; wherein, the convolution layer is used to extract the global appearance feature data of the multi-rotor drone from the image of the area where each multi-rotor drone is located. Preferably, the present invention is provided with four convolution layers, each convolution layer is composed of an input layer, two activation layers, two two-dimensional convolution and BN (Batch Normalization) combined network layers, and any activation layer is provided with a ReLu (Rectified Linear Units) activation function; the average pooling layer is used to perform vector modulus adjustment on the global appearance feature data; the normalization layer is used to convert the adjusted global appearance feature data into an appearance feature vector.
[0047] S104, applying the appearance feature detection results of all the multi-rotor drones contained in each frame of the image to the DeepSort algorithm, and matching the motion trajectories of the plurality of multi-rotor drones frame by frame according to the position detection results of all the multi-rotor drones contained in each frame of the image.
[0048] The implementation process of the present invention includes the following:
[0049] Step 1: Initialize the parameters of the Kalman filter according to the position detection results of all multi-rotor drones contained in the first two frames of the video image by using the DeepSort algorithm to obtain the motion trajectories of the multi-rotor drones.
[0050] Step 2: Use the Kalman filter to predict the position tracking results of all multi-rotor drones contained in the previous frame image to obtain the position prediction tracking results of all multi-rotor drones contained in the current frame image.
[0051] It should be noted that the embodiment of the present invention creates a loop operation starting from step 2, and the loop operation is executed on the third frame image in the video image. At this time, the position tracking results of all multi-rotor drones contained in the second frame image in the video image are actually the position detection results of all multi-rotor drones contained in the second frame image.
[0052] Step 3: According to the position prediction tracking results and position detection results of all multi-rotor drones contained in the current frame image, the motion offset measurement results of all multi-rotor drones contained in the current frame image are calculated using the Mahalanobis distance; wherein the calculation formula for the motion offset measurement result of any multi-rotor drone contained in the current frame image is:
[0053]
[0054] in, is the motion offset measurement result of the i-th multi-rotor drone contained in the current frame image, is the position detection result of the i-th multi-rotor drone, is the position prediction and tracking result of the i-th multi-rotor UAV, is the transpose symbol, is the covariance matrix between the position detection result and the position prediction tracking result of the i-th multi-rotor drone.
[0055] Step 4: extract the appearance features of the current frame image according to the position prediction and tracking results of all the multi-rotor drones contained in the current frame image, and obtain the appearance feature prediction and tracking results of all the multi-rotor drones contained in the current frame image.
[0056] Step 5: According to the appearance feature prediction tracking results and appearance feature detection results of all multi-rotor drones contained in the current frame image, the appearance difference measurement results of all multi-rotor drones contained in the current frame image are calculated using the cosine distance; wherein, the calculation formula for the appearance difference measurement result of any multi-rotor drone contained in the current frame image is:
[0057]
[0058] in, is the appearance difference measurement result of the i-th multi-rotor drone contained in the current frame image, is the appearance feature detection result of the i-th multi-rotor drone, is the set of all appearance feature prediction tracking results that are successfully matched in all frame images arranged before the current frame image. is the real number space.
[0059] Step 6. Combine the motion offset measurement results and appearance difference measurement results of all multi-rotor UAVs contained in the current frame image, and use the existing Hungarian matching algorithm to match the position detection results of all multi-rotor UAVs contained in the current frame image with the motion trajectory formed by each multi-rotor UAV before the current frame image, and then output the position tracking results of all multi-rotor UAVs contained in the current frame image and the updated motion trajectory of each multi-rotor UAV.
[0060] Step 7: Use the position tracking results of all multi-rotor drones contained in the current frame image to update the parameters of the Kalman filter, and then return to the above step 2 to continue the prediction operation for the next frame image until the motion trajectories of all multi-rotor drones matched and output by the last frame image in the video image are obtained.
[0061] In the embodiment of the present invention, by combining the first convolutional neural network, the UAV-FPN neural network and the YOLOHead prediction network, all multi-rotor drones can be more accurately identified and detected from each frame of the video image, solving the problem of multi-rotor drones being missed when performing flight missions in the prior art from the visual field. By combining the second convolutional neural network and the DeepSort algorithm, the appearance characteristics of the multi-rotor drones are fully considered, and the tracking effect of the motion trajectories of several multi-rotor drones can be effectively improved, solving the problem of multi-rotor drones being misdetected when performing flight missions in the prior art.
[0062] In addition, an embodiment of the present invention further provides a trajectory tracking system for a multi-rotor UAV, the system comprising:
[0063] at least one processor;
[0064] at least one memory for storing at least one program;
[0065] When the at least one program is executed by the at least one processor, the at least one processor implements the trajectory tracking method of the multi-rotor drone described in any of the above embodiments.
[0066] The contents of the above method embodiments are all applicable to the present system embodiments. The functions implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are the same as those of the above method embodiments.
[0067] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the trajectory tracking system of the multi-rotor UAV, and uses various interfaces and lines to connect various parts of the trajectory tracking system of the entire multi-rotor UAV that can operate the device.
[0068] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the trajectory tracking system of the multi-rotor drone by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein: the program storage area is used to store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area is used to store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart-Media-Card, SMC), a secure digital (Secure-Digital, SD) card, a flash card (Flash-Card), at least one disk storage device, a flash memory device or other volatile solid-state storage device.
[0069] Although the description of the present application has been quite detailed and specifically describes several described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but should be regarded as providing a broad possible interpretation of these claims by reference to the attached claims, taking into account the prior art, so as to effectively cover the intended scope of the present application. In addition, the above description of the present application is based on the embodiments foreseeable by the inventor, and its purpose is to provide a useful description, and those non-substantial changes to the present application that have not yet been foreseen may still represent equivalent changes to the present application.
Claims
1. A trajectory tracking method for a multi-rotor drone, characterized in that: The method comprises: Obtain video images of several multi-rotor drones performing flight missions; Performing target recognition detection on each frame of the video image to obtain position detection results and the category of the multi-rotor drones contained in each frame of the video image; Extracting the appearance features of each frame of image according to the position detection results of all multi-rotor drones contained in each frame of image, and obtaining the appearance feature detection results of all multi-rotor drones contained in each frame of image; The appearance feature detection results of all multi-rotor drones contained in each frame of the image are applied to the DeepSort algorithm, and the motion trajectories of the multi-rotor drones are matched frame by frame according to the position detection results of all multi-rotor drones contained in each frame of the image; The performing target recognition detection on each frame of the video image to obtain the position detection results and the category of the drones of all multi-rotor drones contained in each frame of the video image includes: Using a pre-built first convolutional neural network to extract backbone features from each frame of the video image, and obtaining low-level feature data and high-level feature data contained in each frame of the video image; The UAV-FPN neural network is used to fuse the low-level feature data and high-level feature data contained in each frame of the image to obtain the fused feature data contained in each frame of the image; The YOLOHead prediction network is used to perform feature conversion on the fused feature data contained in each frame of the image to obtain the position detection results and the category of the drone to which all multi-rotor drones contained in each frame of the image belong.
2. The trajectory tracking method of a multi-rotor UAV according to claim 1, characterized in that: The first convolutional neural network includes an input processing module, a low-level feature extraction module and a high-level feature extraction module which are connected in sequence; wherein the input processing module is used to extract downsampled feature data from any frame image, the low-level feature extraction module is used to extract low-level feature data from the downsampled feature data, and the high-level feature extraction module is used to further extract high-level feature data from the low-level feature data.
3. The trajectory tracking method of a multi-rotor UAV according to claim 2, characterized in that: The input processing module includes an input layer and a Focus structure layer connected in sequence, the low-level feature extraction module includes a first BaseBlock, a second BaseBlock, a first residual convolution layer, a third BaseBlock, a second residual convolution layer, a fourth BaseBlock, a first spatial pyramid pooling layer and a third residual convolution layer connected in sequence, and the high-level feature extraction module includes a fifth BaseBlock, a second spatial pyramid pooling layer and a fourth residual convolution layer connected in sequence.
4. The trajectory tracking method of a multi-rotor UAV according to claim 3, characterized in that: Any BaseBlock includes an input layer, a two-dimensional convolutional layer, a maximum pooling layer, and an activation layer connected in sequence.
5. The trajectory tracking method of a multi-rotor UAV according to claim 1, characterized in that: The method of extracting the appearance features of each frame of image according to the position detection results of all multi-rotor drones contained in each frame of image, and obtaining the appearance feature detection results of all multi-rotor drones contained in each frame of image includes: According to the position detection results of all multi-rotor drones contained in each frame of image, an image of the area where each multi-rotor drone is located is intercepted from each frame of image; The pre-built second convolutional neural network is used to extract features from the image of the area where each multi-rotor drone is located, and the appearance feature vector of each multi-rotor drone is obtained.
6. The trajectory tracking method of a multi-rotor UAV according to claim 5, characterized in that: The second convolutional neural network includes an input layer, a convolution layer, an average pooling layer and a normalization layer connected in sequence; wherein the convolution layer is used to extract the global appearance feature data of the multi-rotor drone from the image of the area where each multi-rotor drone is located, the average pooling layer is used to perform vector modulus adjustment on the global appearance feature data, and the normalization layer is used to convert the adjusted global appearance feature data into an appearance feature vector.
7. A trajectory tracking system for a multi-rotor drone, characterized in that: The system comprises: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the trajectory tracking method for a multi-rotor drone as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-target tracking method and system suitable for embedded terminal
CN113034548A
Image tracking method and device, storage medium and electronic equipment
CN113139442A