Human-like moving object detection unmanned aerial vehicle

By designing a retractable tripod and an improved image processing algorithm on the drone, the problem of insufficient stability and real-time in the traditional drone vision detection system is solved, and high accuracy and efficient object detection and tracking are achieved.

CN120024533APending Publication Date: 2025-05-23NANTONG GRUNNI ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411888754.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the existing drone vision detection system, the commonly used retractable tripod has poor operating stability during work and is prone to shake the equipment, resulting in low accuracy of the visual equipment. At the same time, traditional algorithms have a large amount of calculation when tracking dynamic targets, and insufficient real-time performance.

Method used

A humanoid moving object detection drone was designed, using a retractable tripod, an Arduino UNO R3-based development board and servo expansion board, and the digital servo controls the retractable tripod to ensure the stability of the visual module. At the same time, an improved SIFT feature point matching algorithm and TLD tracking algorithm are adopted to improve the real-time image processing and the stability of target tracking.

Benefits of technology

It realizes high accuracy and high real-time performance of drone visual detection, reduces the interference of the retractable tripod to the vision module, and improves the efficiency and stability of object detection and tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120024533A_ABST
    Figure CN120024533A_ABST
Patent Text Reader

Abstract

The humanoid moving object detection unmanned aerial vehicle comprises an unmanned aerial vehicle body, a development board and a steering engine expansion board are vertically installed on the unmanned aerial vehicle body, and the steering engine expansion board comprises a wireless communication module; the steering engine is controlled by the development board and the steering engine expansion board and is used for controlling the folding and unfolding of a foot stool of the unmanned aerial vehicle body; the foot stools are located on the two sides of the unmanned aerial vehicle body correspondingly and connected with the steering engines correspondingly, and the steering engines control the foot stools to be folded so that the mounting mechanical arms can work conveniently, and the foot stools can be put down to support the unmanned aerial vehicle body; and the visual module is used for the unmanned aerial vehicle body to accurately detect a target and track the detected target. According to the invention, the SIFT feature point matching algorithm is improved, the processing capacity of the ground station is reduced, and the real-time performance is enhanced. And finally, the target is tracked in real time through a TLD algorithm, so that the problem that the moving target is shielded or deformed during target detection and tracking is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an unmanned aerial vehicle, in particular to a humanoid moving object detection unmanned aerial vehicle. Background Art

[0002] Machine vision detection, as a relatively new term, is now in the process of rapid development. Machine vision detection refers to the use of sensors such as cameras and video cameras, combined with machine vision algorithms to give smart devices the function of human eyes, so as to perform functions such as object recognition, detection, and measurement. At present, various hot fields including drones are inseparable from machine vision detection technology. Drone vision detection is like giving drones human eyes, which is an important part of intelligent manufacturing. So far, the development of drones for detecting various static objects or detecting and tracking dynamic objects through vision is still in its infancy in China. The visual modules of drones used in the developed products are generally installed at the bottom of the drone, and drone tripods for supporting the drone are set around the bottom of the drone. Therefore, when the visual module transmits video in real time, the drone tripod will block the lens of the visual module, causing the visual module to shoot the drone tripod into the picture, affecting the picture quality and also causing the shooting range to become smaller. The most important problem is that it increases the difficulty of image processing and target detection and tracking.

[0003] In order to solve the related problems, the traditional method is to set an extendable and shortenable telescopic frame on the drone body, set the equipment on the telescopic frame, and drive the equipment to move up and down relative to the drone body by the length and shortening of the telescopic frame. This technology can reduce the interference of the drone tripod to the drone's visual mounting equipment to a certain extent, but the telescopic frame has poor operating stability during operation and is prone to shaking the equipment, resulting in low accuracy of the visual equipment when performing tasks.

[0004] When it comes to detecting dynamic targets that need to be tracked, traditional algorithms require a lot of calculations when matching feature points of the target, which takes a long time and therefore has defects in real-time performance. Summary of the invention

[0005] Purpose of the invention: The present invention aims to solve one of the technical problems in the related art to at least a certain extent. To this end, the present invention proposes a drone with a simple structure and a stable and reliable movement process. The drone tripod does not affect the operation of the drone vision module. The vision module is first used to identify and detect the target, which enhances the real-time performance of image processing and can track the detected target.

[0006] Technical solution: The humanoid moving object detection drone described in the present invention comprises:

[0007] A drone body, wherein a development board and a steering gear expansion board are vertically mounted on the drone body, and the steering gear expansion board includes a wireless communication module;

[0008] A servo, which is controlled by the development board and the servo expansion board and is used to control the retraction and extension of the UAV body tripod;

[0009] The tripods are located on both sides of the UAV body and are connected to the steering gears respectively. The steering gears control the tripods to be retracted to facilitate the operation of the mounting mechanical arm and to be lowered to support the UAV body.

[0010] A visual module is used by the drone body to accurately detect targets and track the detected targets.

[0011] Furthermore, the development board and the servo expansion board adopt a development board and a servo expansion board based on Arduino UNO R3.

[0012] Furthermore, the wireless communication module adopts a Bluetooth or WIFI communication module.

[0013] Furthermore, the hardware part of the visual module includes a main control module and an image processing module connected to the main control module, and the main control module is also connected to an image transmission module.

[0014] Furthermore, it also includes an image processing module, which first performs grayscale processing on the image, then performs denoising processing on the image, and finally performs image binarization processing.

[0015] Furthermore, it also includes a SIFT target detection algorithm, including: scale space extreme value detection, keyword location, direction determination and key point descriptors.

[0016] Furthermore, the scale space extreme value detection includes:

[0017] A scale space is constructed to detect the blob structure in the image and the interest points called key points in the SIFT framework, and the scale space function is generated by the convolution of the variable scale Gaussian function G(x, y, σ) and the input image I(x, y), as shown in Formula 1:

[0018] L(x,y,σ)=G(x,y,σ)*I(x,y) (Formula 1)

[0019] Among them, σ is the standard deviation of the normal distribution. The larger the σ value, the blurrier (smoother) the image;

[0020] The difference of Gaussian (DoG) image G(x,y,σ) is obtained by subtracting the subsequent scale of each octave, as shown in Equation 2 and Equation 3:

[0021]

[0022] D(x,y,σ)=[D(x,y,kσ)-G(x,y,σ)]*I(x,y)

[0023] = L(x,y,kσ)-L(x,y,σ) (Formula 3)

[0024] The octaves are grouped by the convolved image, and the value of σ is doubled at the same time. The value of k is selected so that a fixed number of blurred images are generated for each octave. Then, by observing the individual image points of G(x, y, σ) and detecting extreme values ​​from them, when the value of a point is the minimum or maximum of all its surrounding neighboring points, the point is determined to be a local minimum or maximum. Each sampling point in G(x, y, σ) is compared with its eight neighboring points in the frame image and the nine neighboring points in the two previous and next frames.

[0025] Furthermore, the keyword positioning is to fit a three-dimensional quadratic function to a local sampling point to determine the interpolation position of the maximum value, using the Taylor expansion of the scale space function G(x, y, σ), with the origin at the sample point.

[0026] Furthermore, the key point descriptor includes:

[0027] First, for the case where the camera is stationary and the drone is only moving at a fixed level, the captured images are defined as I1 and I2 respectively. When processing the images, the images can be placed in a rectangular coordinate system. At this time, the two matching feature points have the same horizontal coordinate in theory. Considering the inevitable errors in practical applications, the error range is set to δ. The horizontal coordinate of the i-th feature point in I1 is loc1(i), and the horizontal coordinate of the j-th feature point in I2 is loc2(j). At this time, loc2(j) should be in the interval [loc1(i)-δ,loc1(i)+δ];

[0028] Secondly, when the camera is stationary and the drone only moves up and down without changing its horizontal position, the captured images are defined as I1 and I2 respectively, and the images are placed in a rectangular coordinate system for processing. At this time, the two matching feature points also have the same ordinate in theory. Considering factors such as errors, the determination of the ordinate detection interval is similar to that of the abscissa detection interval.

[0029] Finally, the processing of the horizontal and vertical coordinates is combined to find the point in I2 that matches the i-th feature point in I1.

[0030] Furthermore, it also includes TLD tracking algorithm:

[0031] The target to be tracked is placed in a two-dimensional space. First, the speed of the target in a certain frame is calculated. Then, the intensity change of the pixel value is calculated by comparing the data of the image change between two frames. The position information of the target is found according to the change. There is no need to predict the movement of the target. As long as some tracking points are clear, the tracker can obtain the position information of the target in the next frame.

[0032] Use the optical flow tracker to track the selected target tracking point until the tracker is at frame t+1, then return the tracker to the initial point, calculate the error FB, and obtain the best target tracking point through the weighted average of several errors;

[0033] Finally, by calculating the change in the position of the target tracking point and the change in the error, the size of the tracking window at the t+1 frame and the scaling ratio are obtained, and the data are all taken as the median.

[0034] Beneficial effects: The present invention promotes the intelligence of the field of UAV visual detection, and adopts a retractable tripod without affecting the support of the UAV, thereby reducing the adverse effects of the commonly used retractable tripod on the visual module when performing tasks;

[0035] In tasks that require detection and tracking of targets, the situation of the target area is transmitted back to the ground station in real time through the lens of the visual module, and the video information of each frame is converted into an image for processing, and then the target is determined through feature point matching detection;

[0036] The SIFT feature point matching algorithm was improved, the processing load of the ground station was reduced, and the real-time performance was enhanced. Finally, the target was tracked in real time through the TLD algorithm, which solved the problems of moving targets being blocked or deformed during target detection and tracking. After the real-time performance of feature matching was enhanced, the tracking effect was better and more stable. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a diagram showing the relationship between the PWM waveform and the steering gear working angle of the present invention;

[0038] Figure 2 This is a diagram showing the relationship between the PWM waveform and the steering gear position of the present invention;

[0039] Figure 3 The structure diagram of the retractable tripod of the UAV of the present invention;

[0040] Figure 4 This is a diagram showing the position of the tripod of the present invention relative to the drone body when the tripod is folded;

[0041] Figure 5 This is a diagram showing the position of the tripod of the present invention when it is lowered and the drone body;

[0042] Figure 6It is the control flow chart of the retractable tripod of the present invention;

[0043] Figure 7 It is the original image in the embodiment of the present invention;

[0044] Figure 8 It is the grayscale processing of the image in the embodiment of the present invention;

[0045] Figure 9 It is the denoising processing of the image in the embodiment of the present invention;

[0046] Figure 10 It is the binarization processing of the image in the embodiment of the present invention;

[0047] Figure 11 It is the process diagram of the SIFT algorithm feature construction of the present invention;

[0048] Figure 12 It is the detection flow chart of the vision module of the present invention;

[0049] Figure 13 It is the TLD target tracking algorithm of the present invention;

[0050] Figure 14 It is the structure diagram of the cascade classifier of the present invention;

[0051] Figure 15 It is the structure diagram of the learning process of the present invention;

[0052] Figure 16 It is the processing flow chart of the TLD algorithm of the present invention. Detailed implementation manners

[0053] The following details the implementation manners of the present invention. The examples of the implementation manners are shown in the accompanying drawings. The implementation manners described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation of the present invention.

[0054] The present invention discloses a humanoid moving object detection drone, and the drone includes:

[0055] A drone body, on which an Arduino UNO R3 development board and a servo extension board are vertically installed. The servo extension board includes a Bluetooth / WiFi module, and in the embodiment of the present invention, the Bluetooth module is preferably adopted;

[0056] Two digital servos, which are controlled by the development board and the extension board and are used to control the retraction and extension of the tripod of the body;

[0057] Two retractable parallel bar tripods, which are respectively located on the left and right sides of the body and are respectively connected to two digital servos. The digital servos control the tripods to retract for convenient mounting of the robotic arm to work and to extend to support the body;

[0058] A visual module is used by the drone to accurately detect targets and track the detected targets.

[0059] 1. Control panel and servo

[0060] The present invention selects Arduino UNO R3 development board and servo expansion board to control digital servos. Arduino is mainly divided into two parts, one is the hardware circuit board, and the other is Arduino IDE. We only need to use Arduino IDE for programming, and then load the program into the development board through the data cable, and then we can work through its own circuit. ArduinoIDE software is developed from Processing IDE, and Arduino language is developed from wiring language, and has an open source function library.

[0061] A digital servo is a position (angle) servo driver, suitable for control systems that require angles to be constantly changed and maintained. Compared with analog servos on the market, the digital servo has a microprocessor and a crystal oscillator, a smaller non-responsive zone, a higher resolution, and a greater fixed force. Common servo control angles are 0-90°, 0-180°, and 0-360°. The present invention uses a digital servo with a control angle of 0-90°. The minimum pulse corresponds to a pulse width of 1ms, i.e., one twentieth of a cycle, corresponding to a working angle of 0° for the digital servo. The maximum pulse corresponds to a pulse width of 2ms and one tenth of a cycle, corresponding to a working angle of 90° for the digital servo. The pins are three compatible wires arranged in the same way, namely, GND (brown), VCC (red), and PWM (yellow). The control method is PWM timing. The relationship between the PWM waveform and the servo working angle is as follows: Figure 1 As shown, the relationship between the PWM waveform and the servo position is as follows Figure 2 shown.

[0062] Since most drones on the market use retractable tripods when mounting electronic remote sensing equipment, and when the electronic remote sensing equipment is installed on the retractable tripod, the movement range is large, which has a significant impact on the electronic remote sensing equipment. The present invention uses the digital servo to control the left and right tripods of the drone. Instead of using a telescopic movement method, the angle between the tripod and the drone body is changed. When the mounted equipment performs an aerial mission, the tripod is controlled by the digital servo alone so that it is 180° with the lower plate of the drone body, and is not connected to the lens of the visual module and the mounting mechanical arm. When landing, the angle is adjusted to about 90° to 100° to support the drone. In this way, while facilitating the operation of the mounting equipment and not affecting the landing support of the drone, it will not cause the visual module lens to shake, and there will be no problem of the tripod blocking the lens or affecting the operation of the mounting mechanical arm during the shooting process. It has stronger practical applicability. The specific retractable tripod body, folding and putting down, the plan views are as follows: Figure 3 , Figure 4 and Figure 5 The control flow chart of the retractable tripod controlled by the digital servo is as follows: Figure 6 shown.

[0063] 2.Bluetooth module

[0064] The communication module gives priority to the BT06 Bluetooth module, which complies with the V3.0 Bluetooth specification, and the Bluetooth module uses the SPP Bluetooth serial port protocol. It uses the UART interface to achieve wireless communication. The characteristics are that it occupies a relatively small space, the material used makes it have low loss during communication, and the cost of the material is not high. It is convenient to send and receive data signals in the hardware module, the signal is sensitive, the function is powerful, and fewer components are used. The data wireless transmission distance of the BT06 Bluetooth module is long, and the data transmission between other hardware modules is convenient. There is no complicated serial port line, which reduces unnecessary serial port connections.

[0065] 3. Vision module

[0066] (1) Hardware

[0067] The visual module system mainly uses STM32 as the main control module. If visual computing is performed, STM32 cannot process images. After design, an image processing module is added to the module part. After many comparisons, we choose to use the Raspberry Pi 4B embedded microcomputer as the image processing module, which specializes in processing the computing part of the visual module.

[0068] The Raspberry Pi 4B microcomputer is based on four powerful ARM A72 chips and uses the new BCM2711 processor. Bluetooth uses the 5.0 specification, which consumes less energy. The USB-C interface is easy to use and has a strong power supply capacity. It has better performance, higher memory specifications, stronger processing power, higher network performance, the Ethernet interface has changed from 330Mbps to Gigabit, and the wireless network supports WiFi, which remains unchanged. The four interfaces are 3.0 standards, and the external storage and read and write capabilities are very strong. The video output interface has become two MICRO HDMI interfaces, which can output two 2K ​​screens at the same time, which is more suitable for work tasks. The 1.2GHz CPU main frequency can meet our computing needs for image processing. Through the connection with the communication part of the Bluetooth Ethernet and other systems, we can better handle related tasks, especially when operating at a distance.

[0069] At the power interface on the drone, a voltage regulator module is connected to the power interface of the Raspberry Pi to ensure its stable operation. There are two connecting channels between the Raspberry Pi and the STM32 microcontroller. One channel is the signal sent from the Raspberry Pi through the conversion module of the USB interface, and finally the data information is transmitted to the STM32 unit; the other connecting channel is the opposite of the previous image, transmitting signal data from the STM32 unit, passing through the USB conversion module to the Raspberry Pi. The Raspberry Pi is equipped with a reading function to save the transmitted flight data.

[0070] The image transmission module uses EWRF 7082VR, which supports SBUS parameter adjustment. The interface channel group, channel, power, frequency, voltage and other parameters can be transmitted back to the display screen. Its output power is stable and the frequency band is clean. It can achieve zero interference power-on and channel switching without interfering with other users. The processing core it uses uses the latest algorithm and is very accurate in function use. It is also equipped with various sensors to achieve precise positioning. Positioning signals can be transmitted in complex terrain and high-speed flight, and the height from the ground can reach 20 meters. In addition, the system's visual sensor uses a USB interface to transmit image information. By configuring the sensor, it can expand the realization of functions such as speed, distance, and depth. This design mainly reads the image information collected by its camera. We perform calculations by reading the quaternion information output by the image transmission module, obtain the position information of the drone from the STM32 unit on the drone, and then analyze the position difference between the drone and the target through image capture. The changes in the quaternion of the image transmission module and the quaternion of the drone body should be consistent. Because the image transmission module is fixed just below the center of the drone body, image processing calculations can be performed in the Raspberry Pi embedded microcomputer, reducing the calculation of the STM32 unit.

[0071] The TLD algorithm used requires the use of a Median-Flow optical flow tracker. When the tracker uses a tracking algorithm to track a target, the tracking path is similar for forward tracking or reverse tracking.

[0072] (2) Image processing

[0073] When the visual module transmits the image back to the ground station, the image needs to be processed. The main task is to process the received data information, improve the image quality and special feature information, and then encode and compress the data according to the image transformation, and finally transmit it to the next processing module. In the image processing part, the image is first grayed, then denoised, and finally binarized. Here, Raspberry Pi is used for wireless connection. After successfully calling the camera, OpenCV is used for image processing. The original image is as follows Figure 7 shown.

[0074] 1) Image grayscale processing

[0075] The grayscale value of an image is obtained mainly by calculating and processing the image. There are three processing methods, namely the maximum value method, the average value method and the weighted average method. The weighted average method is selected by calculation and comparison. This method compares the brightness component value of the color part of the image. Through weighted average processing, a grayscale value of the selected image can be obtained. The value obtained by multiple weighted average calculations is more reasonable and accurate. The image after grayscale processing is as followsFigure 8 shown.

[0076] 2) Image denoising

[0077] There are two ways to denoise an image, namely median filtering and mean filtering. Median filtering is a method that can suppress noise after smoothing. It mainly arranges the signal in order and then performs effective statistics. Its characteristics are that it is easy to use and does not require long-term calculations to get the desired effect. The mean filtering algorithm is also called a linear filtering algorithm. It can average the pixels of several images. The calculation method used can also be called the neighborhood average method. After calculation, an average value can be obtained, which is used as the grayscale value of the image pixel. The former has higher real-time performance and is more suitable for scenes that require real-time feedback. Therefore, the median filtering method is used in image denoising. The steps are as follows: first, determine the center point in a rectangular area or other shaped area, arrange the pixel values ​​calculated by the feature rectangular box in order, and select the middle value as the new grayscale value after statistics, which is used as the center of the image pixel. Then, a template is obtained by median filtering, and the pixel values ​​calculated by the feature rectangular box are placed in the filter template to obtain a related sequence. By observing the fluctuation curve of the sequence, it rises or falls. When the target is moving, the image can be made smoother by median filtering. After the image is smoothed, we can record the original image as f(x,y), and the processed image as g(x,y). The template can move smoothly on the image. Generally, the selected feature rectangle is rectangular or circular. After the median filter sorts the pixels arranged in the template, the median method is used to obtain the required median value to replace the original median value. The image after denoising is as follows Figure 9 shown.

[0078] 3) Image binarization

[0079] After the image is grayed and denoised, it is processed into a binary image using the binarization method. The specific steps are as follows: first determine a gray value, then compare the pixels and thresholds of each image. If the compared value is greater than this gray value, the gray value of the image is recorded as 255. If the compared value is less than this gray value, the gray value of the image is recorded as 0. After the binarization process, the image can clearly show the effect of black and white. The image after binarization is as follows: Figure 10 shown.

[0080] (3) SIFT object detection algorithm

[0081] The present invention selects to use SIFT target detection algorithm after processing the image, because it is invariant not only in terms of rotation, scaling and brightness change, but also in terms of angle change, affine transformation and noise. This method does not require the ratio of the number of feature points to the number of valid points. If there are not many target feature points, the optimized SIFT matching algorithm can also meet the real-time requirements. In addition, it is easy to combine with other forms of feature vectors. The SIFT algorithm consists of four parts. The SIFT feature construction process is as follows: Figure 11 shown.

[0082] 1) Scale space extreme value detection

[0083] By constructing a scale space, the blob structure in the image and the points of interest called key points in the SIFT framework are detected. The scale space function is generated by the convolution of the variable scale Gaussian function G(x,y,σ) and the input image I(x,y), as shown in Formula 1:

[0084] L(x,y,σ)=G(x,y,σ)*I(x,y) (Formula 1)

[0085] Here, σ is the standard deviation of the normal distribution. The larger the σ value, the blurrier (smoother) the image.

[0086] The difference of Gaussian (DoG) image G(x,y,σ) is obtained by subtracting the subsequent scale of each octave, as shown in Equation 2 and Equation 3:

[0087]

[0088] D(x,y,σ)=[D(x,y,kσ)-G(x,y,σ)]*I(x,y)

[0089] = L(x,y,kσ)-L(x,y,σ) (Formula 3)

[0090] Octaves are grouped by the convolved image, the value of σ is doubled, and the value of k is chosen so that a fixed number of blurred images are generated per octave. Then, the extreme values ​​are detected by observing each image point of G(x,y,σ). A point is determined to be a local minimum or maximum when its value is the minimum or maximum of all its surrounding neighbors. Each sample point in G(x,y,σ) is compared with its eight neighbors in the frame it is in and its nine neighbors in the two frames before and after.

[0091] 2) Keyword Targeting

[0092] In this step, keywords will be filtered to a great extent, only stable keywords will be retained, and unstable keywords will be eliminated on a large scale. Brown developed a method to fit a three-dimensional quadratic function to local sampling points to determine the interpolation position of the maximum value. Using the Taylor expansion of the scale space function G(x,y,σ), the origin is at the sample point. The position of the extreme value X' is determined by Formula 4 and Formula 5:

[0093]

[0094] Formula 6 is very helpful in eliminating unstable extreme values ​​with low contrast:

[0095]

[0096] The present invention discards all extreme values ​​where |D(X')| is less than 0.03. The DoG operator has a strong response near the edge of the image, which can cause instability in the key point. The peak in the poorly defined DoG function has a small principal curvature in the vertical direction, but a large principal curvature on the edge. The 2×2 Hessian matrix H can calculate the principal curvature at the position and scale of the key point. H is given by the following formula 7:

[0097]

[0098] The eigenvalues ​​of H are proportional to the principal curvatures of D, but the eigenvalues ​​are not explicitly calculated. Instead, the trace and determinant of H are used to remove key points whose principal curvature ratio is greater than a threshold.

[0099] 3) Direction determination In this step, each key point is assigned one or more directions based on the local image gradient direction. The gradient and direction are calculated by using formula 8 and formula 9:

[0100]

[0101] A set of feature points will be generated after removing the key points with low contrast and edges. Each key point has three parameters: position, gradient and orientation.

[0102] 4) Keypoint Descriptor

[0103] At this stage, a different descriptor is calculated for each keypoint. Once the orientation of a keypoint is selected, the feature descriptor is calculated as a set of orientation histograms of a 16×16 pixel neighborhood. 8 bins are included in each histogram, and a 4×4 histogram array is included in each descriptor. Therefore, it can be concluded that the feature vector contained in each keypoint is 4×4×8=128 dimensions.

[0104] Target recognition refers to matching the target template with the current camera view. The matching between two images is done by comparing each extremum based on the relevant descriptors. First, the SIFT features of the two images are extracted using the SIFT algorithm discussed above. Given the feature point fp11 in image 1, its closest point fp21, the second closest point fp22 and their Euclidean distances ed1 and ed2 are calculated from the feature points in image 2. If the ratio of ed1 / ed2 is less than the threshold, it can be considered that the feature points fp11 and fp21 match. Then the matching score between the two images is determined by confirming the number of matching points, so that the target can be identified. The SIFT algorithm maintains a certain stability and robustness in terms of image rotation, brightness transformation, perspective transformation, scale scaling, affine transformation, etc. The SIFT algorithm has good anti-interference performance.

[0105] Due to the use of SIFT algorithm, most existing programs automatically detect feature points, and the number varies. When detecting smooth targets, there are fewer feature points. Therefore, to address this defect, the number of feature points is defined based on the original program. The procedure is as follows:

[0106] int num_corners =; / / minimum value of feature points

[0107] int max_corners =; / / Maximum value of feature points

[0108] In the original feature matching algorithm, the original target image is defined as I1 and the captured approximate target image is defined as I2. When the SIFT algorithm is used to detect I1 and I2, a large number of feature points will be detected, and each feature point can be represented by a 128-dimensional feature vector, as shown in Formula 10 and Formula 11:

[0109] SIFT(I1)=[N1,des1,loc1] (Formula 10)

[0110] SIFT(I2)=[N2,des2,loc2] (Formula 11)

[0111] Where N1 and N2 represent the number of feature points in I1 and I2 images, des1 and des2 represent the descriptors of feature points in I1 and I2 images, and loc1 and loc2 represent the position information and gradient information of feature points in I1 and I2. In the traditional SIFT feature point matching algorithm, it is necessary to determine whether a feature point des1 (i, 1:128) in I1 has a matching feature point in I2. This requires the feature point in I1 to calculate the Euclidean distance D (i, 1:N2) with all N2 feature points des2 (1:N2, 1:128) in I2, as shown in Formula 12:

[0112] D(i,1:N2)=des(i,1:128)*des2(1:N2,1:128)'(Formula 12)

[0113] Then sort D(i,1:N2) from small to large to get a new D(i,1:N2), record the serial number j of the original D(i,j) currently ranked in D(i,1), and set a threshold d. When D(i,1)<d×D(i,2), it is considered that the feature point des1(i,1:128) in I1 and the feature point des2(j,1:128) in I2 are a pair of matching feature points.

[0114] In the original feature matching algorithm, to determine whether the i-th feature point in I1 matches the j-th feature point in I2, it is necessary to calculate the Euclidean distance between the i-th feature point in I1 and all N2 feature points in I2 and sort them. If all N1 feature points in I1 and N2 feature points in I2 are to be matched, the steps need to be repeated N1 times. The calculation amount is extremely large and the speed is slow. In order to improve the calculation efficiency and reduce the processing amount, the following improvement methods are provided:

[0115] First, for the case where the camera is stationary and the drone is only moving at a fixed level, the captured images are defined as I1 and I2 respectively. When processing the images, the images can be placed in a rectangular coordinate system. At this time, the two matching feature points have the same horizontal coordinate in theory. Considering the inevitable errors in the actual application process, the error range is set as δ. The horizontal coordinate of the i-th feature point in I1 is loc1(i), and the horizontal coordinate of the j-th feature point in I2 is loc2(j). At this time, loc2(j) should be in the interval [loc1(i)-δ,loc1(i)+δ]. In this way, when the i-th feature point in I1 uses the algorithm to match the feature points in I2, it is not necessary to detect the feature points of all rows. It only needs to be matched in the interval of the horizontal coordinate [loc1(i)-δ,loc1(i)+δ], which greatly reduces the matching range and reduces the processing volume.

[0116] Secondly, when the camera is stationary and the drone only moves up and down without changing its horizontal position, the captured images are defined as I1 and I2 respectively, and the images are placed in a rectangular coordinate system for processing. At this time, the two matching feature points theoretically have the same ordinate. Considering factors such as errors, the determination of the ordinate detection interval and the abscissa detection interval is similar.

[0117] Finally, by combining the processing of the horizontal and vertical coordinates, taking the two 640×640 images as an example, we look for points in I2 that match the i-th feature point in I1. Assuming the error δ is 10, the horizontal coordinate detection range is changed from [1,640] to a range with an interval length of only 20, which is reduced by 32 times. If the vertical coordinate detection range is reduced by 4 times, the detection and matching range is reduced by 128 times, which greatly improves the efficiency of the algorithm, reduces the processing volume, and greatly reduces the time consumed for detection and matching.

[0118] The visual module detection flow chart using the improved SIFT algorithm is as follows: Figure 12 shown.

[0119] (4) TLD tracking algorithm

[0120] After using the improved SIFT algorithm to determine the target, the TLD target tracking algorithm can be used to track the previously detected target. The components of the TLD algorithm used include tracking, detection, and learning modules, which can track a selected target for a long time. The relationship between the components is as follows: Figure 13 shown.

[0121] When the tracking module uses the tracking algorithm to track the target, the forward tracking or reverse tracking, the tracking path should be similar. The error of this tracker can be set as FB, then the time t is calculated from the initial position x(t), and after a time process of t+p, the position changes to x(t+p), and the tracker returns from the reached position to the initial position x(t) again, after a time t, a certain error time is generated in the middle, which is FB. The specific tracking method is to place the tracked target in a two-dimensional space, first calculate the speed of the tracked target in a certain frame, and then compare the data of the image change between the two frames to calculate the intensity change of the pixel value, and find the position information of the tracked target according to the change. There is no need to predict the movement of the target, as long as some tracking points are clear, the tracker can obtain the position information of the tracked target in the next frame. Use the optical flow tracker to track the selected target tracking point until the tracker is in the t+1 frame, then let the tracker return to the initial point, calculate the error FB, and obtain the best target tracking point through the weighted average of several errors. Finally, by calculating the change in the position of the target tracking point and the change in the error, the size of the tracking window at the t+1 frame and the scaling ratio are obtained, and the data are all taken as the median.

[0122] The detection module is composed of a combination of classifiers. It detects and classifies the image feature values ​​of the moving target in the video, selects the area where the tracking target is located, and deduces the position where the tracking target will appear in the next frame. The similarity between the previous and next images is determined by formula 13:

[0123] S(P i,P j )=0.5[NCC(P i ,P j )+1](Formula 13)

[0124] NCC refers to the normalized correlation coefficient. If it is not detected in the video, it may be detected again until the moving target is successfully tracked in the next frame. The TLD algorithm uses a cascade classifier to classify the window of the current frame position of the selected tracking target, such as Figure 14 shown.

[0125] The learning module can work together with the detection module to improve the accuracy and effectiveness of tracking targets. The detection results and tracking results when tracking targets will have a certain impact on the learning module. Conversely, the learning module will also affect the detection and tracking effects of the tracking target. The learning process is as follows: Figure 15As shown. The learning method used by the TLD algorithm is PN learning (P-NLearning), which is an algorithm used to monitor machine learning. And it can correct sample detection and classification errors, respectively called P experts and N experts. P experts can change negative samples into positive samples, while N experts can change positive samples into negative samples. By checking windows of different sizes in the module, feature rectangular frames are formed one by one. A target forms a feature rectangular frame, and the position of the rectangular frame can be called an image element. The generation of samples is because the image element is transmitted into the sample set of the machine, so it is necessary to use a classifier to classify it first, and then put it into the sample set to compare the grayscale values ​​of the two images by variance calculation. If the variance is less than the initial time, the sample is recorded as a negative sample, so that more than half of the samples can be excluded. The specific role of the P expert is to make the track of the tracked target at each position continuous and uninterrupted, and its main purpose is to perform structural processing on the target information. When receiving the information of the tracked target, it predicts the position of the target that is about to arrive at the t+1 frame through the tracker. When the position window at this time is classified as a negative sample by the classifier, the P expert can change the negative position to a positive position. The role of the N expert is to initialize the tracking target at the initial position and the position of other frames. Its main role is to perform structured processing on the target information. Compared with the positive sample transmitted by the P expert, the position information of the tracking target is obtained through the detector, and the information of this position can be treated as the tracking result of the TLD algorithm. When the tracking target moves to the position of the t+1 frame, that is, the window frame has been selected, ten windows are selected from the selected windows, and the ratio of the area of ​​the intersection of the windows to the total area must not exceed 0.7. Then, each window is transformed in size, position and direction, and 10% is taken as the standard. After the transformation, twenty image elements and two hundred positive samples can be generated. Then some windows with a long distance are selected, and the ratio of the area of ​​the intersection to the total area is required to be less than 0.2. In this way, the corresponding negative samples can be obtained. After being processed by the classifier, they are put into the sample set to update the parameter value of the classifier in real time.

[0126] The comprehensive module will detect the three large classifiers, find the corresponding windows, and then compare and analyze the similarities between the windows to form clusters. Repeating the comparison task will find the differences, and finally output them through judgment and selection. Secondly, the tracking module obtains the similarity between the two windows and continues to save the processing. Finally, the two saved data are compared and judged in combination with relevant conditions to output the selected target position. The selection of the tracking target frame is completed by judging the information obtained after the detection module and the comprehensive module data processing. This module can be used to solve problems such as occlusion or deformation of moving targets during target tracking. After improving the feature matching algorithm, the tracking effect is more stable when tracking the features of the target.

[0127] The specific steps are as follows:

[0128] First, the target position of the video is initialized, and then the tracked target is calibrated. After the tracked target position data is transmitted to the tracker for tracking, the tracking effect will be fed back to the learning module. Since there are characteristic values ​​of target tracking in the current frame of the initial video sequence, corresponding samples will be generated in the learning module. Then, the corresponding judgment criteria are generated by selecting the training detector of the sample, and the tracker tracks the target frame by frame. After obtaining the tracking results, they are output in the comprehensive module to display the target's position information. Finally, the characteristic values ​​of the tracked target are detected in real time through the learning module, and then acted on the tracking module to track and detect the target. The tracked target in the learning module will be dynamically updated in real time.

[0129] There are four situations for whether the tracking detection is successful:

[0130] ①Successful tracking and successful detection

[0131] Tracking success means that the feature rectangle is reflected on the screen when the target is tracked, and detection success means that the detection module detects the feature rectangle of the target. If two feature rectangles are passed, the similarity of the two feature rectangles needs to be calculated first, and then the judgment process continues. If there are three or more feature rectangles in the window, the similarity between the two feature rectangles needs to be calculated separately, and then the classification process is performed according to the threshold.

[0132] ②Tracking success detection failure

[0133] If the tracking is successful, the feature rectangle can be displayed. If the detection fails, the feature rectangle cannot be displayed after the detector classification. The selected feature rectangle can be directly used as the parameter of the final positioning of the tracking target for output processing.

[0134] ③Tracking failure detection success

[0135] If the tracking fails, the tracker cannot select the feature rectangle, but the detection module can continue to display. If the detection is successful, it means that the detected feature rectangle is displayed on the screen, and clustering processing is required first, and then judgment and output are made according to relevant conditions.

[0136] ④Tracking failure Detection failure

[0137] If both the tracker and the detector fail at the same time, the detection is considered useless and the above detection and tracking are performed again.

[0138] To summarize, when the tracking module and the detection module are detecting, if the tracking module fails, the feature rectangle may still be displayed. However, if multiple results are generated after the rectangular box is clustered, the final detection result is still a failure. When the tracking module detects successfully, the feature rectangular box can be displayed on the screen regardless of whether the detection module detects successfully. On the premise of successful detection, the similarity between the rectangular boxes determines the credibility of the results. The parameters of the tracking module can be modified according to the results obtained after clustering. If the tracked feature rectangle has more than one centroid, the detection is also considered a failure and needs to be re-detected. Repeat the above operations until the tracking detection is successful. The effect diagram after successful tracking detection is as follows. Figure 16 shown.

[0139] The present invention promotes the intelligence of the field of UAV visual detection. It adopts a retractable tripod without affecting the support of the UAV, thereby reducing the adverse effects of the commonly used retractable tripod on the visual module when performing tasks. In the work tasks that require detection and tracking of targets, the situation of the target area is transmitted back to the ground station in real time through the lens of the visual module, and the video information of each frame is converted into an image for processing, and then the target is determined by feature point matching detection. The SIFT feature point matching algorithm is improved, the processing volume of the ground station is reduced, and the real-time performance is enhanced. Finally, the target is tracked in real time through the TLD algorithm, which solves the problems of the moving target being blocked or deformed during target detection and tracking. After the real-time performance of feature matching is enhanced, the tracking effect is better and more stable.

[0140] The above description is only a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A humanoid moving object detection drone, characterized by: include: A drone body, wherein a development board and a steering gear expansion board are vertically mounted on the drone body, and the steering gear expansion board includes a wireless communication module; A servo, which is controlled by the development board and the servo expansion board and is used to control the retraction and extension of the UAV body tripod; The tripods are located on both sides of the UAV body and are connected to the steering gears respectively. The steering gears control the tripods to be retracted to facilitate the operation of the mounting mechanical arm and to be lowered to support the UAV body. A visual module is used by the drone body to accurately detect targets and track the detected targets.

2. The humanoid moving object detection drone according to claim 1, characterized in that: The development board and the servo expansion board adopt a development board servo expansion board based on Arduino UNO R3.

3. The humanoid moving object detection drone according to claim 1, characterized in that: The wireless communication module adopts a Bluetooth or WIFI communication module.

4. The humanoid moving object detection drone according to claim 1, characterized in that: The hardware part of the visual module includes a main control module and an image processing module connected to the main control module, and the main control module is also connected to an image transmission module.

5. The humanoid moving object detection drone according to claim 1, characterized in that: It also includes an image processing module, which first performs grayscale processing on the image, then performs denoising processing on the image, and finally performs image binarization processing.

6. The humanoid moving object detection drone according to claim 1, characterized in that: It also includes the SIFT target detection algorithm, including: scale space extremum detection, keyword localization, direction determination, and key point descriptors.

7. The humanoid moving object detection drone according to claim 6, characterized in that: The scale space extreme value detection includes: A scale space is constructed to detect the blob structure in the image and the interest points called key points in the SIFT framework, and the scale space function is generated by the convolution of the variable scale Gaussian function G(x, y, σ) and the input image I(x, y), as shown in Formula 1: L(x,y,σ)=G(x,y,σ)*I(x,y) (Formula 1) where σ is the standard deviation of the normal distribution. The larger the σ value, the blurrier (smoother) the image. The difference of Gaussian (DoG) image G(x,y,σ) is obtained by subtracting the subsequent scale of each octave, as shown in Equation 2 and Equation 3: The octaves are grouped by the convolved image, and the value of σ is doubled at the same time. The value of k is selected so that a fixed number of blurred images are generated for each octave. Then, by observing the individual image points of G(x, y, σ) and detecting extreme values ​​from them, when the value of a point is the minimum or maximum of all its surrounding neighboring points, the point is determined to be a local minimum or maximum. Each sampling point in G(x, y, σ) is compared with its eight neighboring points in the frame image and the nine neighboring points in the two previous and next frames.

8. The humanoid moving object detection drone according to claim 6, characterized in that: The keyword positioning is to fit a three-dimensional quadratic function to a local sampling point to determine the interpolation position of the maximum value, using the Taylor expansion of the scale space function G(x, y, σ), with the origin at the sample point.

9. The humanoid moving object detection drone according to claim 6, characterized in that: The key point descriptor includes: First, for the case where the camera is stationary and the drone is only moving at a fixed level, the captured images are defined as I1 and I2 respectively. When processing the images, the images can be placed in a rectangular coordinate system. At this time, the two matching feature points have the same horizontal coordinate in theory. Considering the inevitable errors in practical applications, the error range is set to δ. The horizontal coordinate of the i-th feature point in I1 is loc1(i), and the horizontal coordinate of the j-th feature point in I2 is loc2(j). At this time, loc2(j) should be in the interval [loc1(i)-δ,loc1(i)+δ]; Secondly, when the camera is stationary and the drone only moves up and down without changing its horizontal position, the captured images are defined as I1 and I2 respectively, and the images are placed in a rectangular coordinate system for processing. At this time, the two matching feature points also have the same ordinate in theory. Considering factors such as errors, the determination of the ordinate detection interval is similar to that of the abscissa detection interval. Finally, the processing of the horizontal and vertical coordinates is combined to find the point in I2 that matches the i-th feature point in I1.

10. The humanoid moving object detection drone according to claim 1, characterized in that: Also includes TLD tracking algorithm: The target to be tracked is placed in a two-dimensional space. First, the speed of the target in a certain frame is calculated. Then, the intensity change of the pixel value is calculated by comparing the data of the image change between two frames. The position information of the target is found according to the change. There is no need to predict the movement of the target. As long as some tracking points are clear, the tracker can obtain the position information of the target in the next frame. Use the optical flow tracker to track the selected target tracking point until the tracker is at frame t+1, then return the tracker to the initial point, calculate the error FB, and obtain the best target tracking point through the weighted average of several errors; Finally, by calculating the change in the position of the target tracking point and the change in the error, the size of the tracking window at the t+1 frame and the scaling ratio are obtained, and the data are all taken as the median.