Unmanned aerial vehicle visual tracking method and system based on neural network
Through the multi-scale spatiotemporal feature extraction in drone vision tracking technology and the feature fusion of neural network models, real-time motion prediction parameters are generated, which solves the problem of low accuracy in complex environments in the existing technology, and achieves efficient and flexible drone target tracking.
Patent Information
- Application Number
- CN202510314025.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing drone vision tracking technology has low accuracy in complex environments, which is difficult to meet the needs of precise target tracking, and is sensitive to light changes and target posture changes, so it cannot effectively handle complex three-dimensional spatial motion and rapidly changing target trajectory.
The drone is equipped with a camera device to collect dynamic image sequences of target objects in real time, perform multi-scale spatiotemporal feature extraction, obtain global motion characteristics and local texture features, and input them into the pre-trained neural network model for feature fusion, generate real-time motion prediction parameters, adjust the drone's flight attitude, and realize dynamic tracking.
It improves the accuracy and reliability of motion prediction, realizes highly flexible tracking of targets by drones in complex environments, reduces the risk of tracking loss, and improves the overall success rate of visual tracking tasks.
Smart Images

Figure CN120219997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a neural network-based unmanned aerial vehicle visual tracking method and system. Background Art
[0002] With the rapid development of science and technology, drones have been widely used in many fields. Among them, drone visual tracking technology has become one of the research hotspots, playing an important role in security monitoring, logistics distribution, agricultural plant protection, film and television shooting and other scenarios.
[0003] In the early days of drone tracking technology, it mainly relied on pre-set fixed routes or simple sensor feedback. For example, some traditional methods use the GPS positioning system to obtain the approximate location information of the target, and the drone flies according to the set program based on this information for tracking. However, this method has obvious limitations. The GPS positioning accuracy is limited, and the signal is easily interfered in some complex environments (such as urban areas with tall buildings, dense forests, etc.), resulting in a significant decrease in tracking accuracy, making it difficult to meet the demand for accurate target tracking.
[0004] Later, vision-based drone tracking technology gradually emerged. Some existing technologies use simple image matching algorithms to determine the target position by comparing the similarity between the current image and the initial template image, and then adjust the drone flight. However, this method is very sensitive to factors such as changes in lighting and changes in target posture. When the target appearance changes slightly or the ambient lighting conditions change, tracking failure is likely to occur. Moreover, this type of method can only handle relatively simple target motion patterns, and cannot effectively track complex three-dimensional spatial motion and rapidly changing target trajectories.
[0005] In addition, some tracking algorithms based on single features, such as only using the color features or edge features of the target for tracking. These algorithms are too one-sided in feature extraction and cannot fully reflect the motion characteristics of the target. For example, when relying only on color features, if the color of the target is similar to the surrounding environment, or the color changes under different lighting, the tracking will be biased or even lose the target; using only edge features will make it difficult to deal with situations such as partial occlusion of the target, resulting in unstable tracking. Summary of the invention
[0006] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a method for visual tracking of a drone based on a neural network, the method comprising: The camera device carried by the drone collects a dynamic image sequence of the target object in real time, wherein the dynamic image sequence includes the motion trajectory information of the target object at consecutive moments; Perform multi-scale spatio-temporal feature extraction on the dynamic image sequence to obtain the global motion feature and local texture feature of the target image sequence, where the global motion feature is used to characterize the displacement change of the target object in three-dimensional space, and the local texture feature is used to characterize the spatio-temporal continuity of the surface details of the target object; Input the global motion feature and the local texture feature into a pre-trained neural network model for feature fusion to generate the real-time motion prediction parameters of the target object; Adjust the flight attitude parameters of the drone based on the real-time motion prediction parameters to generate corresponding dynamic tracking instructions, where the dynamic tracking instructions are used to control the drone to maintain a preset relative motion relationship with the target object; Drive the drone to perform a visual tracking task according to the dynamic tracking instructions, and at the same time, update the dynamic image sequence in real time and feedback it to the multi-scale spatio-temporal feature extraction step to form a closed-loop tracking control.
[0007] On the other hand, an embodiment of the present invention further provides a drone vision tracking system based on a neural network, including a processor and a machine-readable storage medium, where the machine-readable storage medium is connected to the processor, and the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0008] Based on the above aspects, in the embodiments of the present application, a drone is equipped with a camera device to collect a dynamic image sequence containing the motion trajectory information of the target object in real time. On this basis, multi-scale spatio-temporal feature extraction is performed to obtain both global motion features and local texture features simultaneously. The global motion features can accurately capture the displacement changes of the target object in three-dimensional space, enabling the drone to understand the motion trend of the target from a macroscopic level; while the local texture features capture the spatio-temporal continuity of the surface details of the target object, supplementing the information at the microscopic level. Next, the global motion features and local texture features are input into a pre-trained neural network model for feature fusion and real-time motion prediction parameters are generated. Utilizing the powerful learning and analysis capabilities of the neural network, different types of features are effectively integrated, and then the real-time motion situation of the target object is predicted, greatly improving the accuracy and reliability of motion prediction. Based on the real-time motion prediction parameters, the flight attitude parameters of the drone are adjusted and dynamic tracking instructions are generated to ensure that the drone can respond in real time according to the motion of the target and maintain a preset relative motion relationship with the target object. This enables the drone to no longer simply fly along a preset path, but can intelligently adapt to the dynamic changes of the target, achieving highly flexible tracking. Finally, according to the dynamic tracking instructions, the drone is driven to execute the visual tracking task, and the dynamic image sequence is updated in real time to form a closed-loop tracking control, which can continuously adjust and optimize according to the new state of the target, greatly improving the stability and robustness of the tracking. Even in the case where the motion trajectory of the target object is complex and changeable, the drone can still continuously and accurately track the target, reducing the risk of tracking loss and improving the overall success rate of the visual tracking task. Description of the Drawings
[0009] Figure 1 It is a schematic flowchart of the execution process of the drone vision tracking method based on neural network provided by the embodiments of the present invention.
[0010] Figure 2 It is a schematic diagram of the exemplary hardware and software components of the drone vision tracking system based on neural network provided by the embodiments of the present invention. Detailed Embodiments
[0011] The present invention will be specifically described below in conjunction with the accompanying drawings of the specification. Figure 1 It is a schematic flowchart of the drone vision tracking method based on neural network provided by an embodiment of the present invention. The drone vision tracking method based on neural network will be introduced in detail below.
[0012] Step S110: A camera device carried by a drone is used to collect a dynamic image sequence of the target object in real time, and the dynamic image sequence contains the motion trajectory information of the target object at consecutive moments.
[0013] In this embodiment, taking a rescue scenario as an example, such as in a rescue operation after an earthquake in a mountainous area, the target object is a survivor trapped in a landslide area. After the rescue drone takes off, the camera device carried by it has a set resolution and frame rate, for example, the resolution is 1920×1080 pixels and the frame rate is 30 frames per second. As the drone approaches the target area, the camera device continuously captures images containing the survivor. For example, at the initial moment t1, the survivor is captured at a certain position in the image. As time progresses to consecutive moments such as t2 and t3, since the survivor may be constantly moving, such as trying to move to a safer area or being forced to change position due to the mountain shaking again, the position changes of the survivor in the image at the above different moments constitute the motion trajectory information. The frames of images captured by the camera device are arranged in chronological order to form a dynamic image sequence. This dynamic image sequence completely records the activities of the survivor at the rescue scene, including preliminary information about their moving direction and speed.
[0014] Step S120, perform multi-scale spatio-temporal feature extraction on the dynamic image sequence to obtain the global motion feature and local texture feature of the target image sequence, where the global motion feature is used to characterize the displacement change of the target object in three-dimensional space, and the local texture feature is used to characterize the spatio-temporal continuity of the surface details of the target object.
[0015] In this embodiment, continuing with the example of mountain rescue, after the rescue drone captures the dynamic image sequence of the survivor, multi-scale spatio-temporal feature extraction can be performed. For example, the dynamic image sequence can be decomposed into multiple consecutive time segments, assuming each time segment is 1 second (including 30 frames of images). For the image frames within each time segment, frame alignment processing is performed. During this process, some feature points in each frame of the image need to be matched and corrected to eliminate the jitter effects caused by the drone's flight jitter or slight changes in the shooting angle, resulting in a stable image sequence after eliminating jitter. For example, when processing a 1-second time segment, feature point detection is performed on the images from the 1st frame to the 30th frame, and feature points such as a certain mark on the survivor's clothes or a certain fixed rock in the surrounding environment are found, and alignment correction is performed based on the position changes of these feature points in different frames.
[0016] Then, perform frame-by-frame gray normalization processing on the stable image sequence, mapping the pixel gray values of each frame of the image to a specific range, for example, between 0 and 255. Then, based on the pixel differences between adjacent frames, an optical flow field matrix is constructed. Each element in the optical flow field matrix represents the instantaneous motion vector of the target object (the survivor) in the image plane. For example, by calculating the position change of a certain pixel point in adjacent frames, the motion speed components in the x and y directions are obtained, and then the motion vector of this pixel point is determined.
[0017] Then, the multi - resolution pyramid algorithm is used to hierarchically analyze the optical flow field matrix, and the histogram of motion direction distributions at different scales is extracted. Assume that the highest - resolution layer of the Gaussian pyramid structure is the original optical flow field matrix with a resolution of 1920×1080, and each lower - resolution layer is generated by successive downsampling based on a Gaussian kernel (for example, the Gaussian kernel size is 5×5), with a downsampling ratio of 0.5. Starting from the highest - resolution layer, the optical flow field matrix of each level is divided into uniformly - distributed square grid cells. Assume that each square grid cell has a pixel size of 10×10, and each square grid cell covers a set number of optical flow vectors. The direction angles of all optical flow vectors are statistically analyzed within each square grid cell. The direction angles are divided into multiple direction intervals according to a preset angle interval of 10°, and the histogram of direction angle distributions of each square grid cell is generated. For example, if there are 50 optical flow vectors in a square grid cell, after statistically analyzing the direction angles of these vectors and dividing them according to the 10° interval, the histogram of direction angle distributions of this cell is obtained. The histograms of direction angle distributions of all square grid cells within each level are cumulatively added across cells to obtain the histogram of motion direction distributions at the current level. The maximum - value normalization process is performed on the histogram of motion direction distributions at each level, and the normalized histograms of motion direction distributions at each level are concatenated into a one - dimensional vector in the order of decreasing resolution to form a cross - scale fusion direction - distribution feature vector. Finally, the principal component analysis dimensionality - reduction operation is performed on the direction - distribution feature vector, and the dimensionality - reduced direction - distribution feature vector is weighted and fused layer - by - layer with the average amplitude of the optical flow at each level of the Gaussian pyramid structure to generate the final vector representation of the global motion feature. This global motion feature can accurately characterize the displacement changes of the survivor in three - dimensional space, such as whether the survivor moves to the upper left or the lower right, and the magnitude of the moving speed, etc.
[0018] Meanwhile, a local region centered on the target object (survivor) is intercepted from the stable image sequence. Assume that the intercepted local region has a size of 300×300 pixels. The local region is filtered by a spatial - domain convolution kernel, and the convolution kernel size can be 3×3 to extract a multi - channel texture response map. The multi - channel texture response map is aggregated by a sliding window along the time axis. The sliding window size is 10 frames of images, and the mean and variance of the responses of each channel within the sliding window are statistically analyzed to generate local texture features. These local texture features can characterize the spatio - temporal continuity of the surface details of the survivor, such as the changes in the texture and color of the survivor's clothes at different times, which helps to more accurately identify and track the target in a complex environment.
[0019] Step S130, input the global motion feature and the local texture feature into a pre - trained neural network model for feature fusion to generate the real - time motion prediction parameters of the target object.
[0020] In this embodiment, the previously obtained global motion features and local texture features can be input into a pre-trained neural network model. This neural network model may be pre-trained with a large amount of rescue scenario data. For example, it is trained using the rescue scenario image data collected under multiple different mountain rescue simulation scenarios. The global motion features and local texture features are respectively mapped to the high-dimensional hidden space of the pre-trained neural network model. Assume the dimension of the high-dimensional hidden space is 1024. The global motion features are mapped to obtain the corresponding first hidden vector, and the local texture features are mapped to obtain the second hidden vector.
[0021] The correlation weight matrix between the first hidden vector and the second hidden vector is calculated through the cross-attention mechanism. The cross-attention mechanism can combine the mutual relationship between the global motion features and the local texture features. For example, the movement direction of the survivor may be related to the relative position change of a certain iconic texture on his clothes. The second hidden vector is weighted and reconstructed according to the correlation weight matrix to obtain the enhanced third hidden vector. Then, the first hidden vector and the third hidden vector are added element by element to obtain the fused fourth hidden vector.
[0022] Then, the fourth hidden vector is input into a multi-layer perceptron for non-linear transformation. The first hidden layer of the multi-layer perceptron has 512 neurons. The fourth hidden vector is input into the first hidden layer for linear transformation to obtain the first intermediate feature vector. The first intermediate feature vector is processed by a non-linear activation function, such as using the ReLU function, to generate the activated first non-linear feature vector. The activated first non-linear feature vector is input into the second hidden layer of the multi-layer perceptron. The second hidden layer has 256 neurons. After linear transformation, the second intermediate feature vector is obtained, and then the second intermediate feature vector is processed by a non-linear activation function to generate the activated second non-linear feature vector. The activated second non-linear feature vector is input into the output layer of the multi-layer perceptron for linear transformation to generate the unnormalized original prediction vector.
[0023] The unnormalized original prediction vector is divided into a position sub-vector and a velocity sub-vector. The position sub-vector contains the original data of the three-dimensional space coordinates, and the velocity sub-vector contains the original data of the three-dimensional velocity vector. Normalization scaling processing is performed on the position sub-vector so that its numerical range matches the field-of-view coordinate system of the UAV camera device. For example, the x-coordinate range from -1 to 1 is mapped to the coordinate range corresponding to the left boundary to the right boundary of the camera device's field of view. Normalization scaling processing is performed on the velocity sub-vector so that its numerical range matches the maximum flight speed threshold of the UAV. For example, the maximum horizontal flight speed of the UAV is 20 m / s, and the velocity values in the velocity sub-vector are mapped to the range between 0 and 20. The normalized position coordinate vector and the velocity vector are concatenated to form the final prediction vector. The first three elements are extracted from the final prediction vector as the predicted position coordinates of the survivor at the next moment, and the last three elements are extracted as the predicted velocity vector of the survivor at the next moment. The above predicted position coordinates and velocity vector will provide a key basis for subsequent adjustment of the UAV's flight attitude.
[0024] Step S140, based on the real-time motion prediction parameters, adjust the flight attitude parameters of the UAV to generate corresponding dynamic tracking instructions, and the dynamic tracking instructions are used to control the UAV to maintain a preset relative motion relationship with the target object.
[0025] In this embodiment, based on the previously generated real-time motion prediction parameters, the rescue UAV starts to adjust the flight attitude parameters. For example, first obtain the current flight state data of the UAV. Assume that the current pitch angle of the UAV is 10°, the yaw angle is 30°, the roll angle is 5°, the altitude is 500 meters, and the horizontal speed is 10 m / s. Then, according to the Euclidean distance between the predicted position coordinates and the current position of the UAV, calculate the heading deviation angle in the horizontal plane and the height compensation amount in the vertical direction. For example, if the predicted position of the survivor deviates leftward from the current position of the UAV by a set distance in the horizontal direction, it is calculated that the yaw angle needs to be increased by 15° to align with the new position of the survivor, and in the vertical direction, the altitude needs to be reduced by 50 meters to maintain an appropriate shooting angle, thereby obtaining the height compensation amount in the vertical direction.
[0026] Next, based on the vector difference between the velocity vector and the current horizontal velocity of the UAV, calculate the torque adjustment coefficient of the propulsion motor and the increment of the propeller rotation speed. Assume that it is predicted that the survivor will move left at a speed of 5 m / s, while the current horizontal velocity of the UAV is 10 m / s and the directions are not exactly the same. The horizontal velocity error vector is obtained through vector calculation. Multiply the magnitude of this error vector by a preset propulsion force response coefficient (e.g., 0.5) to get the torque adjustment coefficient of the propulsion motor. At the same time, if it is predicted that the survivor has an upward or downward velocity in the vertical direction, perform a scalar subtraction with the current vertical velocity component of the UAV to obtain the vertical velocity error scalar. Multiply this error scalar by a propeller lift conversion coefficient (e.g., 1.2) to get the increment of the propeller rotation speed.
[0027] Package the adjusted yaw angle target setting value (45°), altitude target setting value (450 m), torque adjustment coefficient, and increment of the propeller rotation speed into the data packet format of the dynamic tracking instruction, and generate the corresponding dynamic tracking instruction. This dynamic tracking instruction will guide the flight control system of the UAV to perform corresponding actions to maintain the preset relative motion relationship with the survivor, ensuring that the position of the survivor can be continuously tracked and clear image information can be obtained during the rescue process.
[0028] Step S150, drive the UAV to perform a visual tracking task according to the dynamic tracking instruction, and at the same time, continuously update the dynamic image sequence and feedback it to the multi-scale spatio-temporal feature extraction step to form a closed-loop tracking control.
[0029] Specifically, after receiving the dynamic tracking instruction, the rescue UAV transmits the instruction to the flight control system through a wireless communication link. The flight control system can synchronously adjust the pitch angle, yaw angle, roll angle, altitude, and propeller rotation speed according to the dynamic tracking instruction. During the adjustment process, the UAV continuously collects the updated dynamic image sequence through the camera device. For example, when the UAV adjusts the yaw angle and altitude according to the instruction, the camera device still captures the images containing the survivor at a frame rate of 30 frames per second.
[0030] Among them, in each frame of the dynamic image after update, the search area of a preset proportion (for example, 20%) can be expanded with the predicted position coordinate as the center. Assuming that the predicted position coordinate is the pixel point (500, 300) in the image, the expanded search area is a rectangular area centered on (500, 300). The color histogram of the pixels in the search area is reversely projected to generate a target probability distribution map. In the target probability distribution map, the connected area with the largest probability density is located. This connected area is likely to be the area where the survivor is located, and then the minimum bounding rectangle of the connected area is calculated. The position parameters of the latest bounding box position are determined based on the vertex coordinates of the minimum bounding rectangle. The position parameters and the position parameters of the historical bounding box are smoothed by Kalman filtering to eliminate jitter noise caused by image noise or slight shaking of the survivor.
[0031] Check the consistency of the latest bounding box position and the predicted position coordinates. Calculate the pixel distance between the center point coordinates of the latest bounding box position and the predicted position coordinates, assuming that the pixel distance is 50 pixels. Perform perspective projection inversion based on the current altitude of the drone (500 meters) and the focal length parameters of the camera device (e.g., 50 mm) to convert the pixel distance into the actual physical distance. If the actual physical distance is greater than the preset threshold (e.g., 10 meters), it is determined that the survivor may have undergone violent movement or been blocked by a new landslide, and the execution of the current tracking instruction is suspended. At this time, clear the historical feature data cached in the neural network model, and re-execute the multi-scale spatiotemporal feature extraction step from the current frame. When re-executing, the position parameters of the latest bounding box are preferentially used as the initial input, skipping the dependency on historical feature data.
[0032] Update the interception range of the local area according to the latest verified bounding box position. Centered on the latest verified bounding box position, expand the interception range according to the preset scaling factor (e.g. 1.5) to form a new local area. Perform Gaussian pyramid downsampling on the new local area to generate a multi-resolution image block. Apply the directional gradient histogram algorithm on the multi-resolution image blocks to extract the spatial gradient distribution features of different resolutions. Superimpose the spatial gradient distribution features of different resolutions according to the weights to generate the updated local texture features. At the same time, re-execute the optical flow field matrix calculation in the expanded interception range, and update the directional distribution histogram of the global motion feature.
[0033] The updated global motion features and local texture features are input into the neural network model to iteratively generate new real-time motion prediction parameters. During each iteration, the change trend of the actual physical distance is recorded. If the decrease rate of the actual physical distance is less than the convergence threshold (e.g., 0.5 meters per iteration) in a continuous preset number of iterations (e.g., 5 times), it is determined that the tracking steady state is entered. In the tracking steady state, the sampling frequency of the camera device is reduced, for example, from 30 frames per second to 15 frames per second, and the operation levels of multi-scale spatio-temporal feature extraction are reduced to reduce the consumption of computing resources. When the actual physical distance continuously remains within the preset threshold for a preset duration (e.g., 10 seconds), the tracking completion flag is activated, and the feature update iteration is stopped according to the tracking completion flag, and the current flight attitude parameters are maintained until a new visual tracking task is received. In this way, the entire system forms a closed-loop tracking control, which can continuously and stably track the target object (survivor) in a complex rescue scenario.
[0034] Based on the above steps, in the embodiment of the present application, the camera device carried by the UAV is used to collect a dynamic image sequence containing the motion trajectory information of the target object in real time. On this basis, multi-scale spatio-temporal feature extraction is performed, and global motion features and local texture features are obtained at the same time. The global motion features can accurately grasp the displacement change of the target object in the three-dimensional space, enabling the UAV to understand the motion trend of the target from a macroscopic level; while the local texture features capture the spatio-temporal continuity of the surface details of the target object, supplementing the information at the microscopic level. Next, the global motion features and local texture features are input into the pre-trained neural network model for feature fusion and generate real-time motion prediction parameters. By using the powerful learning and analysis ability of the neural network, different types of features are effectively integrated, and then the real-time motion situation of the target object is predicted, greatly improving the accuracy and reliability of motion prediction. Based on the real-time motion prediction parameters, the flight attitude parameters of the UAV are adjusted and dynamic tracking instructions are generated to ensure that the UAV can respond in real time according to the motion of the target and maintain a preset relative motion relationship with the target object, so that the UAV no longer simply flies according to the preset path, but can intelligently adapt to the dynamic changes of the target, realizing highly flexible tracking. Finally, according to the dynamic tracking instructions, the UAV is driven to execute the visual tracking task, and the dynamic image sequence is updated in real time to form a closed-loop tracking control, which can continuously adjust and optimize according to the new state of the target, greatly improving the stability and robustness of the tracking. Even in the case where the motion trajectory of the target object is complex and changeable, the UAV can still continuously and accurately track the target, reducing the risk of tracking loss and improving the overall success rate of the visual tracking task.
[0035] In a possible implementation manner, step S120 includes: Step S121: Decompose the dynamic image sequence into multiple consecutive time segments, and perform inter-frame alignment processing on the image frames within each time segment to obtain a stable image sequence with jitter removed.
[0036] For example, each time segment can be set to 5 seconds. Since the frame rate of the imaging device is 30 frames per second, each time segment contains 150 image frames. For these 150 image frames, any image registration algorithm in the related art can be used for inter-frame alignment processing. This image registration algorithm can detect some significant feature points in each frame of the image, such as unique pattern parts on the survivor's clothes or fixed stone contours in the surrounding environment, etc., as reference points. For example, after determining the positions of these reference points in the first frame of the image, the corresponding positions of these reference points can be searched for in the subsequent frames, and then the image is transformed such as translation and rotation according to the deviation of these corresponding positions, so as to eliminate the image jitter caused by the slight shaking during the flight of the drone or the influence of air flow, and obtain a stable image sequence.
[0037] Step S122: Perform frame-by-frame gray normalization processing on the stable image sequence, and construct an optical flow field matrix based on the pixel differences between adjacent frames. The optical flow field matrix is used to quantify the instantaneous motion vector of the target object in the image plane.
[0038] Specifically, the frame-by-frame gray normalization processing is to adjust the pixel gray values of each frame of the image to a unified range. For example, the pixel gray value range of each frame of the image, which is originally distributed arbitrarily between 0 and 255, is normalized to between 0 and 1 through linear transformation and other methods. Then, an optical flow field matrix is constructed based on the pixel differences between adjacent frames. Taking a certain pixel point as an example, in two adjacent frames of the image, calculate the position changes of this pixel point in the horizontal and vertical directions. Suppose the coordinates of a certain pixel point in the nth frame of the image are (x1, y1), and the coordinates of this pixel point become (x2, y2) in the (n + 1)th frame of the image. Then the motion component in the horizontal direction is x2 - x1, and the motion component in the vertical direction is y2 - y1. Such calculations are performed for each pixel point in each frame of the image, and the optical flow field matrix is constructed. This optical flow field matrix is used to quantify the instantaneous motion vector of the survivor in the image plane.
[0039] Step S123: Use the multi-resolution pyramid algorithm to hierarchically analyze the optical flow field matrix, extract the motion direction distribution histograms at different scales, and cascade the motion direction distribution histogram features of each layer to form the global motion feature.
[0040] Among them, Step S123 includes: Step S1231: Perform a Gaussian pyramid generation operation on the optical flow field matrix to obtain a Gaussian pyramid structure including multiple resolution levels. Among them, the highest resolution level corresponds to the original optical flow field matrix, and each lower resolution level is generated by successive downsampling based on a Gaussian kernel.
[0041] Step S1232: Starting from the highest resolution level of the Gaussian pyramid structure, divide the optical flow field matrix of each level into uniformly distributed square grid cells, and each square grid cell covers a preset number of optical flow vectors.
[0042] In this embodiment, the highest resolution level corresponds to the original optical flow field matrix, and its resolution is 1920×1080 (the same as the image resolution collected by the imaging device). Each lower resolution level is generated by successive downsampling based on a Gaussian kernel. For example, the Gaussian kernel size is set to 5×5, and the downsampling ratio is 0.5. Starting from the highest resolution level, divide the optical flow field matrix of each level into uniformly distributed square grid cells. Assume that each square grid cell has a pixel size of 10×10, and each square grid cell covers a set number of optical flow vectors. Taking the highest resolution level as an example, the size of the optical flow field matrix at this level is 1920×1080 pixels. Then, 192 square grid cells can be divided in the horizontal direction (1920÷10 = 192), 108 square grid cells can be divided in the vertical direction (1080÷10 = 108), and a total of 192×108 square grid cells can be divided.
[0043] Step S1233: In each square grid cell, count the direction angles of all optical flow vectors, and divide the direction angles into multiple direction intervals according to a preset angle interval to generate a direction angle distribution histogram for each square grid cell.
[0044] For example, set the preset angle interval to 10°. For the optical flow vectors in a certain square grid cell, calculate the angle between each optical flow vector and the horizontal direction as the direction angle. Assume that there are 30 optical flow vectors in a certain square grid cell. After calculating their direction angles respectively, divide them according to the 10° interval. If there are 5 optical flow vectors with direction angles between 0° and 10°, 8 between 10° and 20°, and so on, a direction angle distribution histogram for this square grid cell is generated.
[0045] Step S1234: Perform cross-cell accumulation on the direction angle distribution histograms of all square grid cells within each level to obtain the motion direction distribution histogram of the current level.
[0046] For example, at a certain level, there are 100 square grid cells. For each square grid cell, a histogram of the angular distribution is obtained, and then the values in the corresponding intervals of these 100 histograms of the angular distribution are accumulated. For example, in the interval from 0° to 10°, the value of the first square grid cell is 5, the second is 3, and so on. Summing up these 100 values gives the accumulated value of this level in the interval from 0° to 10°. Performing the above operation for all intervals yields the histogram of the motion direction distribution of the current level.
[0047] Step S1235: Perform maximum normalization processing on the histogram of the motion direction distribution of each level, and perform one-dimensional vector splicing on the normalized histograms of the motion direction distribution of each level in the order from the highest resolution to the lowest resolution to form a direction distribution feature vector for cross-scale fusion.
[0048] For example, the maximum normalization processing is to divide the value of each interval by the maximum value in the histogram of the motion direction distribution of this level, so that the values of all intervals are between 0 and 1. For example, the value of the histogram of the motion direction distribution of a certain level in the interval from 0° to 10° is 20, and the maximum value of this level is 50. Then the value of this interval after normalization is 20÷50 = 0.4. Perform one-dimensional vector splicing on the normalized histograms of the motion direction distribution of each level in the order from the highest resolution to the lowest resolution. Suppose the histogram of the motion direction distribution of the highest resolution level after normalization results in a vector of length 10, and the next lower resolution level after processing results in a vector of length 8. Splicing them in order gives a direction distribution feature vector for cross-scale fusion with a length of 18.
[0049] Step S1236: Perform principal component analysis and dimensionality reduction operation on the direction distribution feature vector, and perform layer-by-layer weighted fusion on the dimensionality-reduced direction distribution feature vector and the average optical flow amplitude of each level of the Gaussian pyramid structure to generate the final vector representation of the global motion feature.
[0050] Specifically, for the principal component analysis dimensionality reduction operation, the covariance matrix of the direction distribution feature vectors is calculated, and then its eigenvalues and eigenvectors are obtained. The first few main eigenvectors are selected to reduce the dimensionality of the direction distribution feature vectors. Assume that the length of the direction distribution feature vectors after dimensionality reduction becomes 8. For the average optical flow amplitude of each level of the Gaussian pyramid structure, the average value of the amplitudes of all optical flow vectors in the optical flow field matrix of each level is calculated. For example, in the optical flow field matrix of the highest resolution level, there are 192×108×10×10 optical flow vectors (the total number of optical flow vectors is calculated according to the previously divided square grid cells). After adding up the amplitudes of these optical flow vectors and dividing by the total number of optical flow vectors, the average optical flow amplitude of this level is obtained. Then, the direction distribution feature vectors after dimensionality reduction and the average optical flow amplitude of each level are weighted and fused according to the set weights. For example, the weight of the direction distribution feature vectors after dimensionality reduction is 0.6, and the weight of the average optical flow amplitude of each level is 0.4. After multiplying them according to the corresponding elements and adding them up, the final vector representation of the global motion feature is obtained.
[0051] Step S124: Intercept a local area centered on the target object in the stable image sequence, and perform spatial domain convolution kernel filtering on the local area to extract a multi-channel texture response map.
[0052] Specifically, assume that the intercepted local area is 300×300 pixels in size, and this local area is determined centered on the position of the survivor in the image. The spatial domain convolution kernel filtering uses a 3×3 convolution kernel. For each pixel point in the local area, the 9 pixel values covered by the convolution kernel are multiplied by the corresponding weights of the convolution kernel and then added up to obtain the filtered pixel value. For example, for the pixel point in the upper left corner of the local area, the 9 pixel values covered by the convolution kernel are p1, p2, p3, p4, p5, p6, p7, p8, p9 respectively, and the weights of the convolution kernel are w1, w2, w3, w4, w5, w6, w7, w8, w9 respectively. Then the filtered pixel value is p1×w1 + p2×w2 + p3×w3 + p4×w4 + p5×w5 + p6×w6 + p7×w7 + p8×w8 + p9×w9. Such an operation is performed on each pixel point in the local area to obtain the filtered image, which contains multiple channels (such as three RGB channels), thereby extracting the multi-channel texture response map.
[0053] Step S125: Aggregate the multi-channel texture response map along the time axis using a sliding window, and statistically calculate the mean and variance of the responses of each channel within the sliding window to generate the local texture feature.
[0054] For example, set the sliding window size to 10 frames of images, and use the 1st to 10th frames of the multi-channel texture response map as the first sliding window. For each channel (taking the R channel as an example), calculate the mean value of the corresponding pixel positions in the R channel of these 10 frames of images. Suppose at a certain pixel position, the R channel value of the 1st frame is r1, the 2nd frame is r2, and so on, the 10th frame is r10. Then the mean value of the R channel at this pixel position within this sliding window is (r1 + r2 +... + r10) ÷ 10. Similarly, calculate the variance of the R channel at this pixel position within this sliding window. The calculation of the variance is to first calculate the square of the difference between each value and the mean value, such as (r1 - mean)^2, (r2 - mean)^2, etc., and then divide the sum of these squared values by 10. Perform the above mean and variance calculations for each channel to obtain the mean and variance of the responses of each channel within the sliding window. The above values are combined to generate local texture features.
[0055] In a possible implementation manner, step S130 includes: Step S131, map the global motion feature and the local texture feature to the high-dimensional hidden space of a pre-trained neural network model respectively to obtain corresponding first and second hidden vectors.
[0056] In this embodiment, assume that the dimension of the high-dimensional hidden space set by the pre-trained neural network model is 512 dimensions. For the global motion feature, which contains multiple elements, use the trained mapping function to map these elements into the 512-dimensional high-dimensional hidden space. This mapping process is carried out according to the mapping rules determined during the pre-training of the neural network model, and after mapping, the first hidden vector is obtained. Similarly, for the local texture feature, map it into the 512-dimensional high-dimensional hidden space according to the same mapping rules to obtain the second hidden vector.
[0057] Step S132, calculate the correlation weight matrix between the first hidden vector and the second hidden vector through the cross-attention mechanism, and perform weighted reconstruction on the second hidden vector according to the correlation weight matrix to obtain an enhanced third hidden vector.
[0058] Specifically, the cross-attention mechanism can comprehensively consider the relationships between each element in the first hidden vector and the second hidden vector. Taking an element in the first hidden vector as an example, it can be calculated with all elements in the second hidden vector, and the calculation process may involve operations such as dot product. Suppose the first element of the first hidden vector is a1, and the second hidden vector has 512 elements b1, b2, ..., b512 respectively. Calculate the relationship values between a1 and b1, a1 and b2, etc. After a series of calculations (this calculation is based on the predefined cross-attention calculation method of the neural network model), an association weight matrix between the first hidden vector and the second hidden vector is obtained. Each element in this association weight matrix represents the degree of association between the element in the first hidden vector and the element in the second hidden vector. Then, the second hidden vector is weighted and reconstructed according to this association weight matrix, that is, each element in the association weight matrix is multiplied by the corresponding element in the second hidden vector, and then the multiplied results are added together to obtain an enhanced third hidden vector. For example, an element in the association weight matrix is w1, and the corresponding element in the second hidden vector is b1, then the reconstructed element is w1×b1. Performing such operations on all elements in the second hidden vector yields the enhanced third hidden vector.
[0059] Step S133: Add the first hidden vector and the third hidden vector element by element to obtain a fused fourth hidden vector.
[0060] Element-by-element addition means adding the first element of the first hidden vector to the first element of the third hidden vector, the second element of the first hidden vector to the second element of the third hidden vector, and so on. For example, if the first hidden vector is [a1, a2, a3, ...] and the third hidden vector is [b1, b2, b3, ...], then the fused fourth hidden vector is [a1 + b1, a2 + b2, a3 + b3, ...].
[0061] Step S134: Input the fourth hidden vector into a multi-layer perceptron for non-linear transformation, and output the predicted position coordinates and velocity vector of the target object at the next moment.
[0062] For example, step S134 includes: Step S1341: Input the fourth hidden vector into the first hidden layer of the multi-layer perceptron for linear transformation to obtain a first intermediate feature vector.
[0063] In this embodiment, the first hidden layer of the multi-layer perceptron has a specific number of neurons, assumed to be 256 neurons. There are predefined connection weights between each element in the fourth hidden vector and the neurons in the first hidden layer, and these weights are determined during the pre-training stage of the neural network. For the first element of the fourth hidden vector, it will be multiplied by the connection weights of each neuron in the first hidden layer, and then the results of these multiplications will be accumulated on the corresponding neurons. For example, if the fourth hidden vector is [e1, e2, e3, …] and the connection weights between the first neuron in the first hidden layer and the elements of the fourth hidden vector are [w11, w12, w13, …], then the accumulated result on the first neuron is e1×w11 + e2×w12 + e3×w13 + …, and this result is the first element of the first intermediate feature vector. Calculate the other elements of the first intermediate feature vector in the same way to obtain the first intermediate feature vector.
[0064] Step S1342: Perform a non-linear activation function processing on the first intermediate feature vector to generate an activated first non-linear feature vector.
[0065] For example, the ReLU function is used here as the non-linear activation function. For each element of the first intermediate feature vector, if the value of this element is less than 0, then it is set to 0; if the value of this element is greater than or equal to 0, it remains unchanged. For example, if the first intermediate feature vector is [f1, f2, f3, …], if f1 is less than 0, then the first element of the activated first non-linear feature vector is 0; if f2 is greater than or equal to 0, then the second element of the activated first non-linear feature vector is f2, and so on, to obtain the activated first non-linear feature vector.
[0066] Step S1343: Input the activated first non-linear feature vector into the second hidden layer of the multi-layer perceptron for linear transformation to obtain a second intermediate feature vector.
[0067] For example, the second hidden layer also has a predefined number of neurons, assumed to be 128. Similar to the calculation method of the first hidden layer, each element of the activated first non-linear feature vector is multiplied by the connection weights of the neurons in the second hidden layer, and then accumulated on the neurons to obtain the elements of the second intermediate feature vector. For example, if the activated first non-linear feature vector is [g1, g2, g3, …] and the connection weights between the first neuron in the second hidden layer and the elements of this vector are [w21, w22, w23, …], then the first element of the second intermediate feature vector is g1×w21 + g2×w22 + g3×w23 + …, and all elements of the second intermediate feature vector are calculated in this way.
[0068] Step S1344: Perform a non - linear activation function processing on the second intermediate feature vector to generate an activated second non - linear feature vector.
[0069] For example, also using the ReLU function, for each element of the second intermediate feature vector, if it is less than 0, it becomes 0, and if it is greater than or equal to 0, it remains unchanged. For example, if the second intermediate feature vector is [h1, h2, h3,...], if h1 is less than 0, the first element of the activated second non - linear feature vector is 0; if h2 is greater than or equal to 0, the second element of the activated second non - linear feature vector is h2, and so on, to obtain the activated second non - linear feature vector.
[0070] Step S1345: Input the activated second non - linear feature vector into the output layer of the multi - layer perceptron for linear transformation to generate an unnormalized original prediction vector.
[0071] For example, the output layer has a specific number of neurons, assumed to be 6, corresponding to the three dimensions (x, y, z) of the predicted position coordinates and the three dimensions (vx, vy, vz) of the velocity vector respectively. Each element of the activated second non - linear feature vector is multiplied by the connection weight of the output layer neuron and accumulated to obtain the element of the unnormalized original prediction vector. For example, if the activated second non - linear feature vector is [i1, i2, i3,...], the connection weight of the first neuron in the output layer with the elements of this vector is [w31, w32, w33,...], and the first element of the unnormalized original prediction vector is i1×w31 + i2×w32 + i3×w33+..., and all elements of the unnormalized original prediction vector are obtained in this way.
[0072] Step S1346: Split the unnormalized original prediction vector into a position sub - vector and a velocity sub - vector, where the position sub - vector contains the original data of the three - dimensional space coordinates, and the velocity sub - vector contains the original data of the three - dimensional velocity vector.
[0073] For example, if the unnormalized original prediction vector is [j1, j2, j3, j4, j5, j6], then the position sub - vector is [j1, j2, j3], and the velocity sub - vector is [j4, j5, j6].
[0074] Step S1347: Perform a normalization scaling process on the position sub - vector to make its numerical range match the field - of - view coordinate system of the UAV camera device, and generate a normalized position coordinate vector.
[0075] Assume that in the field-of-view coordinate system of the UAV camera device, the range of the x coordinate is from -1 to 1, the range of the y coordinate is from -1 to 1, and the range of the z coordinate is from -1 to 1. For the original value of the x coordinate in the position component sub-vector, assumed to be x0, first find the minimum value min_x and the maximum value max_x of the x coordinate in the position component sub-vector, and then perform normalization by calculating (x0 - min_x) ÷ (max_x - min_x) × 2 - 1. For example, if the minimum value of the x coordinate in the position component sub-vector is -5, the maximum value is 5, and the original value x0 is 0, then the normalized x coordinate is (0 - (-5)) ÷ (5 - (-5)) × 2 - 1 = 0. Normalize the y and z coordinates in the same way to obtain the normalized position coordinate vector.
[0076] Step S1348, perform normalization scaling on the velocity component sub-vector so that its numerical range matches the maximum flight speed threshold of the UAV, and generate the normalized velocity vector.
[0077] Assume that the maximum horizontal flight speed of the UAV is 20 m / s and the vertical flight speed is 10 m / s. For the original value of the horizontal speed in the velocity component sub-vector, assumed to be vx0, first find the maximum value max_vx of the horizontal speed in the velocity component sub-vector, and then perform normalization by calculating vx0 ÷ max_vx × 20. For example, if the maximum value of the horizontal speed in the velocity component sub-vector is 50 m / s and the original value vx0 is 25 m / s, then the normalized horizontal speed is 25 ÷ 50 × 20 = 10 m / s. Normalize the vertical speed in a similar way to obtain the normalized velocity vector.
[0078] Step S1349, splice the normalized position coordinate vector and the velocity vector to form the final prediction vector.
[0079] Step S13410, extract the first three elements from the final prediction vector as the predicted position coordinates of the target object at the next moment, and extract the last three elements as the predicted velocity vector of the target object at the next moment.
[0080] For example, if the normalized position coordinate vector is [k1, k2, k3] and the normalized velocity vector is [k4, k5, k6], then the final prediction vector is [k1, k2, k3, k4, k5, k6]. Extract the first three elements from the final prediction vector as the predicted position coordinates of the survivor at the next moment, and extract the last three elements as the predicted velocity vector of the survivor at the next moment. Assume that the final prediction vector is [l1, l2, l3, l4, l5, l6], then the predicted position coordinates are (l1, l2, l3), and the predicted velocity vector is (l4, l5, l6).
[0081] Step S135, calculate the heading angle adjustment amount, flight altitude correction amount, and propulsion force parameters required for the UAV based on the predicted position coordinates and velocity vector, and generate the real-time motion prediction parameters.
[0082] For example, step S135 includes: Step S1351, obtain the current three-dimensional space coordinates and current velocity vector of the UAV, where the three-dimensional space coordinates include longitude, latitude, and altitude components.
[0083] For example, assume that the current three-dimensional space coordinates of the UAV are (x1, y1, z1), where x1 is the longitude, y1 is the latitude, z1 is the altitude component, and the current velocity vector is (vx1, vy1, vz1).
[0084] Step S1352, subtract the longitude component of the predicted position coordinates from the current longitude component of the UAV to obtain the first lateral displacement deviation, and subtract the latitude component of the predicted position coordinates from the current latitude component of the UAV to obtain the second lateral displacement deviation.
[0085] For example, assume that the predicted position coordinates are (x2, y2, z2), and the first lateral displacement deviation is x2 - x1. Subtract the latitude component of the predicted position coordinates from the current latitude component of the UAV to obtain the second lateral displacement deviation, that is, y2 - y1.
[0086] Step S1353, calculate the comprehensive displacement vector in the horizontal plane based on the first lateral displacement deviation and the second lateral displacement deviation, and subtract the direction angle of the comprehensive displacement vector from the current yaw angle of the UAV to obtain the initial heading deviation angle.
[0087] For example, first calculate the magnitude of the comprehensive displacement deviation. According to the Pythagorean theorem, the magnitude of the comprehensive displacement deviation is sqrt((x2 - x1)^2 + (y2 - y1)^2). For example, if x2 - x1 = 3 and y2 - y1 = 4, then the magnitude of the comprehensive displacement deviation is sqrt(3^2 + 4^2) = 5. Then calculate the direction angle of the comprehensive displacement vector according to trigonometric functions. Assume the direction angle is θ, tanθ = (y2 - y1) ÷ (x2 - x1), and obtain the value of θ through the arctangent function. If x2 - x1 = 3 and y2 - y1 = 4, then tanθ = 4 ÷ 3, and the value of θ obtained through the arctangent function is approximately 53.13°. Subtract the direction angle of the comprehensive displacement vector from the current yaw angle of the UAV to obtain the initial heading deviation angle. Assume the current yaw angle of the UAV is α, and the initial heading deviation angle is θ - α.
[0088] Step S1354: Subtract the altitude component of the predicted position coordinates from the current altitude component of the UAV to obtain the height displacement deviation in the vertical direction, and divide the height displacement deviation by a preset unit time step to generate a flight height correction amount.
[0089] Specifically, subtract the altitude component of the predicted position coordinates from the current altitude component of the UAV to obtain the height displacement deviation in the vertical direction, that is, z2 - z1, and divide the height displacement deviation by a preset unit time step to generate a flight height correction amount. Assume that the preset unit time step is 1 second. If z2 - z1 = 10 meters, then the flight height correction amount is 10÷1 = 10 meters.
[0090] Step S1355: Perform a vector subtraction on the horizontal component of the velocity vector and the horizontal component of the current velocity vector of the UAV to obtain a horizontal velocity error vector, and multiply the magnitude of the horizontal velocity error vector by a preset propulsion force response coefficient to generate a torque adjustment coefficient for the propulsion motor.
[0091] For example, assume that the horizontal component of the velocity vector is (vx2, vy2), and the horizontal velocity error vector is (vx2 - vx1, vy2 - vy1). For example, if vx2 = 15 m / s, vx1 = 10 m / s, vy2 = 5 m / s, and vy1 = 3 m / s, then the horizontal velocity error vector is (15 - 10, 5 - 3) = (5, 2). Multiply the magnitude of the horizontal velocity error vector by a preset propulsion force response coefficient to generate a torque adjustment coefficient for the propulsion motor. First, calculate the magnitude of the horizontal velocity error vector. According to the Pythagorean theorem, the magnitude of the horizontal velocity error vector is sqrt((vx2 - vx1)^2+(vy2 - vy1)^2). If (vx2 - vx1) = 5 and (vy2 - vy1) = 2, then the magnitude of the horizontal velocity error vector is sqrt(5^2 + 2^2) = sqrt(29). Assume that the preset propulsion force response coefficient is 0.5, then the torque adjustment coefficient for the propulsion motor is sqrt(29)×0.5.
[0092] Step S1356: Perform a scalar subtraction on the vertical component of the velocity vector and the vertical component of the current velocity vector of the UAV to obtain a vertical velocity error scalar, and multiply the vertical velocity error scalar by a propeller lift conversion coefficient to generate an increment in the propeller rotation speed.
[0093] For example, the vertical component of the velocity vector can be subtracted scalarially from the vertical component of the current velocity vector of the UAV to obtain a scalar of the vertical velocity error, i.e., vz2 - vz1, and the scalar of the vertical velocity error is multiplied by the propeller lift conversion coefficient to generate an increment in the propeller rotation speed. Suppose vz2 = 8 m / s, vz1 = 6 m / s, and the propeller lift conversion coefficient is 1.2. Then the scalar of the vertical velocity error is 8 - 6 = 2 m / s, and the increment in the propeller rotation speed is 2 × 1.2 = 2.4.
[0094] Step S1357: Input the initial heading deviation angle into an angle smoothing filter for noise suppression to obtain a smoothed heading angle adjustment amount.
[0095] For example, the angle smoothing filter processes the initial heading deviation angle according to a set algorithm (such as a weighted average algorithm). Suppose the initial heading deviation angle is β. The angle smoothing filter performs a weighted average calculation based on the heading deviation angle values at several previous moments. For example, the heading deviation angles at the previous three moments are β-1, β-2, β-3 respectively, and the corresponding weights are w1, w2, w3. Then the smoothed heading angle adjustment amount is (β-1×w1 + β-2×w2 + β-3×w3)÷(w1 + w2 + w3).
[0096] Step S1358: Perform interval truncation processing on the flight altitude correction amount, torque adjustment coefficient, and increment in the propeller rotation speed respectively with corresponding dynamic thresholds to generate a limited flight altitude correction amount, torque adjustment coefficient, and increment in the propeller rotation speed.
[0097] For example, suppose the dynamic threshold for the flight altitude correction amount is from -5 to 5 m. If the calculated flight altitude correction amount is -8 m, then it is truncated to -5 m; if the calculated flight altitude correction amount is 8 m, then it is truncated to 5 m. The torque adjustment coefficient and the increment in the propeller rotation speed are processed in a similar manner. Suppose the dynamic threshold for the torque adjustment coefficient is from 0 to 1. If the calculated torque adjustment coefficient is 1.5, then it is truncated to 1; suppose the dynamic threshold for the increment in the propeller rotation speed is from 0 to 3. If the calculated increment in the propeller rotation speed is 4, then it is truncated to 3.
[0098] Step S1359: Package the smoothed heading angle adjustment amount, limited flight altitude correction amount, torque adjustment coefficient, and increment in the propeller rotation speed into a structured parameter set according to a preset coding rule to generate real-time motion prediction parameters containing multi-dimensional adjustment instructions.
[0099] For example, the preset coding rule stipulates operations such as arranging these parameters in a specific order and adding some check bits to generate real-time motion prediction parameters, which will be used to adjust the flight attitude of the UAV to track the survivor.
[0100] In a possible implementation, step S140 includes: Step S141, obtaining the current flight state data of the drone, where the current flight state data includes pitch angle, yaw angle, roll angle, altitude, and horizontal speed.
[0101] For example, assume that the current pitch angle of the drone is 12°. This angle represents the degree of inclination of the drone body in the front - rear direction relative to the horizontal plane; the yaw angle is 35°, which reflects the deviation angle of the drone's nose relative to the due - north direction; the roll angle is 8°, which reflects the rotation angle of the drone body around the flight direction axis; the altitude is 550 meters, which is the vertical height of the drone relative to the sea level; and the horizontal speed is 12 m / s, which is the flight speed of the drone in the horizontal direction.
[0102] Step S142, calculating the course deviation angle in the horizontal plane and the height compensation amount in the vertical direction according to the Euclidean distance between the predicted position coordinates and the current position of the drone.
[0103] Specifically, the predicted position coordinates are the predicted position of the survivor at the next moment. Assume the predicted position coordinates are (x2, y2, z2), and the current position coordinates of the drone are (x1, y1, z1). First, calculate the Euclidean distance on the horizontal plane. According to the principle of the Euclidean distance formula, the distance d on the horizontal plane is d = sqrt((x2 - x1)^2+(y2 - y1)^2). For example, if x2 - x1 = 5 meters and y2 - y1 = 3 meters, then d = sqrt(5^2+3^2)=sqrt(34) meters. Then calculate the course deviation angle θ in the horizontal plane according to trigonometric functions, tanθ=(y2 - y1)÷(x2 - x1). In the above example, tanθ = 3÷5 = 0.6, and the value of θ obtained through the arctangent function is approximately 31°. The height compensation amount in the vertical direction is z2 - z1. Assume z2 - z1=-20 meters, which means the predicted position is 20 meters lower than the current position in the vertical direction.
[0104] Step S143, adjusting the target set value of the yaw angle based on the course deviation angle and adjusting the target set value of the altitude based on the height compensation amount.
[0105] For example, if the current yaw angle is 35° and the calculated course deviation angle is 31°, then the adjusted target set value of the yaw angle is 35°+31° = 66°. If the current altitude is 550 meters and the height compensation amount is - 20 meters, the adjusted target set value of the altitude is 550 - 20 = 530 meters.
[0106] Step S144: Calculate the torque adjustment coefficient of the propulsion motor and the increment of the propeller rotation speed according to the vector difference between the velocity vector and the current horizontal velocity of the UAV.
[0107] Suppose the predicted velocity vector is (vx2, vy2), and the current horizontal velocity of the UAV is (vx1, vy1). The vector difference of velocities is (vx2 - vx1, vy2 - vy1). For example, if vx2 = 15 m / s, vx1 = 12 m / s, vy2 = 4 m / s, and vy1 = 3 m / s, then the vector difference of velocities is (15 - 12, 4 - 3) = (3, 1). First, calculate the magnitude of the vector difference of velocities. According to the Pythagorean theorem, the magnitude is sqrt((3)^2+(1)^2)=sqrt(10) m / s. The torque adjustment coefficient of the propulsion motor is related to the magnitude of the vector difference of velocities. Suppose there is a preset proportional coefficient k1 (this proportional coefficient is determined according to the characteristics of the UAV's power system), for example, k1 = 0.3, then the torque adjustment coefficient of the propulsion motor is sqrt(10)×0.3. The increment of the propeller rotation speed is related to the velocity change in the vertical direction. Suppose the vertical component of the predicted velocity vector vz2 = 6 m / s, the current vertical velocity of the UAV vz1 = 5 m / s, and the velocity difference is vz2 - vz1 = 1 m / s. There is a coefficient k2 related to the characteristics of the propeller, for example, k2 = 1.5, then the increment of the propeller rotation speed is 1×k2 = 1.5.
[0108] Step S145: Package the adjusted yaw angle target set value, altitude target set value, torque adjustment coefficient, and propeller rotation speed increment into the data packet format of the dynamic tracking instruction, and generate the corresponding dynamic tracking instruction.
[0109] Specifically, according to a specific data packet format, first convert the yaw angle target set value into a digital form with a specific precision (such as reserved to two decimal places), suppose it is 66.00°, the altitude target set value is 530 m, the torque adjustment coefficient is reserved with an appropriate precision according to the previous calculation result, suppose it is 0.95, and the increment of the propeller rotation speed is 1.5. Arrange these data in a predetermined order and add some necessary check information (such as checksum, etc.), so as to generate the dynamic tracking instruction. This dynamic tracking instruction will be sent to the flight control system of the UAV to guide the UAV to adjust the flight attitude, so as to continuously track the position of the survivor in the mountain earthquake rescue scenario.
[0110] In a possible implementation manner, step S150 includes: Step S151: Transmit the dynamic tracking instruction to the flight control system of the UAV through a wireless communication link, and trigger the flight control system to synchronously adjust the pitch angle, yaw angle, roll angle, altitude, and propeller rotation speed.
[0111] In this embodiment, the wireless communication link adopts a specific communication protocol, such as a wireless communication protocol based on the IEEE 802.11 standard, to ensure the accurate transmission of instructions. After receiving the instructions, the flight control system adjusts relevant flight attitude parameters according to the yaw angle target setting value, altitude target setting value, torque adjustment coefficient, and propeller speed increment in the instructions. For example, if the yaw angle target setting value in the instructions is 66.00°, the flight control system will adjust components such as the rudder of the UAV to gradually approach this target value; for an altitude target setting value of 530 meters, the lift is changed by controlling the speed of the propellers to adjust the altitude of the UAV; the output torque of the propulsion motor is adjusted according to the torque adjustment coefficient, and the speed of the propellers is adjusted according to the propeller speed increment. At the same time, the pitch angle and roll angle are also adjusted accordingly to maintain the stable flight attitude of the UAV.
[0112] Step S152, during the attitude adjustment process of the UAV, continuously collect the updated dynamic image sequence through the imaging device, and mark the latest bounding box position of the target object in the dynamic image.
[0113] Specifically, the imaging device collects images at a fixed frame rate, such as 30 frames per second. In each frame of the image, the latest bounding box position of the survivor is determined through a target detection algorithm. This target detection algorithm can search for possible target areas in the image based on the previous feature learning of the survivor (such as color features, texture features, etc.), and then determine the smallest rectangular area containing the survivor as the bounding box. For example, in a frame of the image, the body contour of the survivor is detected through the target detection algorithm, and its upper left corner coordinates are determined as (x3, y3), and the lower right corner coordinates are (x4, y4). This rectangular area (x3, y3, x4, y4) is the latest bounding box position.
[0114] Step S153, perform a consistency check on the latest bounding box position and the predicted position coordinates. If the consistency check result indicates that the position deviation exceeds a preset threshold, re-initialize the multi-scale spatio-temporal feature extraction step.
[0115] In detail, the coordinates of the center point of the latest bounding box position are calculated, assuming that it is ((x3+x4)÷2, (y3+y4)÷2), and the deviation is calculated between it and the predicted position coordinate (x2, y2). First, the horizontal deviation is calculated, that is, ((x3+x4)÷2-x2), and then the vertical deviation is calculated ((y3+y4)÷2-y2). Then, the total position deviation is calculated according to the Pythagorean theorem, that is, sqrt(((x3+x4)÷2-x2)^2+((y3+y4)÷2-y2)^2). Assuming that the preset threshold is 10 meters, if the calculated position deviation is greater than 10 meters, this may mean that the survivor has undergone sudden violent movement or a new occlusion has occurred. At this point, the multi-scale spatiotemporal feature extraction step needs to be reinitialized, which means clearing all temporary data stored in this process, such as the previous optical flow field matrix calculation results, the intermediate data in the local texture feature calculation process, etc., and then re-performing the multi-scale spatiotemporal feature extraction operation starting from the current frame image.
[0116] Step S154, updating the clipping range of the local area according to the verified latest bounding box position, and recalculating the global motion feature and the local texture feature to obtain updated global motion feature and local texture feature.
[0117] In detail, the latest bounding box position after verification can be used as the center. Assuming that the previous local area is a 300×300 pixel area intercepted with the predicted position coordinate as the center, it is now adjusted according to the latest bounding box position. For example, the coordinate change of the latest bounding box position requires the local area to be enlarged. According to the set rules, such as adding 100 pixels in the horizontal and vertical directions, the new local area size becomes 500×500 pixels. The previous operations are repeated on this new local area to calculate the global motion features and local texture features. For the global motion features, the optical flow field matrix is recalculated, including processing the image frames in the new local area, constructing the optical flow field matrix, and then using the multi-resolution pyramid algorithm for hierarchical analysis to extract the motion direction distribution histogram at different scales. After a series of operations, the updated global motion features are obtained; for the local texture features, the new local area is re-filtered with a spatial convolution kernel to extract a multi-channel texture response map, and then a sliding window aggregation is performed along the time axis, and the mean and variance of the responses of each channel in the sliding window are statistically calculated, so as to obtain the updated local texture features.
[0118] Step S155, inputting the updated global motion features and local texture features into the neural network model to iteratively generate new real-time motion prediction parameters until the relative motion relationship between the drone and the target object meets the preset tracking accuracy.
[0119] In this embodiment, the neural network model performs operations such as feature fusion and non-linear transformation based on the new input, and re-outputs new predicted position coordinates and velocity vectors. Then, based on these values, real-time motion prediction parameters such as the new heading angle adjustment amount, flight altitude correction amount, and propulsion force parameters are calculated. This process is continuously iterated, and the flight attitude of the UAV is adjusted according to the new predicted position and velocity each time. For example, in one iteration, if the new predicted position coordinates are closer to the actual position of the survivor, the adjustment amplitude of the UAV's flight attitude will become smaller. This process continues until the relative position deviation and speed matching between the UAV and the survivor meet the preset tracking accuracy requirements, such as the position deviation is within 5 meters and the speed matching error is within 1 meter / second. At this time, it indicates that the UAV can stably and accurately track the survivor and complete the visual tracking task.
[0120] In a possible implementation manner, step S152 includes: Step S1521, in each updated dynamic image, expand a search area with a preset ratio centered on the predicted position coordinates.
[0121] Specifically, in each updated dynamic image, expand a search area with a preset ratio centered on the predicted position coordinates. Assume that the predicted position coordinates are (x2, y2), the preset ratio is 20%, and the resolution of the image is 1920×1080 pixels. Then, in the horizontal direction, the left and right boundaries of the search area are x2-(1920×0.1) and x2+(1920×0.1) respectively; in the vertical direction, the upper and lower boundaries of the search area are y2-(1080×0.1) and y2+(1080×0.1) respectively, thus determining a search area centered on the predicted position coordinates.
[0122] Step S1522, perform color histogram backprojection on the pixels within the search area to generate a target probability distribution map.
[0123] For example, the distribution characteristics of the survivor's clothing color have been analyzed before. By comparing the color of each pixel within the search area with this color distribution characteristic, the probability that each pixel belongs to the survivor is calculated. For each pixel within the search area, assume its color is c and the corresponding probability in the color histogram is p(c), then the value of this pixel in the target probability distribution map is p(c). By performing such calculations on all pixels within the search area, a target probability distribution map is generated.
[0124] Step S1523, locate the connected region with the maximum probability density in the target probability distribution map, and calculate the minimum bounding rectangle of the connected region.
[0125] In the target probability distribution map, by traversing all pixels, find the pixel with the largest probability density as the starting point, and then start searching around from this pixel to form a connected area with a probability density greater than a certain threshold (such as 0.5) and interconnected pixels. This connected area may contain the whole or part of the survivor's body. Then calculate the minimum bounding rectangle of the connected area, that is, find the minimum rectangle that can completely contain the connected area. Assume that the coordinates of the upper left vertex of the minimum bounding rectangle are (x3, y3), and the coordinates of the lower right vertex are (x4, y4).
[0126] Step S1524: determining the position parameters of the latest bounding box position according to the vertex coordinates of the minimum circumscribed rectangle.
[0127] For example, the vertex coordinates (x3, y3, x4, y4) of the minimum bounding rectangle are the position parameters of the latest bounding box position, which determine the approximate position range of the survivors in this frame of the image.
[0128] Step S1525, performing Kalman filtering smoothing processing on the position parameters and the position parameters of the historical boundary box to eliminate jitter noise.
[0129] Kalman filtering is an optimal estimation algorithm based on the state space model of a linear system. Assume that the state vector of the historical bounding box position parameters is [x_last, y_last, vx_last, vy_last], where x_last and y_last are the center coordinates of the bounding box in the previous frame, and vx_last and vy_last are the estimated velocities of the center coordinates in the x and y directions. The bounding box position parameters of the current frame are [x_current, y_current]. First, the state vector of the current frame is predicted according to the state transfer equation, which may be [x_predicted=x_last+vx_last, y_predicted=y_last+vy_last, vx_predicted=vx_last, vy_predicted=vy_last]. Then, the error between the measured value and the predicted value is calculated according to the measurement equation, which is [z=[x_current, y_current]-[x_predicted, y_predicted]]. Next, the updated state vector is calculated according to the Kalman gain, which is calculated based on parameters such as the error covariance matrix. Finally, the smoothed position parameters are obtained. This process effectively eliminates the noise caused by factors such as jitter during image acquisition or slight shaking of the survivor.
[0130] In a possible implementation, step S153 includes: Step S1531, calculate the pixel distance between the center point coordinates of the latest bounding box position and the predicted position coordinates.
[0131] For example, the center point coordinates of the latest bounding box position are ((x3 + x4)÷2, (y3 + y4)÷2), and the pixel distance from the predicted position coordinates (x2, y2) is calculated as follows. First, calculate the pixel distance in the horizontal direction as the absolute value of ((x3 + x4)÷2 - x2), and the pixel distance in the vertical direction as the absolute value of ((y3 + y4)÷2 - y2). Then, calculate the total pixel distance according to the Pythagorean theorem, that is, sqrt(((x3 + x4)÷2 - x2)^2 + ((y3 + y4)÷2 - y2)^2).
[0132] Step S1532, convert the pixel distance to the actual physical distance, specifically through perspective projection inversion using the current altitude of the UAV and the focal length parameter of the camera device.
[0133] Assume that the current altitude of the UAV is h meters and the focal length of the camera device is f millimeters. According to the perspective projection principle, the actual physical distance d = (h × pixel distance × sensor size)÷(f × image resolution). Among them, the sensor size is the actual size of the camera device sensor. For example, the width of the sensor is w millimeters and the height is h_s millimeters. Assuming that the image is captured according to the aspect ratio of the sensor width and height, then for the horizontal direction calculation, the sensor size is taken as w millimeters, and for the vertical direction or comprehensive calculation, it can be converted according to the actual situation.
[0134] Step S1533, if the actual physical distance is greater than the preset threshold, it is determined that the target object has undergone violent movement or occlusion, and the execution of the current tracking instruction is paused.
[0135] Assume that the preset threshold is 10 meters. If the calculated actual physical distance is greater than 10 meters, this means that the position of the survivor deviates significantly from the predicted position, which may be because the survivor has made a sudden rapid movement (such as running away from the UAV) or is blocked by a new landslide or other objects. At this time, the execution of the current tracking instruction is paused to avoid the UAV adjusting its flight attitude based on incorrect predictions.
[0136] Step S1534, clear the historical feature data cached in the neural network model, and restart the multi-scale spatio-temporal feature extraction step from the current frame.
[0137] Step S1535, after restarting the multi-scale spatio-temporal feature extraction step, preferentially use the position parameters of the latest bounding box as the initial input, skipping the dependence on historical feature data.
[0138] In this embodiment, the neural network model caches some historical feature data during the previous tracking process, such as the intermediate results of the global motion features and local texture features calculated previously. Clear all these cached data, and then start from the current frame image to re - extract multi - scale spatio - temporal features. After re - executing the multi - scale spatio - temporal feature extraction step, preferentially use the position parameters of the latest bounding box as the initial input and skip the dependence on historical feature data. For example, when recalculating the local texture features, intercept the local area centered on the position of the latest bounding box instead of relying on the position of the previous local area; when calculating the global motion features, use the pixel points near the position of the latest bounding box as reference points for calculating the optical flow field matrix, etc., so as to more quickly and accurately adapt to the sudden change of the survivor's position or respond to re - tracking after occlusion.
[0139] In a possible implementation manner, step S154 includes: Step S1541, with the position of the verified latest bounding box as the center, expand the interception range according to a preset scaling factor to form a new local area.
[0140] In this embodiment, assume that the upper - left vertex coordinate of the position of the verified latest bounding box is (x3, y3), the lower - right vertex coordinate is (x4, y4), and the preset scaling factor is 1.5. Then the boundary calculation of the new local area in the horizontal direction is: the new left boundary is x3 - ((x4 - x3)×0.25), and the new right boundary is x4 + ((x4 - x3)×0.25); the boundary calculation in the vertical direction is: the new upper boundary is y3 - ((y4 - y3)×0.25), and the new lower boundary is y4 + ((y4 - y3)×0.25). In this way, a new local area centered on the position of the latest bounding box and with an expanded range is determined.
[0141] Step S1542, perform Gaussian pyramid down - sampling on the new local area to generate multi - resolution image patches.
[0142] Specifically, Gaussian pyramid down - sampling is a method of down - sampling an image through a Gaussian kernel function. First, for the new local area image, starting from the highest - resolution layer, perform a convolution operation using a Gaussian kernel (for example, the Gaussian kernel size is 5×5). This convolution operation can perform weighted summation on the pixels around each pixel according to the Gaussian kernel function to obtain the image after Gaussian filtering. Then, down - sample the filtered image according to the set down - sampling ratio (for example, 0.5), that is, take every other pixel to obtain the image of the next - layer resolution. Repeat this process to generate image patches of multiple resolution levels, and each image patch is one layer in the multi - resolution image patches.
[0143] Step S1543: Apply the Histogram of Oriented Gradients (HOG) algorithm to the multi-resolution image blocks respectively to extract the spatial gradient distribution features at different resolutions.
[0144] For example, for each resolution of the image blocks, the HOG algorithm calculates the gradient direction and gradient magnitude of each pixel in the image. When calculating the gradient direction, it is determined by calculating the gray-scale change rate of the pixel in the horizontal and vertical directions. For example, for a certain pixel (x, y), the gray-scale change rate in the horizontal direction is Gx, and the gray-scale change rate in the vertical direction is Gy. Then the gradient direction θ = arctan(Gy / Gx). The gradient magnitude is sqrt(Gx^2 + Gy^2). Then the image is divided into several small regions (such as 8×8 small regions), and the histogram of the gradient direction is statistically calculated in each small region. The gradient direction is divided into several intervals (such as 0° - 20°, 20° - 40°, etc.), and the cumulative value of the gradient magnitude in each interval is statistically calculated. In this way, the histogram of oriented gradients of each small region is obtained, and the combination of the histograms of oriented gradients of all small regions is the spatial gradient distribution feature of the image block at this resolution.
[0145] Step S1544: Superimpose the spatial gradient distribution features at different resolutions according to weights to generate the updated local texture features. At the same time, re-perform the calculation of the optical flow field matrix within the expanded intercepted range and update the direction distribution histogram of the global motion features.
[0146] For example, the spatial gradient distribution features at different resolutions have different importance for representing local textures. Therefore, different weights are assigned to the spatial gradient distribution features at each resolution. Assume that the weight of the highest resolution is 0.5, the weight of the next lower resolution is 0.3, and the weight of the next lower resolution is 0.2. For each corresponding feature component, multiply the feature components at different resolutions by their respective weights and then add them. For example, a certain component of the spatial gradient distribution feature at the highest resolution is f1, the corresponding component at the next lower resolution is f2, and the corresponding component at the next lower resolution is f3. Then the updated local texture feature of this component is f1×0.5 + f2×0.3 + f3×0.2. Calculate all the feature components in this way to obtain the updated local texture features.
[0147] In a possible implementation manner, step S155 includes: Step S1551: Record the change trend of the actual physical distance during each iteration.
[0148] Step S1552, if the decreasing rate of the actual physical distance is less than the convergence threshold in a continuous preset number of iterations, it is determined that the tracking has reached a steady state. In this tracking steady state, reduce the sampling frequency of the camera device and decrease the operation levels of the multi-scale spatio-temporal feature extraction.
[0149] Step S1553, when the actual physical distance continuously remains within the preset threshold for a preset duration, activate the tracking completion flag, and stop the feature update iteration according to the tracking completion flag, and maintain the current flight attitude parameters until a new visual tracking task is received.
[0150] For example, within the expanded intercepted range, similar to the previous calculation of the optical flow field matrix, an optical flow field matrix is constructed based on the pixel differences between adjacent frames. For each frame image in the new local area, calculate the position change of each pixel point in the adjacent frame to obtain the motion components in the horizontal and vertical directions, thereby constructing the optical flow field matrix. Then, use the multi-resolution pyramid algorithm to perform hierarchical analysis on the optical flow field matrix to update the direction distribution histogram of the global motion features. Specifically, perform Gaussian pyramid generation operation on the optical flow field matrix to obtain structures at multiple resolution levels. Starting from the highest resolution level, divide the optical flow field matrix at each level into uniformly distributed square grid cells (for example, each square grid cell is 10×10 pixel size). In each square grid cell, count the direction angles of all optical flow vectors, and divide the direction angles into multiple direction intervals according to a preset angle interval (for example, 10°) to generate the direction angle distribution histogram of each square grid cell. Accumulate the direction angle distribution histograms of all square grid cells within each level across cells to obtain the motion direction distribution histogram of the current level. Perform maximum value normalization processing on the motion direction distribution histogram of each level, and splice the normalized motion direction distribution histograms of each level in a one-dimensional vector in the order from high to low resolution to form a cross-scale fusion direction distribution feature vector. This direction distribution feature vector is a part of the direction distribution histogram of the updated global motion features, and the direction distribution histogram of the updated global motion features is obtained through complete calculation.
[0151] Furthermore, during each iteration, record the change trend of the actual physical distance. The actual physical distance is obtained by converting the pixel distance between the latest bounding box position and the predicted position coordinates, and the conversion process is based on the perspective projection inversion according to the current altitude of the UAV and the focal length parameters of the camera device. In each iteration, compare the current actual physical distance with the actual physical distance in the previous iteration, calculate the difference, and thus obtain the change trend of the actual physical distance. For example, if the current actual physical distance is 8 meters and the actual physical distance in the previous iteration is 9 meters, then the actual physical distance is decreasing, and the decreasing value is 1 meter.
[0152] If the decreasing rate of the actual physical distance is less than the convergence threshold in a continuous preset number of iterations, it is determined that the tracking has entered a steady state. Assume the preset number is 5 times and the convergence threshold is 0.5 meters per iteration. In 5 consecutive iterations, calculate the decrease value of the actual physical distance for each iteration. For example, in the first iteration, the decrease is 1 meter, in the second iteration, the decrease is 0.8 meters, in the third iteration, the decrease is 0.4 meters, in the fourth iteration, the decrease is 0.3 meters, and in the fifth iteration, the decrease is 0.2 meters. The average decreasing rate is (1 + 0.8 + 0.4 + 0.3 + 0.2) ÷ 5 = 0.54 meters per iteration. Since 0.54 meters per iteration is less than the convergence threshold of 0.5 meters per iteration, it is determined that the tracking has entered a steady state. In the tracking steady state, reduce the sampling frequency of the camera device and decrease the operation levels of multi-scale spatio-temporal feature extraction. For example, the original sampling frequency of the camera device is 30 frames per second and it is reduced to 15 frames per second. For multi-scale spatio-temporal feature extraction, if previously three-layer Gaussian pyramid downsampling was performed to calculate global motion features and local texture features, now it is reduced to two layers. This can reduce the consumption of computing resources and still be able to maintain effective tracking of the target object (survivor) while the tracking is relatively stable.
[0153] When the actual physical distance continuously remains within the preset threshold for a preset duration, activate the tracking completion flag and stop the feature update iteration according to the tracking completion flag, maintaining the current flight attitude parameters until a new visual tracking task is received. Assume the preset threshold is 5 meters and the preset duration is 10 seconds. When the actual physical distance continuously remains within 5 meters for 10 seconds, for example, the actual physical distances within these 10 seconds are 4 meters, 3 meters, 3.5 meters, 4.2 meters, 3.8 meters, 4.5 meters, 4 meters, 3.2 meters, 3.6 meters, 4 meters respectively, meeting the condition of continuously remaining within 5 meters. At this time, activate the tracking completion flag. According to this tracking completion flag, stop the update iteration of the global motion features and local texture features, and the drone maintains the current flight attitude parameters, such as the current pitch angle, yaw angle, roll angle, altitude, and propeller speed, etc., until a new visual tracking task is received, such as discovering a new survivor or being required by the command center to re-adjust the tracking position, etc.
[0154] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of a neural network-based UAV vision tracking system 100 that can implement the idea of the present application provided by some embodiments of the present application. For example, the processor 120 can be used on the neural network-based UAV vision tracking system 100 and is used to execute the functions in the present application.
[0155] The neural network-based UAV vision tracking system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the neural network-based UAV vision tracking method of this application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0156] For example, the neural network-based UAV vision tracking system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the neural network-based UAV vision tracking system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of this application can be implemented according to these program instructions. The neural network-based UAV vision tracking system 100 also includes an input / output (I / O) interface 150 between the computer and other input / output devices.
[0157] For ease of explanation, only one processor is described in the neural network-based UAV vision tracking system 100. However, it should be noted that the neural network-based UAV vision tracking system 100 in this application may also include multiple processors. Therefore, the steps executed by one processor described in this application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the neural network-based UAV vision tracking system 100 executes step A and step B, it should be understood that step A and step B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.
[0158] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned neural network-based UAV vision tracking method is implemented.
[0159] It should be noted that, in order to simplify the expression of the disclosure of the present invention and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.
Claims
1. A method for visual tracking of unmanned aerial vehicles based on neural networks, characterized in that: The method comprises: The camera device carried by the drone collects a dynamic image sequence of the target object in real time, wherein the dynamic image sequence includes the motion trajectory information of the target object at consecutive moments; Performing multi-scale spatiotemporal feature extraction on the dynamic image sequence to obtain global motion features and local texture features of the target image sequence, wherein the global motion features are used to characterize the displacement changes of the target object in three-dimensional space, and the local texture features are used to characterize the spatiotemporal continuity of surface details of the target object; Inputting the global motion features and the local texture features into a pre-trained neural network model for feature fusion to generate real-time motion prediction parameters of the target object; Adjusting the flight attitude parameters of the UAV based on the real-time motion prediction parameters, and generating corresponding dynamic tracking instructions, wherein the dynamic tracking instructions are used to control the UAV to maintain a preset relative motion relationship with the target object; The UAV is driven to perform visual tracking tasks according to the dynamic tracking instructions, and the dynamic image sequence is updated in real time and fed back to the multi-scale spatiotemporal feature extraction step to form a closed-loop tracking control.
2. The neural network-based drone visual tracking method according to claim 1, characterized in that: The step of extracting multi-scale spatiotemporal features from the dynamic image sequence to obtain global motion features and local texture features of the target image sequence includes: Decomposing the dynamic image sequence into a plurality of continuous time segments, and performing inter-frame alignment processing on the image frames in each time segment to obtain a stable image sequence after eliminating jitter; Performing frame-by-frame grayscale normalization processing on the stable image sequence, and constructing an optical flow field matrix based on pixel differences between adjacent frames, wherein the optical flow field matrix is used to quantify the instantaneous motion vector of the target object in the image plane; A multi-resolution pyramid algorithm is used to perform hierarchical analysis on the optical flow field matrix, to extract motion direction distribution histograms at different scales, and the motion direction distribution histogram features of each layer are cascaded to form the global motion feature; intercepting a local area centered on the target object in the stable image sequence, and performing spatial convolution kernel filtering on the local area to extract a multi-channel texture response map; Aggregate the multi-channel texture response map along the time axis using a sliding window, and count the mean and variance of each channel response in the sliding window to generate the local texture feature; The multi-resolution pyramid algorithm is used to perform hierarchical analysis on the optical flow field matrix, extract the motion direction distribution histograms at different scales, and cascade the motion direction distribution histogram features of each layer to form the global motion feature, including: Performing a Gaussian pyramid generation operation on the optical flow field matrix to obtain a Gaussian pyramid structure including multiple resolution levels, wherein the highest resolution level corresponds to the original optical flow field matrix, and each lower resolution level is generated by step-by-step downsampling based on a Gaussian kernel; Starting from the highest resolution layer of the Gaussian pyramid structure, dividing the optical flow field matrix of each level into evenly distributed square grid cells, each square grid cell covers a preset number of optical flow vectors; Counting the direction angles of all optical flow vectors in each square grid unit, and dividing the direction angles into multiple direction intervals according to a preset angle interval, to generate a direction angle distribution histogram for each square grid unit; The direction angle distribution histograms of all square grid cells in each level are accumulated across units to obtain the motion direction distribution histogram of the current level; The maximum value normalization processing is performed on the motion direction distribution histogram of each level, and the normalized motion direction distribution histograms of each level are concatenated into one-dimensional vectors in the order of high to low resolution to form a cross-scale fused direction distribution feature vector; A principal component analysis dimensionality reduction operation is performed on the directional distribution feature vector, and the directional distribution feature vector after dimensionality reduction is weightedly fused with the average amplitude of the optical flow of each level of the Gaussian pyramid structure to generate a final vector representation of the global motion feature.
3. The neural network-based drone visual tracking method according to claim 2, characterized in that: The step of inputting the global motion features and the local texture features into a pre-trained neural network model for feature fusion to generate real-time motion prediction parameters of the target object includes: Mapping the global motion feature and the local texture feature to the high-dimensional latent space of the pre-trained neural network model respectively to obtain a corresponding first latent vector and a second latent vector; Calculating an association weight matrix between the first latent vector and the second latent vector through a cross attention mechanism, and weighted reconstructing the second latent vector according to the association weight matrix to obtain an enhanced third latent vector; Add the first latent vector and the third latent vector element by element to obtain a fused fourth latent vector; Inputting the fourth latent vector into a multi-layer perceptron for nonlinear transformation, and outputting the predicted position coordinates and velocity vector of the target object at the next moment; The heading angle adjustment amount, flight altitude correction amount and propulsion force parameter required by the UAV are calculated according to the predicted position coordinates and velocity vector to generate the real-time motion prediction parameters.
4. The neural network-based drone visual tracking method according to claim 3, characterized in that: The step of adjusting the flight attitude parameters of the UAV based on the real-time motion prediction parameters and generating corresponding dynamic tracking instructions includes: Acquire current flight status data of the UAV, wherein the current flight status data includes pitch angle, yaw angle, roll angle, altitude, and horizontal speed; Calculate the heading deviation angle in the horizontal plane and the height compensation amount in the vertical direction according to the Euclidean distance between the predicted position coordinates and the current position of the drone; adjusting the target setting value of the yaw angle based on the heading deviation angle, and adjusting the target setting value of the altitude based on the altitude compensation amount; Calculating a torque adjustment coefficient of a propulsion motor and a propeller speed increment according to a vector difference between the velocity vector and a current horizontal velocity of the UAV; The adjusted yaw angle target setting value, altitude target setting value, torque adjustment coefficient and propeller speed increment are packaged into a data packet format of the dynamic tracking instruction to generate a corresponding dynamic tracking instruction.
5. The neural network-based UAV visual tracking method according to claim 4 is characterized in that: The method of driving the UAV to perform visual tracking tasks according to the dynamic tracking instructions, and updating the dynamic image sequence in real time and feeding it back to the multi-scale spatiotemporal feature extraction step to form a closed-loop tracking control includes: Transmitting the dynamic tracking instruction to the flight control system of the UAV via a wireless communication link, triggering the flight control system to synchronously adjust the pitch angle, yaw angle, roll angle, altitude and propeller speed; During the process of the drone performing attitude adjustment, the camera device continuously collects updated dynamic image sequences and marks the latest bounding box position of the target object in the dynamic image; Performing a consistency check between the latest bounding box position and the predicted position coordinates, and if the consistency check result indicates that the position deviation exceeds a preset threshold, reinitializing the multi-scale spatiotemporal feature extraction step; updating the interception range of the local area according to the latest verified bounding box position, and recalculating the global motion feature and the local texture feature to obtain updated global motion feature and local texture feature; The updated global motion features and local texture features are input into the neural network model to iteratively generate new real-time motion prediction parameters until the relative motion relationship between the drone and the target object meets the preset tracking accuracy.
6. The neural network-based unmanned aerial vehicle visual tracking method according to claim 5, characterized in that: During the process of the drone performing attitude adjustment, continuously collecting updated dynamic image sequences through the camera device and marking the latest bounding box position of the target object in the dynamic image include: In each frame of the updated dynamic image, a search area of a preset proportion is expanded with the predicted position coordinate as the center; Performing color histogram backprojection on pixels within the search area to generate a target probability distribution map; Locate the connected area with the largest probability density in the target probability distribution map, and calculate the minimum circumscribed rectangle of the connected area; Determine the position parameters of the latest bounding box position according to the vertex coordinates of the minimum circumscribed rectangle; The position parameters and the position parameters of the historical boundary box are subjected to Kalman filtering smoothing to eliminate jitter noise.
7. The neural network-based drone visual tracking method according to claim 5, characterized in that: The step of performing a consistency check on the latest bounding box position and the predicted position coordinates, and reinitializing the multi-scale spatiotemporal feature extraction step if the consistency check result indicates that the position deviation exceeds a preset threshold, comprises: Calculating the pixel distance between the center point coordinates of the latest bounding box position and the predicted position coordinates; Convert the pixel distance into an actual physical distance, specifically by performing perspective projection inversion according to the current altitude of the drone and the focal length parameter of the camera device; If the actual physical distance is greater than the preset threshold, it is determined that the target object is in violent motion or blocked, and the execution of the current tracking instruction is suspended; Clear the historical feature data cached in the neural network model, and re-execute the multi-scale spatiotemporal feature extraction step from the current frame; After re-executing the multi-scale spatiotemporal feature extraction step, the position parameters of the latest bounding box are preferentially used as the initial input, skipping the reliance on historical feature data.
8. The neural network-based UAV visual tracking method according to claim 7, characterized in that: The updating of the interception range of the local area according to the latest verified bounding box position, and recalculating the global motion feature and the local texture feature, comprises: Taking the latest bounding box position after verification as the center, the interception range is expanded according to the preset scaling factor to form a new local area; Performing Gaussian pyramid down-sampling on the new local area to generate a multi-resolution image block; Applying a directional gradient histogram algorithm to the multi-resolution image blocks respectively to extract spatial gradient distribution features of different resolutions; The spatial gradient distribution features of different resolutions are superimposed according to weights to generate updated local texture features, and at the same time, the optical flow field matrix calculation is re-executed within the expanded interception range, and the direction distribution histogram of the global motion feature is updated.
9. The neural network-based unmanned aerial vehicle visual tracking method according to claim 7, characterized in that: The updated global motion features and local texture features are input into the neural network model to iteratively generate new real-time motion prediction parameters until the relative motion relationship between the UAV and the target object meets the preset tracking accuracy, including: During each iteration, recording a changing trend of the actual physical distance; If the decreasing rate of the actual physical distance in a preset number of consecutive iterations is less than a convergence threshold, it is determined that a tracking steady state has been entered, and in the tracking steady state, the sampling frequency of the camera device is reduced, and the operation level of the multi-scale spatiotemporal feature extraction is reduced; When the actual physical distance continues to remain within the preset threshold for a predetermined period of time, the tracking completion flag is activated, and the feature update iteration is stopped according to the tracking completion flag, and the current flight attitude parameters are maintained until a new visual tracking task is received.
10. A UAV visual tracking system based on neural network, characterized in that: The neural network-based drone visual tracking system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the neural network-based drone visual tracking method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Numerical control machine tool workpiece anomaly detection method and device based on machine vision
CN110728655A
Unmanned aerial vehicle target tracking method and system, medium, equipment and terminal
CN115619829A
Unmanned aerial vehicle ground target tracking method based on improved twin neural network
CN116189019A
Micro unmanned aerial vehicle target identification and tracking control method based on computer vision
CN117036989A
Target tracking method and system for unmanned aerial vehicle, unmanned aerial vehicle gimbal, and unmanned aerial vehicle
WO2024051574A1
Cited By
Tracking method, device and equipment for self-balancing robot and medium
CN120742909A
Unmanned aerial vehicle control system based on monocular vision
CN121297815A