Monocular vision-based unmanned aerial vehicle control system

By using a monocular vision-based UAV control system, combined with laser rangefinders and inertial data, the UAV can accurately track and strike targets, solving the problems of system complexity and high cost in existing technologies, and improving positioning accuracy and system efficiency.

CN121297815BActive Publication Date: 2026-04-03UNIV OF CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing multi-rotor UAV control systems have complex system frameworks, high costs, and large computational loads. Furthermore, their monocular vision positioning accuracy is insufficient, which limits their application scope and efficiency in complex environments.

Method used

A monocular vision-based UAV control system is adopted, which utilizes a monocular acquisition device, a laser rangefinder, a flight control module, a vision perception module, and a control module, combined with deep neural networks and inertial data, to achieve accurate distance and position determination between the UAV and the target object. The laser rangefinder and inertial data assist in positioning, reducing the dependence on the RTK module.

Benefits of technology

This technology improves the accuracy of UAVs in tracking and engaging targets without using an RTK module, reduces system costs and computational load, simplifies the hardware and software framework, and enhances positioning accuracy and system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121297815B_ABST
    Figure CN121297815B_ABST
Patent Text Reader

Abstract

This invention provides a monocular vision-based unmanned aerial vehicle (UAV) control system, comprising: a monocular acquisition device acquiring target images; a laser rangefinder detecting a first distance between the UAV and the ground; a flight control module detecting a second distance between the UAV and the ground and inertial data, and controlling the deployment of a servo motor based on control signals or controlling the UAV to track based on the tracking position; a visual perception module performing target detection on the target image using a deep neural network; extracting multi-scale fusion features of image features using a decoder; extracting inertial features of inertial data using a first sensor; fusing and scaling the multi-scale fusion features and inertial features to obtain a third distance; and a control module determining the target distance between the UAV and the target object based on the first, second, and third distances, determining the target position of the target object based on the target distance, target detection results, and coordinate mapping relationship, and outputting control signals or tracking positions based on the target position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) control technology, and more specifically, to a UAV control system based on monocular vision. Background Technology

[0002] In practical applications, there is a need to use multi-rotor drone control systems to control drones to track or strike target objects.

[0003] Existing multi-rotor UAV control systems combine a Real-Time Kinematic (RTK) module from a Real-Time Differential Global Positioning System (GPS) with a binocular vision module to improve the UAV's positioning accuracy and spatial awareness. However, the use of the RTK module makes the system framework complex, costly, and places high demands on the communication link. Compared to a monocular vision module, the binocular vision module has drawbacks such as high cost and high computational resource requirements. On the other hand, replacing the binocular vision module with a monocular vision module results in insufficient target positioning accuracy.

[0004] Therefore, the existing multi-rotor UAV control systems suffer from complex system frameworks, high costs, large computational loads, and insufficient monocular vision positioning accuracy, which are technical problems that urgently need to be solved. Summary of the Invention

[0005] In view of this, the present invention provides a drone control system based on monocular vision.

[0006] Embodiments of the present invention provide a monocular vision-based unmanned aerial vehicle (UAV) control system, the system comprising:

[0007] The system comprises: a monocular acquisition device for acquiring target images including the target object; a laser rangefinder for detecting the first distance between the UAV and the ground; a flight control module for detecting the second distance between the UAV and the ground, the UAV's inertial data, and controlling the deployment of the servo motor based on control signals or controlling the UAV to track the target object based on the tracking position; a visual perception module for using a deep neural network to perform target detection on the target image and obtain the target detection result; using a decoder to extract multi-scale fusion features of the image features output by the feature extraction layer of the deep neural network; using a first sensor to extract inertial features of inertial data for a predetermined duration; fusing and scaling the multi-scale fusion features and inertial features to obtain the third distance; and a control module for determining the target distance between the UAV and the target object based on the first, second, and third distances, and determining the target position of the target object in the navigation coordinate system corresponding to the servo motor based on the target distance, the target detection result, and the coordinate mapping relationship between the monocular acquisition device and the servo motor deployment device, and outputting control signals or tracking positions based on the target position.

[0008] The monocular vision-based UAV control system provided by this invention can accurately determine the target distance between the UAV and a target object without using an RTK module. This is achieved by using a control module based on a first distance from the ground detected by a laser rangefinder, a second altitude of the UAV from the ground detected by a flight control module, and a third distance output by a visual perception module based on a target image acquired by a monocular acquisition device and inertial data output by the flight control module. Furthermore, the control module accurately determines the target position of the target object in the navigation coordinate system corresponding to the servo deployment device based on the target distance, the target detection results output by the visual perception module, and the coordinate mapping relationship between the monocular acquisition device and the servo deployment device. This allows the control module to output highly accurate control signals or tracking positions based on the target position, enabling the flight control module to precisely control the deployment of the servo deployment device or precisely control the UAV to track the target object based on the tracking position. This achieves precise tracking or engagement of the target object. The system has a simple hardware and software framework, low cost, low computational load, and high monocular vision positioning accuracy. Attached Figure Description

[0009] The above and other objects, features and advantages of the present invention will become more apparent from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0010] Figure 1 A schematic diagram of a monocular vision-based unmanned aerial vehicle (UAV) control system according to an embodiment of the present invention is shown.

[0011] Figure 2 A schematic diagram of the structure of a visual perception module according to an embodiment of the present invention is shown;

[0012] Figure 3 A flowchart of a monocular vision-based unmanned aerial vehicle (UAV) control method according to an embodiment of the present invention is shown;

[0013] Figure 4 A flowchart of a monocular vision-based unmanned aerial vehicle (UAV) control method according to another embodiment of the present invention is shown. Detailed Implementation

[0014] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0015] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0016] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0017] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0018] In the embodiments of this invention, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to maintain the security of user personal information and network security.

[0019] Existing multi-rotor UAV control systems utilize RTK modules to receive differential signals from ground base stations to achieve precise target tracking and strike missions, further improving the UAV's positioning accuracy to the centimeter level. Binocular vision modules enhance the UAV's spatial perception capabilities. However, RTK modules place high demands on onboard hardware performance, resulting in a complex and costly hardware-software system framework with stringent communication link requirements. Binocular vision modules, compared to monocular vision modules, suffer from higher costs and computational resource requirements. When replacing binocular vision modules with traditional monocular vision modules, the latter exhibits significant shortcomings in target positioning accuracy, struggles to directly acquire target depth information, and is prone to target loss or misidentification in complex scenarios, impacting system reliability and real-time performance, thus limiting the UAV's application in high-precision tracking and strike missions. These issues also increase the UAV's weight and power consumption, limiting its endurance and applicability.

[0020] Therefore, the existing multi-rotor UAV control systems suffer from complex system frameworks, high costs, large computational loads, and insufficient monocular vision positioning accuracy, which are pressing technical problems that need to be solved. These problems limit the application scope and efficiency of multi-rotor UAV control systems in complex environments.

[0021] In view of this, embodiments of the present invention provide a monocular vision-based unmanned aerial vehicle (UAV) control system, which can be applied to the field of UAV control technology.

[0022] Figure 1 A schematic diagram of a monocular vision-based unmanned aerial vehicle (UAV) control system according to an embodiment of the present invention is shown.

[0023] like Figure 1 As shown, the monocular vision-based UAV control system 100 may include a monocular acquisition device 110, a laser rangefinder 120, a flight control module 130, a visual perception module 140, and a control module 150. The visual perception module 140 and the control module 150 are both data processing modules within an onboard computer.

[0024] The monocular acquisition device 110 can be used to acquire target images including target objects. The number of target objects in the target image can be 1, 2, 3, 4, or 5, etc. The target objects can be people, vehicles, etc. Therefore, the monocular acquisition device 110 can acquire target images including at least one target object.

[0025] For example, the monocular acquisition device 110 can be a monocular camera. The monocular camera model could be an IMX219. The IMX219 camera has a high resolution of 8 megapixels, providing clear image data and offering low-cost, high-quality input for subsequent image processing and target recognition. Compared to the binocular modules used in related technologies, the monocular acquisition device 110 has advantages such as lower cost, lighter weight, and lower computational resource requirements.

[0026] For example, the monocular acquisition device 110 can be connected to the camera serial interface (CSI) in the airborne computer, and transmit the acquired target image to the airborne computer in real time through the CSI, ensuring high-speed transmission and stable reading of the target image data.

[0027] The laser rangefinder 120 can be used to detect the initial distance of a drone from the ground.

[0028] For example, the laser rangefinder 120 can be connected to the Universal Asynchronous Receiver / Transmitter (UART) interface of the flight control module 130, and send the detected high-precision first distance to the flight control module 130 via the UART to assist the flight control module 130 in flight control.

[0029] The flight control module 130 can be used to detect the second distance between the UAV and the ground, the inertial data of the UAV, and to control the deployment of the servo release device 160 according to the control signal or to control the UAV to track the target object according to the tracking position.

[0030] According to an embodiment of the present invention, the monocular vision-based unmanned aerial vehicle control system 100 may further include a servo motor deployment device 160.

[0031] For example, flight control module 130 can be a PX4 flight control module. The PX4 flight control module offers high flexibility and scalability, adapting to various sensors and actuators to ensure stable and reliable flight control. Flight control module 130 is a core component of flight control.

[0032] For example, the flight control module 130 may include an inertial measurement unit (IMU). The IMU can be used to detect the inertial data of the UAV. The inertial data includes linear acceleration data and angular velocity data collected by the IMU at the same time. The inertial data can be represented as a data sequence. .in, The linear acceleration data is at time t. Here are the angular velocity data at time t.

[0033] The flight control module 130 can communicate with the onboard computer based on a lightweight communication protocol (Micro Air Vehicle Link, MAVLink). For example, the flight control module 130 and the onboard computer can use Ethernet as the underlying communication medium and MAVLink as the communication protocol to encapsulate and parse the transmitted message data to achieve communication.

[0034] For example, the servo-launching device 160 may include a servo, a mechanical structure, and a payload delivery device. The servo is mounted on the bottom of the UAV and connected to the payload delivery device via the mechanical structure. The flight control module 130 receives control signals from the onboard computer and generates corresponding pulse width modulation (PWM) signals to control the rotation of the servo, thereby achieving automatic payload delivery from the payload delivery device. This structural design ensures the accuracy and reliability of the delivery, enabling precise release of the payload at a specified time and location.

[0035] The visual perception module 140 can be used to perform target detection on a target image using a deep neural network to obtain the target detection result; extract multi-scale fusion features of the image features output by the feature extraction layer of the deep neural network using a decoder; extract inertial features of inertial data for a predetermined duration using a first perceptron; and fuse and scale-transform the multi-scale fusion features and inertial features to obtain a third distance. The third distance is the distance between the UAV and the ground obtained by the visual perception module 140.

[0036] According to an embodiment of the present invention, the target object is located on the ground. The first distance, the second distance, and the third distance can also be the distances between the UAV and the target object in the direction perpendicular to the ground, obtained using different methods.

[0037] For example, a decoder can be used to extract multi-scale features from the image features output by the feature extraction layer of a deep neural network, resulting in multi-scale fused features. A first perceptron can be used to extract features from inertial data of a predetermined duration, yielding inertial features.

[0038] For example, the deep neural network could be the YOLOv8 network corresponding to the object detection algorithm You Only Look Once version 8 (YOLOv8). The decoder could be a U-shaped network (UNet) architecture in deep learning. The first perceptron could be a multilayer perceptron (MLP).

[0039] According to an embodiment of the present invention, the visual perception module 140 performs target detection on the target image using a YOLOv8 network to obtain target detection results. This enables efficient analysis of the target image acquired by the monocular acquisition device 110 and accurate identification of the target detection result corresponding to the target object. The target detection result may include the position information of the target object in the target image.

[0040] According to an embodiment of the present invention, the visual perception module 140 performs target detection on the target image using the YOLOv8 network to obtain the target detection result, performs multi-scale feature extraction on the image features output by the feature extraction layer of the deep neural network using the UNet decoder to obtain multi-scale fusion features, performs feature extraction on the inertial data of a predetermined duration using the first perceptron to obtain inertial features, and performs fusion and scale transformation on the multi-scale fusion features and inertial features to obtain the third distance, thereby enabling spatial positioning and semantic recognition of the target object.

[0041] According to embodiments of the present invention, a scene distance map can be obtained by fusing multi-scale fusion features and inertial features and performing scale transformation. Based on the pixel position of the target object in the scene distance map, a third distance is obtained, achieving pixel-level distance estimation. The pixel position of the target object in the scene distance map can be obtained based on the target detection results.

[0042] The control module 150 can determine the target distance between the UAV and the target object based on the first distance, the second distance, and the third distance. It can also determine the target position of the target object in the navigation coordinate system corresponding to the servo deployment device based on the target distance, the target detection result, and the coordinate mapping relationship between the monocular acquisition device and the servo deployment device. Finally, it can output control signals or tracking positions based on the target position. The target distance is the distance between the UAV and the target object in the direction perpendicular to the ground.

[0043] According to an embodiment of the present invention, the airborne computer, as the "brain" of the entire monocular vision-based unmanned aerial vehicle (UAV) control system 100, plays a crucial role. As the "brain" of the entire system, the airborne computer communicates with the flight control module 130, indirectly controlling the UAV's flight behavior and logic. The airborne computer is equipped with a high-concurrency system based on coroutines and asynchronous programming, enabling the visual perception module 140 and the control module 150 to process data in parallel, greatly improving the efficiency of data processing and task execution.

[0044] For example, the control module 150 can interact with the flight control module 130 through the application programming interface (API) in the software development kit (MAVSDK) for UAV application development based on the MAVLink communication protocol, so as to send precise control commands to the flight control module 130, thereby directing the UAV and other actuators to complete a series of tasks such as high-precision target tracking, positioning and attack.

[0045] The monocular vision-based UAV control system provided by this invention can accurately determine the target distance between the UAV and a target object without using an RTK module. This is achieved by using a control module based on a first distance from the ground detected by a laser rangefinder, a second altitude of the UAV from the ground detected by a flight control module, and a third distance output by a visual perception module based on a target image acquired by a monocular acquisition device and inertial data output by the flight control module. Furthermore, the control module accurately determines the target position of the target object in the navigation coordinate system corresponding to the servo deployment device based on the target distance, the target detection results output by the visual perception module, and the coordinate mapping relationship between the monocular acquisition device and the servo deployment device. This allows the control module to output highly accurate control signals or tracking positions based on the target position, enabling the flight control module to precisely control the deployment of the servo deployment device or precisely control the UAV to track the target object based on the tracking position. This achieves precise tracking or engagement of the target object. The system has a simple hardware and software framework, low cost, low computational load, and high monocular vision positioning accuracy.

[0046] The monocular vision-based UAV control system 100 provided according to embodiments of the present invention aims to improve data processing efficiency and operational flexibility. The core of the monocular vision-based UAV control system 100 revolves around a shared message queue, which facilitates efficient data exchange between various modules and devices. All data receiving tasks from sensors, real-time data generated by the visual perception module 140 and the monocular acquisition device 110 are integrated into this message queue, ensuring data consistency and timeliness. Based on this comprehensive information, the control module 150 can make precise decisions and generate control commands to guide the UAV in completing specific tasks such as takeoff, tracking, and target engagement.

[0047] According to an embodiment of the present invention, the control module 150 is further configured to: obtain a first distance, a second distance, a third distance, a target detection result, and a coordinate mapping relationship between the monocular acquisition device and the servo deployment device from the message queue. The visual perception module 140 is further configured to obtain a target image and inertial data from the message queue. The first distance, the second distance, the third distance, the target image, the inertial data, and the target detection result in the message queue all have the same timestamp.

[0048] According to an embodiment of the present invention, by using a control module to obtain the first distance, second distance, third distance, target detection result, and coordinate mapping relationship between the monocular acquisition device and the servo motor deployment device from the message queue, and using a visual perception module to obtain the target image and inertial data from the message queue, the first distance, second distance, third distance, target image, inertial data, and target detection result in the message queue have the same timestamp, enabling the control module and the visual perception module to quickly obtain data with the same timestamp from the message queue, ensuring that the data is aligned in time, and thus ensuring that the control signal or tracking position corresponding to the same timestamp can be accurately obtained subsequently.

[0049] According to an embodiment of the present invention, the monocular vision-based UAV control system 100 may further include a message queue and serialization module. The message queue and serialization module is designed to achieve efficient data management and processing. This module includes a message queue, timed serialization tasks, and a unified interface to ensure seamless data exchange and storage among the various components within the system.

[0050] According to an embodiment of the present invention, the message queue, as a core component, supports the operation of the UAV by integrating various information received from the flight control module 130, such as flight model, unlock status, inertial data, second distance, quaternions and position translation information corresponding to the navigation coordinate system, and the first distance obtained by the laser rangefinder 120. This information is acquired in real time and written to the message queue via MAVSDK, ensuring that all tasks relying on this data can access the latest and most accurate information. Furthermore, real-time data generated by the visual perception module 140 and the monocular acquisition device 110, such as target detection results, third distance, and target images, are also integrated into the message queue for use by other modules within the system.

[0051] According to an embodiment of the present invention, the timed serialization task is an independently running task used to periodically access all data in the message queue, obtain the latest data, and directly serialize and store it in a predetermined file. The data storage format of the predetermined file can be, for example, a bag. Unlike the traditional telemetry plugin approach, the timed serialization task directly utilizes the comprehensive information in the message queue for serialization processing, ensuring data consistency and integrity. This mechanism not only meets the system's needs for real-time data processing but also enables effective management of historical data, allowing users to easily review various information during flight and providing rich historical data resources for subsequent data analysis. The interval of the timed serialization can be flexibly adjusted according to specific application scenarios to adapt to different data storage requirements.

[0052] To improve system maintainability and scalability, the entire system adopts a unified interface design, combined with a lightweight message publishing and subscription mechanism. Both data subscribers (such as control module 150) and publishers (such as visual perception module 140 and monocular acquisition device 110) interact through the same set of APIs. This design simplifies the communication process between modules, reduces interface complexity, and improves the system's flexibility and scalability.

[0053] The message publishing operation involves each sensor, vision perception module 140, and monocular acquisition device 110 publishing the collected data to the message queue. All data follows a unified format for easy subsequent processing. The message subscription operation involves the control module 150, flight control module 130, vision perception module 140, and other modules requiring data subscribing to the necessary information through a unified interface. This includes not only flight control data but also real-time analysis results from the vision perception module 140. The serialization and storage operation involves a timed serialization task retrieving all relevant data from the message queue at preset time intervals, serializing it, and storing it in a bag file for subsequent analysis and use.

[0054] According to embodiments of the present invention, this integrated data management design simplifies the system's software architecture, allowing developers to easily add new functional modules or improve existing modules through simple interface calls without significantly modifying the overall system software structure. Simultaneously, it enhances the system's response speed, accuracy, and scalability, providing strong support for the development and application of UAV technology.

[0055] Figure 2 A schematic diagram of the structure of a visual perception module according to an embodiment of the present invention is shown.

[0056] like Figure 2 As shown, the visual perception module 140 may include a deep neural network 141, a decoder 142, a first perceptron 143, a second perceptron 144, and a fully connected network 145. The deep neural network 141 may include a feature extraction layer 1411, a feature fusion layer 1412, and a decision layer 1413.

[0057] According to an embodiment of the present invention, the image features include a first image feature, a second image feature, and a third image feature. The first image feature is obtained based on the target image, the second image feature is obtained based on the first image feature, and the third image feature is obtained based on the second image feature.

[0058] For example, the feature extraction layer 1411 can extract features from the target image to obtain a first image feature E1, extract features from the first image feature E1 to obtain a second image feature E2, and extract features from the second image feature E2 to obtain a third image feature E3. The sizes of the first image feature E1, the second image feature E2, and the third image feature E3 decrease sequentially.

[0059] The feature fusion layer 1412 can perform feature fusion and enhancement on the third image feature E3 to obtain fused enhanced features. For example, the feature fusion layer 1412 can perform top-down and bottom-up feature fusion on the third image feature E3 to obtain fused enhanced features. The decision layer 1413 can map the fused enhanced features to obtain the target detection result. The target detection result includes the 2D (2D) bounding box information and class probability of the target object in the target image.

[0060] The decoder 142 includes a first upsampling layer 1421, a first fusion layer 1422, a second upsampling layer 1423, a second fusion layer 1424, a third upsampling layer 1425, and a third fusion layer 1426.

[0061] The visual perception module 140 uses the decoder 142 to perform multi-scale feature extraction on the image features output by the feature extraction layer 1411 of the deep neural network 141, obtaining multi-scale fused features, including: using a first upsampling layer 1421 to perform bilinear upsampling on the third image feature E3 multiple times to obtain a first upsampled feature; using a first fusion layer 1422 to concatenate the first upsampled feature and the second image feature E2 through channels to obtain a first concatenated feature; and performing convolution, normalization, and nonlinear operations on the first concatenated feature multiple times to obtain a first fused feature; using a second upsampling layer 1423 to perform bilinear upsampling on the first fused feature multiple times to obtain a second upsampled feature; using a second fusion layer 1424 to concatenate the second upsampled feature and the first image feature E1 through channels to obtain a second concatenated feature; and performing convolution, normalization, and nonlinear operations on the second concatenated feature multiple times to obtain a second fused feature; and using a third upsampling layer 1425 to perform bilinear upsampling on the second fused feature multiple times to obtain a third upsampled feature. By using the third fusion layer 1426 to perform convolution, normalization and nonlinear operations on the third upsampled features multiple times, multi-scale fusion features are obtained.

[0062] For example, the first upsampling layer 1421 can perform bilinear upsampling on the third image feature E3 twice consecutively to obtain the first upsampled feature. That is, after the first upsampling layer 1421 performs bilinear upsampling on the third image feature E3 once, it performs bilinear upsampling on the obtained feature again to obtain the first upsampled feature. The first fusion layer 1422 can perform convolution, normalization, and nonlinear operations on the first stitched feature twice consecutively to obtain the first fused feature. That is, after the first fusion layer 1422 performs convolution, normalization, and nonlinear operations on the first stitched feature, it performs convolution, normalization, and nonlinear operations on the obtained feature again to obtain the first fused feature. This process can be repeated to understand the specific operations for performing bilinear upsampling on features multiple times consecutively, as well as the operations for performing convolution, normalization, and nonlinear operations on features multiple times consecutively. The nonlinear operation can be implemented based on the Rectified Linear Unit (ReLU).

[0063] The second upsampling layer 1423 performs bilinear upsampling twice on the first fused feature to obtain the second upsampled feature. The second fusion layer 1424 performs convolution, normalization, and nonlinear operations twice on the second concatenated feature to obtain the second fused feature. The third upsampling layer 1425 performs bilinear upsampling twice on the second fused feature to obtain the third upsampled feature. The third fusion layer 1426 performs convolution, normalization, and nonlinear operations twice on the third upsampled feature to obtain the multi-scale fusion feature. The multi-scale fusion feature is a high-resolution deep feature.

[0064] The inertial features are obtained by extracting features from the inertial data of a predetermined duration using the first perceptron 143, including: extracting features from the inertial data of a predetermined duration using the first perceptron 143 to obtain initial inertial features; and then fully connecting the initial inertial features using a fully connected network 145 to obtain the final inertial features.

[0065] The visual perception module 140 fuses and scales multi-scale fusion features and inertial features to obtain a third distance, including: concatenating channels of the multi-scale fusion features and inertial features to obtain a third concatenated feature; continuously performing convolution and nonlinear operations on the third concatenated feature to obtain a relative depth feature; pre-integrating inertial data for a predetermined duration to obtain a depth change; using the second perceptron 144 to perform regression calculation on the depth change to obtain a scale factor; and obtaining the third distance based on the scale factor and the relative depth feature. The scale fusion features and inertial features have the same size and the same number of channels.

[0066] For example, the scale factor can be multiplied by the relative depth feature to obtain the third distance.

[0067] According to an embodiment of the present invention, the relative depth feature is the distance of the UAV from the ground predicted by the visual perception module 140. The scale factor reflects the multiple conversion relationship between the actual distance of the UAV from the ground and the distance of the UAV from the ground predicted by the visual perception module 140. Since monocular visual depth is scale-free, the relative depth feature differs significantly from the actual distance of the UAV from the ground. Therefore, the scale factor can be used to calibrate the dimensions of the relative depth feature to obtain a more accurate third distance.

[0068] The visual perception module 140 continuously performs convolution and non-linear operations on the third stitched feature to obtain the relative depth feature, including: continuously performing 1×1 convolution on the third stitched feature ( Figure 2 (represented by 1×1 Conv in Chinese), 3×3 convolution ( Figure 2 The relative depth features are obtained by using 3×3 Conv, 1×1 convolution, and nonlinear operations.

[0069] For example, nonlinear operations can be implemented based on the Rectified Linear Unit (ReLU) function. Figure 2 (ReLu is used in Chinese).

[0070] According to an embodiment of the present invention, after using a first perceptron 143 to extract features from inertial data of a predetermined duration to obtain inertial features, and using a decoder 142 to extract multi-scale features from the image features output by the feature extraction layer 1411 of the deep neural network to obtain multi-scale fused features, the multi-scale fused features and inertial features are channel-stitched using a visual perception module 140 to obtain third stitched features. The third stitched features are then continuously convolved and subjected to nonlinear operations to obtain relative depth features. Pre-integration is performed on the inertial data of the predetermined duration to obtain depth change. The second perceptron 144 performs regression calculation on the depth change to obtain a scale factor. Based on the scale factor and the relative depth features, a third distance is obtained. This method can fuse inertial data obtained from the IMU and multi-scale fused features, perform scale calibration on the relative depth features, and obtain a relatively accurate third distance, thus solving the problem of scale-free depth in monocular vision.

[0071] According to embodiments of the present invention, the visual perception module 140 not only possesses excellent detection performance but also operates stably in complex environments. In terms of both real-time performance and accuracy, this module provides strong support for UAV tracking and strike missions, ensuring the efficient execution and successful completion of subsequent tracking and strike tasks. The visual perception module 140 is a key component of the UAV system, significantly enhancing the overall system's intelligence level and operational flexibility. Based on the fusion technology of deep learning and inertial data, the visual perception module 140 can improve the accuracy of third-distance estimation and output complete three-dimensional positioning information of the target object in complex environments.

[0072] According to embodiments of the present invention, extensive data augmentation techniques are employed during training to enhance the robustness and adaptability of the various models in the visual perception module 140. These include background replacement, which enhances the model's understanding of complex scenes by placing the target object in different environmental backgrounds. Simultaneously, methods such as random cropping, scaling, color jittering, horizontal flipping, and rotation are applied to increase the diversity of the dataset and reduce the risk of overfitting. In particular, hybrid augmentation techniques (such as MixUp and Mosaic) are also used to further enrich the training samples and improve the model's learning efficiency. The application of pre-trained weights and the optimization of the custom loss function, along with the learning rate scheduling strategy, collectively promote faster and more stable model convergence.

[0073] According to an embodiment of the present invention, the target detection result corresponding to the target location includes detection box information corresponding to the target location and a semantic category. The detection box information may be the 2D bounding box information of the target object in the target image.

[0074] For example, the detection bounding box information can be represented as .in, Let x be the x-coordinate of the target corner point of the i-th detection box. Let be the ordinate of the target corner point of the i-th detection box. Let be the width of the i-th detection box. Let be the height of the i-th detection box. The target corner point can be the top left, top right, bottom left, or bottom right corner of the i-th detection box.

[0075] According to an embodiment of the present invention, and Given the coordinates of the target corner point in the pixel coordinate system of the monocular acquisition device 110, the coordinate mapping relationship between the monocular acquisition device 100 and the servo motor deployment device can be used to... and Convert to coordinates in the navigation coordinate system to precisely control the drone based on the coordinates in the navigation coordinate system.

[0076] According to an embodiment of the present invention, coordinate transformation is key to achieving automatic tracking and strike in the monocular vision-based UAV control system 100. The monocular vision-based UAV control system 100 includes a pixel coordinate system where the monocular acquisition device 110 is located, a camera coordinate system, a servo coordinate system corresponding to the servo deployment device, and a navigation coordinate system.

[0077] According to an embodiment of the present invention, the pixel coordinate system has its origin at the upper left corner of the target image captured by the monocular acquisition device 110, with the positive x-axis pointing to the right and the positive y-axis pointing downwards. The camera coordinate system is typically centered on the optical center of the camera, with the optical axis pointing to the positive z-axis, and the x and y axes aligned with the pixel coordinate system. The relationship between points (x, y) in the pixel coordinate system and points (X, Y, Z) in the camera coordinate system can be described by the intrinsic parameter matrix of the monocular acquisition device 110, such as a camera. Specifically, the monocular acquisition device 110, such as a camera, can be calibrated using the Zhang Zhengyou calibration method to obtain its intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficients.

[0078] The transformation steps between the pixel coordinate system and the camera coordinate system are as follows: 1) Correct the image distortion using the camera's intrinsic parameter matrix and distortion coefficients, where the pixel coordinates after distortion correction are (x′, y′). 2) Determine the target distance z (unit: meters) between the UAV and the target object. 3) Using formula (1), based on the target distance and the focal length in the intrinsic parameter matrix... Calculate the coordinates (X,Y,Z) of each distortion-corrected pixel (x′,y′) in the camera coordinate system.

[0079] (1);

[0080] in, This represents the abscissa after distortion correction in the pixel coordinate system. X is the distorted ordinate in the pixel coordinate system, X is the x-coordinate in the camera coordinate system, Y is the ordinate in the camera coordinate system, and Y is the depth in the camera coordinate system.

[0081] According to an embodiment of the present invention, through the above steps 1) to 3) and formula (1), the coordinates in the pixel coordinate system can be accurately transformed into the camera coordinate system, providing a basis for subsequent coordinate transformation and target positioning.

[0082] According to an embodiment of the present invention, the second distance includes a fused distance and a satellite positioning distance. The satellite positioning distance is the distance between the UAV and the ground output by the satellite positioning unit included in the flight control module, and the fused distance is the distance between the UAV and the ground obtained by the flight control module based on inertial data and the satellite positioning distance. The satellite positioning unit can be, for example, a positioning unit based on a Global Navigation Satellite System (GNSS).

[0083] The control module 150 determines the target distance between the UAV and the target object based on a first distance, a second distance, and a third distance, which may include: determining the first distance as the target distance if a first condition is met; determining the third distance as the target distance if the first condition is not met but the second condition is met; determining the satellite positioning distance as the target distance if neither the first nor the second condition is met but the third condition is met; and determining the fused distance as the target distance if none of the first, second, and third conditions are met but a fourth condition is met.

[0084] According to an embodiment of the present invention, the first condition is that the laser rangefinder 120 is operating normally and the fused distance is less than the first maximum range of the laser rangefinder. The second condition is that the fused distance is less than the second maximum range of the monocular acquisition device 110 and the absolute value of the difference between the third distance and the fused distance is less than a first preset threshold. The third condition is that the absolute value of the difference between the satellite positioning distance and the fused distance is less than a second preset threshold. The fourth condition is that none of the first, second, and third conditions are met. Therefore, it can also be understood that when none of the first, second, and third conditions are met, the fused distance is determined as the target distance.

[0085] According to an embodiment of the present invention, the first maximum range characterizes the maximum effective range of the laser rangefinder 120. The second maximum range characterizes the upper limit of the optimal range of the third distance estimated by the visual perception module 140.

[0086] According to embodiments of the present invention, the first maximum range, the second maximum range, the first preset threshold, and the second preset threshold can be selected according to actual conditions, and are not limited herein. For example, the first maximum range can be 12 meters. The second maximum range can be 10 meters. The first preset threshold and the second preset threshold can both be 3 meters.

[0087] According to an embodiment of the present invention, the accuracy of the first distance is greater than the accuracy of the third distance, the accuracy of the third distance is greater than the accuracy of the satellite positioning distance, and the accuracy of the satellite positioning distance is greater than the accuracy of the fused distance. The accuracy of the fused distance is the lowest; therefore, when the fused distance is less than the first maximum range of the laser rangefinder, it is ensured that the first distance is less than the first maximum range of the laser rangefinder. Since the fused distance is less than the second maximum range of the monocular acquisition device 110, it is ensured that the third distance is less than the second maximum range of the monocular acquisition device 110.

[0088] According to an embodiment of the present invention, when the first condition is met, the first distance is determined as the target distance; when the first condition is not met but the second condition is met, the third distance is determined as the target distance; when neither the first nor the second condition is met but the third condition is met, the satellite positioning distance is determined as the target distance; and when neither the first, second, nor the third condition is met but the fourth condition is met, the fused distance is determined as the target distance. This ensures that the control module 150 can continuously output the target distance and output the target distance with relatively high accuracy.

[0089] According to an embodiment of the present invention, after passing through steps 1) to 3) and formula (1) above, and After converting to the camera coordinate system and obtaining the corresponding coordinates in the camera coordinate system, we can then convert back to the servo coordinate system based on the coordinate mapping relationship between the camera coordinate system and the servo coordinate system, and then convert back to the navigation coordinate system based on the coordinate mapping relationship between the servo coordinate system and the navigation coordinate system, to obtain the target position of the target object in the navigation coordinate system.

[0090] According to an embodiment of the present invention, the direction of the servo coordinate system is the same as the direction of the UAV's local coordinate system, and the origin is the servo mounting and deployment point. The UAV's local coordinate system is a right-handed Cartesian coordinate system with the UAV's front, left, and bottom axes being the XYZ axes, respectively.

[0091] According to an embodiment of the present invention, in actual installation, there is a certain offset between the installation positions of the servo motor and the monocular acquisition device 110. Therefore, the camera coordinate system cannot be directly applied to the servo motor deployment device 160. The coordinates in the camera coordinate system need to be translated and rotated to transform to the servo motor coordinate system. The camera coordinate system and the servo motor coordinate system have the same scale, and the coordinate transformation between the two coordinate systems is a rigid transformation.

[0092] According to embodiments of the present invention, a coordinate transformation algorithm based on multiple points can be used to achieve accurate transformation from the camera coordinate system to the servo coordinate system. The core idea of ​​the coordinate transformation algorithm is to solve for the coordinate transformation parameters by using the Singular Value Decomposition (SVD) method, based on the known coordinates of multiple points in the target coordinate system and their corresponding coordinates in the servo coordinate system.

[0093] According to an embodiment of the present invention, the coordinate mapping relationship between the monocular acquisition device 110 and the servo deployment device is determined according to the following operations: The monocular acquisition device 110 acquires a checkerboard calibration board placed on the ground to obtain a checkerboard image, and simultaneously determines the projection position of the servo deployment device on the checkerboard calibration board, defining the projection position as the origin position of the servo coordinate system corresponding to the servo deployment device. Based on the corner points of each checkerboard grid, the side length of each checkerboard grid, the origin position, and the checkerboard image, the coordinates of each checkerboard grid corner point in the servo coordinate system and the camera coordinate system where the monocular acquisition device 110 is located are determined. Based on the coordinates of each checkerboard grid corner point in the servo coordinate system and the camera coordinate system where the monocular acquisition device 110 is located, a covariance matrix is ​​constructed and decomposed to obtain rotation and translation matrices used to transform the coordinates in the camera coordinate system where the monocular acquisition device is located to the coordinates in the servo coordinate system. The quaternion and position translation information corresponding to the navigation coordinate system are obtained from the flight control module. Based on the rotation matrix, translation matrix, quaternion, and position translation information, the coordinate mapping relationship between the monocular acquisition device and the servo motor deployment device is determined.

[0094] For example, before calculating the coordinate mapping relationship between the camera coordinate system and the servo coordinate system, the UAV system can be calibrated with a gyroscope to ensure its attitude angle is 0°, i.e., parallel to the ground. Then, the UAV equipped with the monocular vision-based UAV control system 100 is fixed on a specially designed frame to ensure stability. A checkerboard calibration plate is placed on a level surface, ensuring it is flat and unobstructed, wherein the size and dimensions of the checkerboard grid on the calibration plate are known. The distance between the checkerboard calibration plate and the monocular acquisition device 110 is a predetermined distance, which is a known quantity.

[0095] When calculating the coordinate mapping relationship between the camera coordinate system and the servo coordinate system, a monocular acquisition device 110 is used to acquire a checkerboard calibration board placed on the ground to obtain a checkerboard image. Simultaneously, the projection position of the servo deployment device on the checkerboard calibration board is determined, and this projection position is defined as the origin of the servo coordinate system corresponding to the servo deployment device. The projection position of the servo deployment device on the checkerboard calibration board can be the projection position of the servo on the checkerboard calibration board.

[0096] According to an embodiment of the present invention, each corner point of each chessboard grid in the chessboard calibration plate can be used to determine the coordinate mapping relationship between the camera coordinate system and the servo coordinate system.

[0097] For example, the bottom left corner point in the checkerboard calibration board can be the first corner point, and the points can be gradually increased upwards. Then, the points can be increased to the second column. Based on this pattern, a set of points can be obtained to determine the coordinate mapping relationship between the camera coordinate system and the servo coordinate system.

[0098] The checkerboard image can be processed using corner extraction algorithms from a pre-defined image processing library to obtain the coordinates of each checkerboard corner in the camera coordinate system. The pre-defined image processing library could be, for example, OpenCV. Based on the origin of the servo coordinate system and the side length of each checkerboard square, the coordinates of each checkerboard corner in the calibration board are calculated. The coordinates of each checkerboard corner in the camera coordinate system and its coordinates in the servo coordinate system form a multi-pair coordinate correspondence. These correspondences form the basis for subsequent coordinate transformation parameter calculation.

[0099] The point cloud registration algorithm (Kabsch) can then be used to solve for the rotation matrix and translation vector through SVD decomposition.

[0100] For example, formula (2) can be used to calculate the centroid coordinates corresponding to the coordinates of the corner points of each chessboard square in the camera coordinate system. Using formula (3), calculate the centroid coordinates corresponding to the coordinates of the corner points of each chessboard square in the servo coordinate system. .

[0101] (2);

[0102] (3);

[0103] in, These are the coordinates in the camera coordinate system. Let be the coordinates in the servo coordinate system, and n be the total number of corner points on the chessboard. Number the corner points of the chessboard.

[0104] The coordinates in the camera coordinate system and the servo coordinate system can be decentered using formulas (4) and (5) respectively.

[0105] (4);

[0106] (5);

[0107] in, These are the decentralized coordinates corresponding to the camera coordinate system. These are the decentralized coordinates corresponding to the servo coordinate system.

[0108] Formula (6) can be used to determine the correct method based on the given information. and Construct the covariance matrix H. Perform SVD decomposition on the covariance matrix H using formula (7).

[0109] (6);

[0110] (7);

[0111] Where U is a 3×3 orthogonal matrix, and its column vectors are the left singular vectors of H. Σ is a 3×3 diagonal matrix, and the elements on the diagonal are the singular values ​​of H, arranged in descending order. V is a 3×3 orthogonal matrix, and its column vectors are the right singular vectors of H.

[0112] Formulas (8) and (9) can be used to calculate the rotation matrix R and translation matrix t used to transform the coordinates in the camera coordinate system where the monocular acquisition device is located to the coordinates in the servo coordinate system based on the results of SVD decomposition.

[0113] (8);

[0114] (9).

[0115] According to embodiments of the present invention, the points in the camera coordinate system can also be transformed to the servo coordinate system using the solved rotation matrix R and translation matrix t. Simultaneously, the accuracy and reliability of the transformation are evaluated by calculating the error between the transformed points and the points in the actual servo coordinate system.

[0116] According to an embodiment of the present invention, in UAV control, the local coordinate system of the UAV is usually a coordinate system with the UAV takeoff point as the origin, which is also the servo coordinate system, and its axis is consistent with the axis of the UAV body. The navigation coordinate system uses the North-East-Down (NED) coordinate system, which is a geographic coordinate system with north, east, and ground as axes, and is widely used in aviation and maritime fields.

[0117] According to an embodiment of the present invention, the flight control module 130 can acquire quaternions and position translation information corresponding to the navigation coordinate system. The quaternions and position translation information corresponding to the navigation coordinate system are simultaneously the quaternions and position translation information corresponding to the UAV and the quaternions and position translation information corresponding to the monocular vision-based UAV control system 100. Quaternions are a mathematical tool used to represent rotation in three-dimensional space, possessing advantages such as computational stability and avoidance of gimbal lock.

[0118] According to an embodiment of the present invention, the rotation matrix between the camera coordinate system and the servo coordinate system can be calculated using formula (10) based on quaternions. .

[0119] (10);

[0120] in, , , , These are the components of the quaternion.

[0121] According to embodiments of the present invention, it is possible to utilize The coordinates in the servo coordinate system are rotated to align with the axis of the NED coordinate system. Then, translation is performed based on the position translation information to obtain the final NED coordinates, which are the coordinates in the navigation coordinate system.

[0122] For example, for any coordinate in the servo coordinate system ( , , ), which can be converted to the navigation coordinate system using formula (11).

[0123] (11);

[0124] in, , and These represent the position translation information along the horizontal axis, the vertical axis, and the depth direction, respectively. , and These represent the x-coordinate, y-coordinate, and depth in the navigation coordinate system.

[0125] According to an embodiment of the present invention, a checkerboard image is obtained by acquiring a checkerboard calibration board placed on the ground using a monocular acquisition device 110. Simultaneously, the projection position of the servo motor deployment device on the checkerboard calibration board is determined, and this projection position is defined as the origin of the servo motor coordinate system corresponding to the servo motor deployment device. Based on the corner points of each checkerboard grid, the side length of each checkerboard grid, the origin position, and the checkerboard image, the coordinates of each checkerboard grid corner point in the servo motor coordinate system and the camera coordinate system where the monocular acquisition device 110 is located are determined. The coordinates in the camera coordinate system are used to construct a covariance matrix and decompose it to obtain the rotation and translation matrices used to transform the coordinates in the camera coordinate system of the monocular acquisition device to the coordinates in the servo coordinate system. The quaternion and position translation information corresponding to the navigation coordinate system are obtained from the flight control module. Based on the rotation matrix, translation matrix, quaternion and position translation information, the coordinate mapping relationship between the monocular acquisition device and the servo deployment device is determined. This prepares for accurately determining the target position of the target object based on the coordinate mapping relationship, so that the accurate target position of the target object can be obtained without using the RTK module.

[0126] According to an embodiment of the present invention, to ensure the accuracy of coordinate transformation, it is necessary to align the time when the monocular acquisition device 110 generates the target image with the time recorded in the odometer of the flight control module 130, wherein the odometer is used to record quaternions and position translation information in real time. This can be achieved by adding timestamps to the system and synchronizing them during data processing. Time alignment can reduce control latency and coordinate transformation errors.

[0127] For example, the factors affecting the conversion accuracy from the servo coordinate system to the navigation coordinate system include the following two aspects. The first aspect is the NED coordinate fusion accuracy on the flight control module 130. The flight control module 130 determines the position and attitude of the UAV by fusing data from multiple sensors such as IMU and GNSS. The IMU provides high-frequency attitude and acceleration data, while GNSS provides absolute position information. The fusion accuracy of these two directly affects the accuracy of the NED coordinates.

[0128] The second aspect involves the time alignment of the onboard computer and flight control module 130, which can also be described as the time alignment of data acquired via API and target image acquired by the monocular acquisition device 110. The control module 150 in the onboard computer acquires various information output by the flight control module 130 using APIs provided in the MAVSDK library, while the target image acquired by the monocular acquisition device 110 is used by the visual perception module 140 for target detection and tracking. The flight control module 130 performs coordinate transformation based on data from both the flight control module 130 and the visual perception module 140. This time alignment ensures that data from the same moment is used during coordinate transformation, thereby improving transformation accuracy.

[0129] According to an embodiment of the present invention, through the above coordinate transformation operation and time alignment operation, the target position of the target object in the navigation coordinate system can be accurately determined. Based on the target position, it is possible to quickly track each target object in the target area and accurately control the position of the UAV.

[0130] According to an embodiment of the present invention, when the UAV is in the target tracking and attack phase, the primary task of the monocular vision-based UAV control system 100 is to conduct a comprehensive and detailed detection of the target area where the target object is located, and to improve the accuracy of target recognition and positioning by comprehensively analyzing multi-frame target image data.

[0131] According to an embodiment of the present invention, after the UAV flies to the preset target area under the control of the monocular vision-based UAV control system 100, the monocular vision-based UAV control system 100 begins to perform tracking and attack tasks.

[0132] According to an embodiment of the present invention, the control module 150 outputs the control signal or the tracking position based on the target position, comprising: when it is determined that the target object is to be tracked, clustering at least one target position corresponding to each of multiple frames of target images to obtain at least one cluster set; obtaining at least one initial tracking position based on the at least one cluster set; when there is only one initial tracking position, determining the initial tracking position as the tracking position and outputting the tracking position; when there are multiple initial tracking positions, obtaining multiple tracking positions based on the order of the multiple initial tracking positions and outputting the multiple tracking positions. Each cluster set represents a potential tracking position.

[0133] For example, the number of target images in a multi-frame format can be 10 to 15 frames.

[0134] For example, for each cluster, the average value of all target locations within that cluster can be calculated as the final coordinate estimate, i.e., the initial tracking location.

[0135] According to embodiments of the present invention, when there is only one target object in the target region and the target object is tracked, each frame of the target image includes one target object, and each frame of the target image corresponds to one target location. Clustering can be performed on the target locations corresponding to each of the multiple frames of the target image to obtain at least one cluster set. When the number of target objects in the target region is greater than one and the target object is tracked, each frame of the target image includes at least two target objects, and each frame of the target image corresponds to at least two target locations. Clustering can be performed on the at least two target locations corresponding to each of the multiple frames of the target image to obtain at least two cluster sets.

[0136] According to an embodiment of the present invention, after obtaining multiple initial tracking positions, the UAV will sequentially fly to these initial tracking positions for detailed detection in order to track the target object in real time. Considering that the multi-rotor UAV flies at a high altitude and there are no excessively tall obstacles in the target area, the flight environment can be regarded as a simple plane. Path planning is performed within this plane based on multiple initial tracking positions to improve the efficiency and reliability of the tracking task.

[0137] According to an embodiment of the present invention, after obtaining multiple initial tracking positions, in order to simplify the path planning process and ensure efficiency, a method based on two sorting steps can be used to generate a logical path scanning from left to right. Specifically, firstly, all initial tracking positions are arranged in ascending order according to their horizontal coordinates along the horizontal axis (e.g., the X-axis) of the navigation coordinate system, forming a basic left-to-right order. Next, within each column (i.e., multiple initial tracking positions with the same horizontal coordinate), they are further arranged in ascending order according to their vertical coordinates along the vertical axis (e.g., the Y-axis) of the navigation coordinate system, ensuring that points within the same column are arranged in top-to-bottom order, resulting in multiple tracking positions. In this way, a path is generated that scans row by row from left to right, and the UAV will visit each tracking position sequentially along this path. If the last tracking position in the current column has been visited, the UAV will move to the first tracking position in the next column to continue tracking, thereby ensuring that all tracking positions are covered and each tracking position is visited only once. This method not only simplifies the complexity of path planning but also improves the efficiency and reliability of task execution.

[0138] Throughout the tracking process, the drone continuously collected and transmitted high-resolution imagery and other sensor data back to the command center for subsequent analysis. This data not only verified the existence and characteristics of the initial tracking location but also provided crucial information for developing the next course of action.

[0139] According to embodiments of the present invention, when it is determined that a target object is to be tracked, at least one target location corresponding to each of multiple target images is clustered to obtain at least one cluster set. Based on the at least one cluster set, at least one initial tracking position is obtained. This achieves depth-aware spatiotemporal adaptive clustering of at least one target location corresponding to each of multiple target images, grouping different target locations belonging to the same actual target location but identified in different frames together to form a more accurate initial tracking position. This method can effectively reduce the impact of single-frame recognition errors and improve the accuracy of initial tracking position estimation.

[0140] According to an embodiment of the present invention, the control module 150 clusters at least one target location corresponding to each of multiple target images to obtain at least one cluster set, including: for each target image frame, calculating the mean and standard deviation of multiple target locations corresponding to the target image; determining the cluster radius based on the mean, standard deviation, a first preset coefficient, and a second preset coefficient; for any target location among the multiple target locations corresponding to the target image, calculating the positional deviation between any target location and the multiple target locations corresponding to the multiple target images; and clustering the target locations adjacent to any target location among the multiple target locations corresponding to the multiple target images based on the multiple positional deviations and the cluster radius to obtain a cluster set corresponding to any target location.

[0141] According to an embodiment of the present invention, when each frame of the target image includes multiple target locations, the average value can be calculated for the multiple target locations corresponding to the target image. and standard deviation .

[0142] For example, the cluster radius can be determined based on the mean, standard deviation, first preset coefficient, and second preset coefficient according to formula (12). .

[0143] (12);

[0144] in, The first preset coefficient, This is the second preset coefficient.

[0145] According to an embodiment of the present invention, the control module 150 calculates the positional deviation between any target position and the multiple target positions corresponding to multiple frames of target images for any target position among multiple target positions corresponding to a target image, including: determining the acquisition time corresponding to each of the multiple frames of target images; determining the detection box information corresponding to each of the multiple target positions based on the target detection results corresponding to each of the multiple frames of target images; and determining the positional deviation between any target position and the multiple target positions corresponding to the multiple frames of target images based on any target position, the acquisition time corresponding to each of the multiple frames of target images, the target position, and the detection box information.

[0146] For example, the detection bounding box information can be represented as The area of ​​each detection frame can be calculated based on the detection frame information. and take To mitigate data discrepancies. Among these, the normalized area... .

[0147] For example, formula (13) can be used to determine the positional deviation between any target position and the multiple target positions corresponding to the multiple target images, based on any target position, the acquisition time corresponding to each of the multiple target images, the target position, and the detection box information. .

[0148] (13);

[0149] in, The coordinates of the navigation coordinate system corresponding to the i-th detection box. The coordinates of the navigation coordinate system corresponding to the j-th detection box. and These are time weights and scale weights, respectively. The acquisition time of the target image corresponding to the i-th detection box. The acquisition time is the time of the target object corresponding to the j-th detection box.

[0150] According to an embodiment of the present invention, for any target location It can collect all that meet the requirements. Neighbor set Set of points The cluster set corresponding to any target location is determined. Wherein, if Then any target location Mark the location as the core location; otherwise, temporarily label it as a noise location. Calculate the average value of all target locations in the neighbor set corresponding to the core location to obtain the initial tracking location corresponding to the core location. Noise locations, if visited in subsequent expansions, are assigned to the corresponding cluster set. The number of target locations in the neighboring point set. This is the minimum number threshold. At least one initial tracking position is the initial tracking position corresponding to at least one core position.

[0151] According to embodiments of the present invention, a unique target ID can be assigned to each cluster set, and the average three-dimensional coordinates of all target locations within the cluster set can be calculated as the optimal target location. Each target ID and its optimal location are then passed to subsequent tracking or strike decision-making processes to achieve high-precision positioning and action planning.

[0152] According to an embodiment of the present invention, after determining the clustering radius, the control module 150 determines the acquisition time corresponding to each of the multiple target images, determines the detection box information corresponding to each of the multiple target positions based on the target detection results corresponding to each of the multiple target images, determines the positional deviation between each target position and the multiple target positions corresponding to the multiple target images based on any target position, the acquisition time corresponding to each of the multiple target images, the target position, and the detection box information, and clusters the target positions adjacent to any target position among the multiple target positions corresponding to the multiple target images based on the multiple positional deviations and the clustering radius, and obtains the cluster set corresponding to any target position. This operation is a depth-sensing spatiotemporal adaptive clustering algorithm, which performs three-dimensional spatiotemporal density clustering on the set of target observation points of multiple frames by combining the detection box, semantic category, and monocular depth estimation results. Unlike traditional clustering algorithms that rely solely on two-dimensional spatiotemporal distance, this method incorporates the target bounding box area (representing scale information) and the three-dimensional coordinates after depth back-projection, i.e., the target location, into the clustering metric. This enables the effective differentiation of occluded targets and suppression of false detection noise in complex viewpoint changes and multi-target object scenarios. Furthermore, it automatically adapts to the target distance distribution, achieving higher accuracy and stronger robustness in target localization and clustering.

[0153] According to an embodiment of the present invention, during the target tracking and strike phase, the monocular vision-based UAV control system 100 captures and processes the target position corresponding to each frame of the target image in real time, ensuring that the UAV can respond quickly and adjust its flight attitude precisely, thereby achieving efficient target tracking and accurate strike.

[0154] According to an embodiment of the present invention, the visual perception module 140 captures target image frames in real time and generates a target detection result and a third distance for each detected target object. This information is sent to the control module 150 as the basis data for tracking and localization.

[0155] According to an embodiment of the present invention, the control module 150 determines the target distance between the UAV and the target object based on a first distance, a second distance, and a third distance, and determines the target position of the target object in the navigation coordinate system corresponding to the servo deployment device based on the target distance, the target detection result, and the coordinate mapping relationship between the monocular acquisition device and the servo deployment device. Based on the target position, the relative distance between the UAV and the target object in the navigation coordinate system is determined.

[0156] If the current relative distance is less than the preset distance threshold or the strike operation timeout occurs, a strike operation is triggered. In this case, the control module 150 immediately executes the strike operation and outputs a control signal corresponding to the strike operation. This allows the flight control module 130 to control the servo to release according to the control signal, and the monocular vision-based UAV control system 100 exits the tracking and strike mode. If the current relative distance does not reach the preset distance threshold and the timeout has not occurred, the system enters the PID (Proportional-Integral-Derivative) control stage.

[0157] According to an embodiment of the present invention, the control module 150 outputs a control signal or tracks a position based on the target position, including: when it is determined that a target object will be attacked, using a feedback adjustment controller to calculate the flight speed and direction of the UAV based on the target position and the current reference position of the UAV in the navigation coordinate system, and outputs a control signal indicating the flight speed and direction, wherein the current reference position is obtained by the flight control module based on a first distance, inertial data and positioning information output by the satellite positioning unit included in the flight control module, and the positioning information includes the satellite positioning distance.

[0158] For example, the feedback control controller can be a PID controller.

[0159] According to an embodiment of the present invention, when it is determined that a target object will be attacked, a PID controller is used to calculate the flight speed and direction of the UAV based on the target position and the current reference position of the UAV in the navigation coordinate system, and output control signals indicating the flight speed and direction. This enables the UAV to control its flight speed and direction using the PID controller, gradually approaching the location of the target object, and ensuring that the UAV can smoothly approach the target object.

[0160] According to an embodiment of the present invention, during the tracking and strike process, the control module 150 continuously receives target detection results, a third distance, a first distance from the laser rangefinder 120, a second distance from the flight control module 130, and the current reference position from the visual perception module 140, and continuously calculates the target position. It then dynamically adjusts the flight path based on the latest target position and the current reference position. Furthermore, considering the relatively low frame rate of the monocular acquisition device 110, and consequently the relatively low frequency of data output by the visual perception module 140, the control module 150 employs predictive algorithms, such as PID control methods, to compensate for possible delays, ensuring that the UAV can respond promptly to the movement of the target object.

[0161] For example, within the time frame during which the visual perception module 140 outputs relevant data four times, the control module 150 can adjust the flight speed and direction of the drone 20 times. Within the time frame during which the visual perception module 140 outputs relevant data between two consecutive times, the number of adjustments made by the control module 150 varies according to the actual situation.

[0162] Figure 3 A flowchart of a monocular vision-based unmanned aerial vehicle (UAV) control method according to an embodiment of the present invention is shown.

[0163] Figure 3 The monocular vision-based UAV control method shown can be applied to... Figure 1 The monocular vision-based UAV control system 100 shown is used to control the UAV to track and strike target objects.

[0164] like Figure 3 As shown, the UAV control method based on monocular vision can include operations S310 to S340.

[0165] During operation of S310, the control module receives the target detection result and third distance from the visual perception module, the first distance from the laser rangefinder, and the second distance and current reference position from the flight control module. Based on the first distance, second distance, and third distance, it determines the target distance between the UAV and the target object, and calculates the target position based on the target distance, target detection result, coordinate mapping relationship between the monocular acquisition device and the servo delivery device, and determines the target position of the target object in the navigation coordinate system corresponding to the servo delivery device.

[0166] When operating the S320, the control module determines the relative distance between the UAV and the target object in the navigation coordinate system based on the target's location, and determines whether the relative distance is less than a preset distance threshold or whether the strike operation timeout has occurred.

[0167] If not, in operation S330, the control module uses the feedback adjustment controller to calculate the UAV's flight speed and direction based on the target position and the UAV's current reference position in the navigation coordinate system, and outputs control signals indicating the flight speed and direction. Then, it returns to operation S310. If yes, in operation S340, the control module outputs a control signal corresponding to the strike operation. The flight control module can then use this control signal to release the servos and strike the target object. The monocular vision-based UAV control system exits the tracking and strike mode. Specifically, if the relative distance is less than a preset distance threshold, it is determined that the UAV has reached above the target object.

[0168] According to embodiments of the present invention, the above design enables efficient tracking and precise strike of target objects. The rich information provided by the visual perception module, laser rangefinder, and flight control module not only supports accurate target positioning but also enhances the robustness and adaptability of the system. Simultaneously, the high-frequency operation of the control module ensures the rapid response capability of the UAV, thereby improving the overall system performance and reliability.

[0169] Figure 4 A flowchart of a monocular vision-based unmanned aerial vehicle (UAV) control method according to another embodiment of the present invention is shown.

[0170] like Figure 4 As shown, the monocular vision-based UAV control method includes operations S401 to S406. Operation S402 includes operations S4021 to S4023. Operation S403 includes operation S4031. Operation S404 includes operations S4041 to S4044.

[0171] In S401, connect the drone. In S402, create a data parsing task.

[0172] In the data parsing task, in operation S4021, a data parsing function is registered. In operation S4022, a real-time flight control data acquisition task is created, which writes data to a message queue. In operation S4023, the task waits for completion. For example, the data parsing function can be a Topic parsing function corresponding to the Topic communication mechanism. The flight control data acquired in operation S4022 includes data output from the monocular acquisition device 110, laser rangefinder 120, flight control module 130, visual perception module 140, and control module 150.

[0173] In operation S403, a serialization task is created. Within the serialization task, in operation S4031, data is periodically serialized to a file. In operation S404, the main task begins polling until it completes. The main task includes the tracking task executed in operation S4041 and the strike task executed in operation S4042.

[0174] During the execution of operation S4041 or S4042, operation S4043, target perception and recognition, and operation S4044 will be triggered simultaneously.

[0175] In step S405, cancel all tasks and wait for completion. In step S406, save the log file and exit the program.

[0176] According to an embodiment of the present invention, the main program in the monocular vision-based UAV control method is executed by an onboard computer. Operations S4021-S4023 and S4031 are operations implemented by the message queue and serialization module. Operations S4041-S4042 are operations executed by the control module 150. Operation S4043 is an operation executed by the visual perception module 140. Operation S4044 is an operation executed by the flight control module 130.

[0177] According to an embodiment of the present invention, the visual perception module 140 can directly receive the target image output by the monocular acquisition device 110. In this case, the visual perception module 140 needs to obtain inertial data from the message queue whose acquisition time is the same as that of the target image. The visual perception module 140 can also directly obtain the target image and inertial data with the same timestamp from the message queue, and perform target perception and recognition based on this data.

[0178] According to embodiments of the present invention, the monocular vision-based UAV control system provided by the present invention operates based on the coordination of a primary task. The primary task is responsible for initiating connections, creating necessary data processing and flight tasks, and saving results upon completion of the tasks. This design not only significantly improves the system's response speed and accuracy but also enhances its scalability and flexibility, allowing for easy addition of new functional modules or improvement of existing modules in the future without major modifications to the overall architecture.

[0179] According to embodiments of the present invention, UAV tracking and strike systems in related technologies often rely on multi-sensor (including sensors in RTK modules) fusion, resulting in complex hardware and high costs. Traditional multi-sensor systems require precise calibration and data synchronization; even slight deviations affect performance, and they also increase weight and energy consumption, limiting endurance and applicability. Hardware costs and performance requirements for airborne hardware are both high. The monocular vision-based UAV control system 100 provided in this invention uses only the GNSS positioning module 130, monocular acquisition device 110, and laser rangefinder 120, significantly reducing the number of hardware types, lowering costs, simplifying integration, reducing UAV weight and power consumption, and improving endurance and applicability.

[0180] As related technologies expand the functionality of unmanned aerial vehicles (UAVs), software systems become increasingly complex. Other software frameworks, such as MAVROS, have complex architectures, are cumbersome to develop, have long development cycles, and consume significant hardware resources. Under limited hardware resources, the system is prone to lag and delays, affecting flight safety and mission execution. This invention uses the lightweight MAVSDK development library, simplifying the development process, reducing difficulty and cycle time. It has low hardware performance requirements, can fully utilize hardware performance, improve the real-time performance and response speed of UAVs, and ensure efficient mission execution.

[0181] According to embodiments of the present invention, the design of the message pool and serialization module in the monocular vision-based UAV control system 100 enables efficient data management and processing, ensuring seamless data exchange and storage among the system's components. The message queue integrates data from the flight control module 130, the visual perception module 140, and the control module 150, ensuring all components access the latest information. A timed serialization task periodically serializes and stores data in a bag file, ensuring data consistency and integrity. A unified interface design simplifies communication between modules, improving system maintainability and scalability. This design reduces hardware costs and computational complexity, enhances system real-time performance and response speed, and provides strong support for the development of UAV technology. This modular design not only reduces hardware costs and computational complexity but also improves data processing efficiency and system scalability, solving the problems of high cost and high computational complexity inherent in traditional multi-sensor fusion systems.

[0182] Traditional monocular vision systems in related technologies suffer from insufficient positioning accuracy and are prone to target loss or misidentification in UAV tracking and strike missions. Traditional target recognition methods either rely on single visual information, resulting in low accuracy, or involve large computational demands due to multi-sensor fusion. Furthermore, in complex scenarios such as multiple targets, occlusion, and varying lighting conditions, recognition accuracy and stability are difficult to guarantee. The visual perception module 140 of this invention integrates YOLOv8, UNet, and IMU sensor data to achieve an efficient combination of target detection, semantic recognition, and depth estimation, improving recognition accuracy and stability. The depth-sensing spatiotemporal adaptive clustering algorithm can also effectively distinguish occluded targets, suppress noise, and adapt to target distance distribution, providing reliable information for precision strikes.

[0183] In unmanned aerial vehicle (UAV) systems, the transformation between different coordinate systems is a crucial step in achieving autonomous tracking and strike capabilities. Traditional coordinate transformation relies on complex geometric models and precise synchronization of multi-sensor data, involving cumbersome parameter calibration and calculation steps, resulting in significant coordinate deviations after transformation, which impacts flight safety and mission execution. This invention designs a simple and complete coordinate transformation process, employing targeted methods at each stage, such as distortion correction, laser ranging, and SVD algorithms, to ensure transformation accuracy and provide reliable flight control commands for high-precision tracking and strike missions.

[0184] Drone tracking and strike missions require a scientifically sound process. Traditional processes are prone to problems such as inaccurate target identification and imprecise positioning, and also have shortcomings in path planning and target tracking and strike phases. Traditional processes lack comprehensive analysis of multi-frame data in the tracking phase, have simplistic path planning, and suffer from long execution times and low efficiency. Furthermore, they lack the ability to dynamically adjust to the target during tracking and strike, affecting strike accuracy. This invention tightly integrates each stage of the tracking and strike process. Multi-frame data clustering improves accuracy in the tracking phase, path planning simplifies complexity and improves efficiency, and PID control and predictive algorithms in the tracking and strike phases ensure rapid response and precise adjustments, thereby improving mission execution efficiency and success rate.

[0185] Those skilled in the art will understand that the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.

[0186] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.

Claims

1. A monocular vision-based unmanned aerial vehicle (UAV) control system, characterized in that, include: A monocular acquisition device used to acquire target images including the target object; Laser rangefinders are used to detect the initial distance between a drone and the ground. The flight control module is used to detect the second distance between the UAV and the ground, the inertial data of the UAV, and to control the deployment of the servo motor according to the control signal or to control the UAV to track the target object according to the tracking position. The visual perception module is used to perform target detection on the target image using a deep neural network and obtain the target detection result; to extract multi-scale fusion features of the image features output by the feature extraction layer of the deep neural network using a decoder; and to extract inertial features of inertial data for a predetermined duration using a first perceptron. The multi-scale fusion feature and the inertial feature are concatenated to obtain a third concatenated feature; the third concatenated feature is continuously convolved and nonlinearly operated to obtain a relative depth feature; the inertial data of the predetermined duration is pre-integrated to obtain the depth change. The depth change is regressed using a second perceptron to obtain a scale factor; a third distance is obtained based on the scale factor and the relative depth feature. The control module is used to determine the target distance between the UAV and the target object based on the first distance, the second distance and the third distance, and to determine the target position of the target object in the navigation coordinate system corresponding to the servo deployment device based on the target distance, the target detection result, the coordinate mapping relationship between the monocular acquisition device and the servo deployment device, and to output control signals or tracking positions based on the target position.

2. The unmanned aerial vehicle control system according to claim 1, characterized in that, The image features include a first image feature, a second image feature, and a third image feature, wherein the first image feature is obtained based on the target image, the second image feature is obtained based on the first image feature, and the third image feature is obtained based on the second image feature; The visual perception module is also used for: The third image feature is bilinearly upsampled multiple times to obtain the first upsampled feature; The first upsampled feature and the second image feature are concatenated by channels to obtain the first concatenated feature. The first concatenated feature is then subjected to convolution, normalization and nonlinear operations multiple times to obtain the first fused feature. The first fused feature is subjected to bilinear upsampling multiple times to obtain the second upsampled feature; The second upsampled feature and the first image feature are concatenated to obtain the second concatenated feature. The second concatenated feature is then subjected to convolution, normalization and nonlinear operations multiple times to obtain the second fused feature. The second fused feature is subjected to bilinear upsampling multiple times to obtain the third upsampled feature; The third upsampled feature is subjected to convolution, normalization, and nonlinear operations multiple times to obtain the multi-scale fused feature.

3. The unmanned aerial vehicle control system according to claim 1, characterized in that, The second distance includes a fused distance and a satellite positioning distance, wherein the satellite positioning distance is the distance between the UAV and the ground output by the satellite positioning unit included in the flight control module, and the fused distance is the distance between the UAV and the ground obtained by the flight control module based on the inertial data and the satellite positioning distance; The control module is also used for: If the first condition is met, the first distance is determined as the target distance; If the first condition is not met but the second condition is met, the third distance is determined as the target distance; If neither the first condition nor the second condition is met, but the third condition is met, the satellite positioning distance is determined as the target distance. If the first condition, the second condition, and the third condition are all false, but the fourth condition is true, the fusion distance is determined as the target distance; The first condition is that the laser rangefinder is operating normally and the fusion distance is less than the first maximum range of the laser rangefinder; the second condition is that the fusion distance is less than the second maximum range of the monocular acquisition device and the absolute value of the difference between the third distance and the fusion distance is less than the first preset threshold; and the third condition is that the absolute value of the difference between the satellite positioning distance and the fusion distance is less than the second preset threshold.

4. The unmanned aerial vehicle control system according to claim 1, characterized in that, The control module is also used for: When it is determined that the target object is to be tracked, at least one target location corresponding to each of the multiple frames of target images is clustered to obtain at least one cluster set; Based on the at least one cluster set, at least one initial tracking location is obtained; If the number of initial tracking positions is one, the initial tracking position is determined as the tracking position, and the tracking position is output. When there are multiple initial tracking positions, multiple tracking positions are obtained according to the sorting of the multiple initial tracking positions, and the multiple tracking positions are output.

5. The unmanned aerial vehicle control system according to claim 4, characterized in that, The control module is also used for: For each frame of the target image, the mean and standard deviation are calculated for multiple target locations corresponding to the target image; The cluster radius is determined based on the mean, the standard deviation, the first preset coefficient, and the second preset coefficient. For any target position among multiple target positions corresponding to the target image, calculate the positional deviation between the any target position and the multiple target positions corresponding to the multiple frames of target images; Based on the multiple positional deviations and the clustering radius, the target positions that are adjacent to any target position among the multiple target positions corresponding to the multiple frames of target images are clustered to obtain a cluster set corresponding to any target position.

6. The unmanned aerial vehicle control system according to claim 5, characterized in that, The control module is also used for: Determine the acquisition time corresponding to each of the multiple target images; Based on the target detection results corresponding to each of the multiple target images, determine the detection box information corresponding to each of the multiple target locations; Based on any target location, the acquisition time corresponding to each of the multiple target images, the target location, and the detection box information, determine the positional deviation between any target location and the multiple target locations corresponding to the multiple target images.

7. The unmanned aerial vehicle control system according to claim 1, characterized in that, The control module is also used for: When it is determined that the target object will be attacked, the feedback adjustment controller calculates the flight speed and direction of the UAV based on the target position and the current reference position of the UAV in the navigation coordinate system, and outputs control signals indicating the flight speed and direction. The current reference position is obtained by the flight control module based on the first distance, the inertial data and the positioning information output by the satellite positioning unit included in the flight control module. The positioning information includes the satellite positioning distance.

8. The unmanned aerial vehicle control system according to claim 1, characterized in that, The coordinate mapping relationship between the monocular acquisition device and the servo motor deployment device is determined according to the following operation: The monocular acquisition device is used to acquire a checkerboard calibration board placed on the ground to obtain a checkerboard image. At the same time, the projection position of the servo motor deployment device on the checkerboard calibration board is determined, and the projection position is determined as the origin position of the servo motor coordinate system corresponding to the servo motor deployment device. Based on the corner points of each chessboard grid included in the chessboard calibration board, the side length of each chessboard grid, the position of the origin, and the chessboard grid image, determine the coordinates of the corner points of each chessboard grid in the servo coordinate system and the camera coordinate system where the monocular acquisition device is located. Based on the coordinates of the corner points of each chessboard grid in the servo coordinate system and the camera coordinate system where the monocular acquisition device is located, a covariance matrix is ​​constructed and decomposed to obtain a rotation matrix and a translation matrix used to transform the coordinates in the camera coordinate system where the monocular acquisition device is located to the coordinates in the servo coordinate system. The flight control module obtains the quaternion and position translation information corresponding to the navigation coordinate system; The coordinate mapping relationship between the monocular acquisition device and the servo motor deployment device is determined based on the rotation matrix, the translation matrix, the quaternion, and the position translation information.

9. The unmanned aerial vehicle control system according to claim 1, characterized in that, The control module is also used for: Obtain the first distance, the second distance, the third distance, the target detection result, and the coordinate mapping relationship between the monocular acquisition device and the servo motor deployment device from the message queue; The visual perception module is also used to obtain the target image and the inertial data from the message queue; The first distance, the second distance, the third distance, the target image, the inertial data, and the target detection result in the message queue all have the same timestamp.

Citation Information

Patent Citations

  • System and method for monocular airborne target detection

    CN108364304A

  • Unmanned aerial vehicle small target automatic tracking method and device based on monocular vision

    CN117830353A

  • BEV elevation estimation method and system based on binocular data

    CN120388230A