Intelligent unmanned aerial vehicle real-time space perception and target precise detection system

By combining three-dimensional visual RGB space extraction and deep learning object detection methods, the real-time spatial perception and object detection problems of drones in unknown environments are solved, real-time obstacle avoidance and path planning of drones are realized, and the intelligent flight capabilities of drones are improved.

CN120495616APending Publication Date: 2025-08-15赵海山
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510396604.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing drones are difficult to achieve real-time spatial perception and target detection under unknown environments, cannot effectively complete obstacle avoidance and path planning, and lack of real-time three-dimensional map provision and target position identification, resulting in the unmanned aerial vehicle being unable to achieve intervention-free or semi-intervention flight.

Method used

The three-dimensional visual RGB space extraction method and deep learning-based object detection method are adopted, combined with fast deep convolutional neural network and ORB-visual real-time spatial perception modeling, a fast deep convolutional neural network feature extractor is built, a sample library and feature library of interest in indoor targets is established, and a microcomputer embedded in a drone is used to achieve real-time object detection and path planning.

Benefits of technology

Real-time spatial awareness and target detection of drones in unknown environments are realized, space perception efficiency and target detection speed are improved, obstacle avoidance and path planning are ensured, and no intervention or semi-intervention flight of drones is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495616A_ABST
    Figure CN120495616A_ABST
Patent Text Reader

Abstract

According to the intelligent unmanned aerial vehicle real-time space perception and target precise detection system, the space perception module of the unmanned aerial vehicle is improved by integrating a three-dimensional visual RGB space extraction method, and the target detection method is optimized, so that a real-time target precise positioning and target precise recognition capability improvement content extraction module is added to the unmanned aerial vehicle; and finally, navigation and path planning basis is provided for the unmanned aerial vehicle by fusing the two methods, and an intelligent unmanned aerial vehicle system including software and hardware is established. The unmanned aerial vehicle and the JETSON TX2 are used as carrying and development platforms, the combination of visual real-time space sensing modeling and target detection methods of the unmanned aerial vehicle in a specific scene is realized, the method is used for the unmanned aerial vehicle with the flight speed limited by the scene, the space sensing efficiency is high, the target detection speed is high, the accuracy is high, and the method is suitable for popularization and application. Obstacle avoidance and path planning are successfully completed under the condition that the environment is unknown and the unmanned aerial vehicle needs to fly while exploring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a UAV space perception target detection system, and in particular to an intelligent UAV real-time space perception and target precision detection system, belonging to the field of UAV target detection technology. Background Art

[0002] As the value of drone applications continues to grow, their performance is also improving. Drones are no longer equipped with simple cameras for video recording; they are now equipped with specialized equipment such as depth cameras and lidar to perform specific tasks. Current drone applications include vegetation protection, street photography, power inspections, and disaster relief, but these applications are largely basic. Obstacle avoidance and path planning have long been challenging challenges in drone applications. These challenges require the drone's spatial perception of the surrounding environment and content extraction based on object recognition. Currently, most drones address these challenges by manually planning and controlling flight in a known environment. However, successfully achieving obstacle avoidance and path planning in unknown environments, requiring the drone to explore while flying, remains a work in progress. The emergence of content extraction methods based on convolutional neural network object detection offers a promising opportunity to address this problem. Providing a real-time 3D map for a drone can help it accurately locate itself, while content extraction methods based on object detection can identify objects in the scene and determine their locations. Combined, these two methods provide meaningful reference data for obstacle avoidance and path planning, enabling hands-free or semi-handled flight. More importantly, if both methods achieve real-time processing, they could be used to establish an intelligent drone-based processing system using spatial scene perception methods embedded in a microcomputer and target detection-based content extraction methods. Furthermore, the implementation of real-time spatial perception and target recognition-based content extraction methods onboard drones is not limited to solving obstacle avoidance and path planning problems; it can also efficiently utilize the collected data and expand into other applications. These insights into drone applications are of great value.

[0003] In terms of spatial scene perception, with the improvement of hardware performance in visual cameras, lidar cameras, and other technologies, scene perception not only obtains more reliable data sources but also enables information to be complemented through the collaborative work of multiple devices. LiDAR and real-time spatial perception modeling each have their own characteristics, and each has certain limitations when used alone. However, when used together, their strengths and weaknesses can complement each other. For example, vision systems operate relatively stably in texture-rich, dynamic environments and can provide highly accurate point cloud matching for lidar, while the relatively precise direction and distance information provided by lidar can assist in correcting the point cloud image. Furthermore, in environments with low lighting or significant texture loss, lidar can leverage its advantages to assist real-time spatial perception modeling in recording scenes with minimal information. Furthermore, lidar and real-time spatial perception modeling systems are not limited to a single solution. They generally utilize auxiliary positioning tools such as inertial sensors, satellite positioning systems, and indoor base station positioning systems, creating a complementary environment. This has been a recent research trend towards data fusion and collaboration between lidar systems and other sensors. The current trend is towards tightly coupled fusion based on nonlinear global optimization, compared to the previous loosely coupled fusion method based on Kalman filtering. For example, the fusion of real-time spatial perception modeling and IMU (inertial navigation system) can achieve real-time mutual calibration, allowing the vision module to maintain a certain positioning accuracy during sudden acceleration, deceleration, or rotation, preventing tracking loss and significantly reducing positioning and map construction errors.

[0004] In content scene perception, or object detection, the current trend is to balance accuracy and speed. The key lies in starting with candidate box-based object detection. Specifically, this requires maximizing computational sharing between different ROIs, eliminating redundant computations, and efficiently utilizing CNN-derived features, thereby improving overall detection speed. However, candidate box object detection still suffers from artifacts and missed detections, two key issues that need to be addressed in specific object detection applications.

[0005] The problems that need to be solved in existing UAV spatial perception target detection and the key technical difficulties of this application include:

[0006] (1) Obstacle avoidance and path planning are the current challenges in drone applications. These two challenges involve the drone's spatial perception of the surrounding environment and content extraction based on target recognition. Currently, most drones solve the obstacle avoidance and path planning problems by manually making plans and controlling the drone's flight when the environment is known. The successful completion of obstacle avoidance and path planning when the environment is unknown and the drone needs to explore while flying is still under exploration. The emergence of content extraction methods based on convolutional neural network target detection has provided a certain opportunity to solve this problem, but the method is not mature enough and cannot provide real-time three-dimensional maps and cannot help drones accurately locate. The content extraction of target detection cannot fully identify the target in the scene and give the target's position. The combination of the two can provide some meaningful reference data for drone obstacle avoidance and path planning, but there are many problems in the combination process. The technology is not mature enough and cannot achieve non-intervention or semi-intervention flight of drones. Moreover, these two methods cannot achieve real-time processing. It is impossible to use the spatial scene perception method embedded in the microcomputer on the drone and the content extraction method based on target detection to establish a drone onboard intelligent processing system. The inability to efficiently solve the obstacle avoidance and path planning problems restricts the application of drones.

[0007] (2) The existing technology lacks the integration of 3D visual RGB spatial extraction methods to improve the spatial perception module of drones, lacks target detection methods to increase the real-time target precision positioning and target precision identification capabilities of drones to improve the content extraction module, and cannot provide the drone with navigation and path planning basis by integrating the two methods. There is no establishment of an intelligent drone system including software and hardware, resulting in the drone being unable to complete real-time spatial perception and target precision detection. The existing technology lacks the construction of a fast deep convolutional neural network feature extractor, which cannot learn and extract the features of all targets of interest for specific target detection and recognition; lacks the use of single-purpose ORB-visual real-time spatial perception modeling to achieve rapid composition of specific scenes; lacks an integrated target detection module to mark the target of interest in the visual real-time spatial perception modeling diagram and provide specific location information, resulting in poor real-time spatial perception capabilities, low target detection accuracy, poor obstacle avoidance capabilities, and poor application security.

[0008] (3) The existing technology lacks a sample library and a feature library of indoor targets of interest, and cannot use an end-to-end fast neural network to realize real-time target detection and matching. It lacks a real-time composition system for the detected targets using a three-dimensional visual RGB space extraction system, cannot mark the identified targets of interest in the simulation map, cannot provide specific location information, lacks a real-time spatial perception modeling method based on vision, cannot provide a three-dimensional point cloud map for the UAV, cannot provide the spatial positioning of the UAV, and cannot provide data support for the navigation of the UAV; lacks target detection content extraction based on deep learning, cannot provide the spatial position of various targets for the UAV, and cannot use spatial relationships to do obstacle avoidance and path planning for the UAV; lacks a real-time perception and target detection system for the UAV, lacks the module to be embedded in a microcomputer, lacks an intelligent system for real-time perception and target detection for the UAV including software and hardware, cannot solve the obstacle avoidance and path planning problems of the UAV well, and cannot realize the non-intervention or semi-intervention flight of the UAV. Summary of the Invention

[0009] This application constructs a fast deep convolutional neural network feature extractor to learn and extract the features of all targets of interest for use in specific target detection and recognition; uses single-purpose ORB-visual real-time spatial perception modeling to achieve rapid composition of specific scenes; integrates the target detection module to mark the targets of interest in the visual real-time spatial perception modeling diagram and gives specific location information; the target detection method of this application has real-time performance while ensuring a certain accuracy rate, applies the machine vision target detection method to content-based scene perception, integrates the space-based scene perception method, and establishes an intelligent drone system. Using drones and JETSON TX2 as the carrier and development platform, it realizes the combination of visual real-time spatial perception modeling and target detection methods of drones in specific scenarios. It is used for drones whose flight speed is restricted by the scene, with high spatial perception efficiency, fast target detection speed and high accuracy, and successfully completes obstacle avoidance and path planning when the environment is unknown and the drone needs to explore while flying.

[0010] To achieve the above technical effects, the technical solutions adopted in this application are as follows:

[0011] The intelligent UAV real-time spatial perception and target precision detection system improves the UAV's spatial perception module by integrating a 3D visual RGB spatial extraction method. The optimized target detection method adds real-time target precision positioning and target precision recognition capabilities to the UAV, improving the content extraction module. Ultimately, by integrating the two methods, the UAV is provided with a basis for navigation and path planning, and an intelligent UAV system, including software and hardware, is established. First, a fast deep convolutional neural network feature extractor is constructed to learn and extract the features of all targets of interest for specific target detection and recognition. Then, single-purpose ORB-based real-time spatial perception modeling is used to achieve rapid composition of specific scenes. Finally, the integrated target detection module marks the targets of interest in the visual real-time spatial perception modeling diagram and provides specific location information.

[0012] This application establishes a sample library and feature library of indoor targets of interest, uses an end-to-end fast neural network to achieve real-time target detection and matching, and uses a three-dimensional visual RGB space extraction system to perform real-time composition while detecting targets. The identified targets of interest are marked in the simulation map and their specific location information is given. The core methods include:

[0013] (1) Based on the visual real-time spatial perception modeling method: the depth camera is mounted on the UAV to provide RGB acquisition images and depth acquisition images for real-time spatial perception modeling. The real-time spatial perception modeling, i.e., 3D visual RGB space extraction, is used to complete the perception of the spatial scene. Finally, a 3D point cloud map is provided for the UAV, which gives the spatial positioning of the UAV and provides data support for the UAV navigation.

[0014] (2) Object detection content extraction based on deep learning: Based on the speed of UAV flight, the object detection network optimizes real-time processing capabilities, combines the conversion relationship obtained by scene perception, transforms the target detected into three-dimensional space, provides the spatial position of various targets for the UAV, and uses the spatial relationship to perform obstacle avoidance and path planning.

[0015] (3) Establish a real-time perception and target detection system for UAVs: Based on the optimization of real-time scene perception and real-time target detection, the implementation of the two modules is embedded in a microcomputer to establish an intelligent system including software and hardware for real-time perception and target detection suitable for UAVs.

[0016] Preferably, the target real-time precise detection network data format: the training data sample includes a large collection image of multiple objects. For each object in the collection image, the training label includes not only the class of the object but also the coordinates of each corner point of the bounding box. The number of objects varies between different training collection images. By introducing a fixed three-dimensional label format, the problem that the selection of label formats of different lengths and dimensions makes it difficult to define the loss function is solved. The defined format can input collection images of any size including any number of objects.

[0017] The collected image is divided into regular grids with a grid size slightly smaller than the smallest object expected to be detected. Each grid has two key pieces of information, including the category of the object and the coordinates of the corner points of the grid that contains the object. In addition, when there is no object in the grid, a special custom class, the "dontcare" class, is used to uniformly maintain a fixed size in data representation, and an object coverage value represented by 0 or 1 is also set to indicate whether there is an object in the grid. For the case where many objects are in the same grid, the object that occupies the most pixels in the grid is selected. In the case of overlapping objects, the object with the bounding box with the smallest Y value is used.

[0018] Preferably, the target real-time precise detection network training is divided into three steps:

[0019] Step 1: The data layer obtains training collection images and labels, and the conversion layer performs online data enhancement;

[0020] Step 2: The fully convolutional network extracts and predicts the object class and bounding box of each grid;

[0021] Step 3: Predict the object category and target bounding box for each grid separately, and then use the loss function to calculate the error of the two prediction tasks simultaneously;

[0022] The prediction process consists of two steps: first, during the validation process, a clustering function is used to generate the final bounding box set; second, the model performance is measured by a simplified mAP calculation value on the validation dataset;

[0023] The network accepts input images of varying sizes and effectively applies a CNN in a sliding window fashion with a stride, outputting a multidimensional array that is superimposed on the image. Using GoogLeNet with the final pooling layer removed, the CNN uses a sliding window of up to 555×555 pixels with a stride of 16 pixels.

[0024] The final optimized loss function is generated using a linear combination of two independent loss functions, which include the sum of squares of the difference between the true and predicted object coverage of all grids in the training data samples, and the mean absolute difference loss between the true and predicted corner points of the bounding box of the object covered by each grid.

[0025] Preferably, the process of real-time spatial perception modeling includes:

[0026] Process 1: Reading sensor image data: This involves reading and preprocessing the image information collected by the drone camera in real-time spatial perception modeling. The depth camera data includes the RGB image and its corresponding depth map.

[0027] Process 2: Visual odometry modeling: Visual odometry estimates the rotation and translation relationship between two adjacent captured images to calculate the camera's pose change and local map. The key to this step is feature point extraction and captured image matching.

[0028] Process 3: Back-end global optimization: The back-end uses a nonlinear global optimization algorithm to optimize the camera position and posture from the front-end and the past-and-back detection results obtained by another thread, correcting them to a globally unified trajectory map and point cloud map.

[0029] Process 4: Past-location detection: This process determines whether the sensor or the drone carrying the sensor has passed through a certain scene before. If it is detected that the sensor has been there before, the information is provided to the backend to correct the position and posture.

[0030] Process 4, composition: Based on the estimated camera trajectory, a drone cruise map that meets the mission requirements is created.

[0031] Preferably, the real-time spatial perception modeling framework:

[0032] 1) Visual odometry modeling

[0033] The front-end receives the camera's video stream, or captured image sequences, and estimates the camera's motion between adjacent frames using feature matching methods. This allows for the initial acquisition of mileage information with a certain degree of error accumulation. Visual odometry modeling consists of four parts:

[0034] First, the acquisition frame: The information carried includes the drone camera pose, RGB acquisition image, and depth map when the frame was captured;

[0035] Second, the camera model: corresponds to the camera used in the actual shooting, and only includes internal parameters;

[0036] Third, local map: including keyframes and landmark information points. Keyframes and landmark information points that meet the matching rules will be added to the map. The map is only a local map, not a global map. It only includes landmark information points near the current location, and landmark information points farther away are deleted.

[0037] Fourth, landmark information points: These are map points with known information. The known information included in the landmark information points is the feature description corresponding to them. The acquisition method is to use feature matching algorithms to extract them in batches.

[0038] 2) Backend global optimization

[0039] The global process analyzes the data and handles noise problems, including linear global optimization and nonlinear global optimization algorithms. Linear global optimization assumes that there is a linear relationship between each frame of the acquisition image during the shooting process, and uses the Kalman filter algorithm to estimate the state. If it is assumed that only the previous frame and the next frame have a linear relationship, the state estimation is completed by the extended Kalman filter. The difference between the observed value and the algorithm estimate is calculated, that is, the error value between the pixel coordinates and the pixel coordinates of the corresponding 3D point projected onto the two-dimensional plane through the camera position. Linear global optimization assumes that the error between the camera position and the spatial point has a causal relationship. The camera position and posture are first calculated, and then the position of the spatial point is further calculated based on the camera position. Nonlinear global optimization directly puts all data into the same model for optimization and solution, downplaying the relationship between the data.

[0040] 3) Testing of travel to and from the old place

[0041] The key to correcting the accumulated errors of the visual odometry for past-time detection lies in building a bag-of-words model. This model abstracts features into individual words. The detection process involves matching the words that appear in two images to determine whether they depict the same scene. To classify features into words, it is necessary to train a dictionary that includes all possible word sets. Building this dictionary requires massive amounts of data for training. Dictionary building is a clustering process. Assuming that a total of 100 million features are extracted from all images, they are clustered into 100,000 words using the K-means clustering method. During the dictionary training process, a tree with k branches and a depth of d is constructed. The upper nodes of the tree provide coarse classification, and the lower nodes provide fine classification, extending all the way to the leaf nodes. Using this tree, the time complexity is reduced to logarithmic level, accelerating feature matching.

[0042] 4) Composition

[0043] The data collected by the camera is optimized and the camera posture is corrected to convert the two-dimensional plane points into three-dimensional space, forming three-dimensional space point cloud information. In addition to the point cloud map, the camera posture optimization process is represented in the g2o tool to form a posture graph. Depending on the specific situation, you can also define your own map for description.

[0044] Preferably, the three-dimensional visual RGB space extraction method:

[0045] The sensor uses a monocular camera to obtain depth information, and the data source includes RGB acquisition image and depth map.

[0046] Use the depth camera to obtain the color acquisition image and the depth acquisition image, and use the geometric model to convert the 2D plane data into a 3D stereo space. From the pixel coordinates to the acquisition image coordinates: the center points of the coordinate systems are different, and there is only an offset relationship between the two. From the acquisition image coordinate system to the camera coordinate system, the coordinate axes are parallel, and there is only a scaling relationship. The conversion relationship from the pixel coordinate system to the camera coordinate system is expressed as:

[0047]

[0048] Where u and v are the offsets between the origins of the coordinate system and the center of the acquisition plane, d x , d y is the scaling ratio between pixel coordinates and the actual imaging plane, and the scaling ratio is d x =z c / f x , d y =z c / f y , f x , f y is the focal length of the camera on the x and y axes, written in matrix form:

[0049]

[0050] When the camera moves, the camera coordinate system and the world coordinate system are not parallel, and there is a rotation and translation relationship. In the subsequent visual odometry calculation, the relationship between the previous and next frames is the same as above, and the matrix relationship is given as:

[0051]

[0052] Convert the points on the two-dimensional plane to three-dimensional space, and finally obtain a series of point cloud data, which are assigned RGB color attributes to initially obtain a color three-dimensional map.

[0053] Preferably, the 3D visual RGB space extraction is implemented as follows:

[0054] (1) Front-end visual odometry

[0055] Initialization is based on the first frame of the acquisition image, and the search for key frames begins. The ORB algorithm is used to extract key points between the acquisition images, and then the BRIEF descriptor is calculated for each key point. Finally, the fast approximate nearest neighbor algorithm is used for fast matching. The ORB corner point extraction algorithm adds scale and rotation descriptions to the FAST corner point extraction algorithm, adding feature information, richer feature descriptions, higher matching accuracy, and more accurate and reliable composition. The BRIEF descriptor uses a binary descriptor and uses randomly selected points for comparison.

[0056] After matching is completed, the 2D points are projected into 3D space based on the depth acquisition image to obtain the 2D coordinates and corresponding 3D coordinates of a series of points. The camera's position is estimated by solving the PnP problem. The actual calculation result is the rotation and translation matrix between the two frames of acquisition. All data are matched pairwise in sequence and the camera pose is calculated to obtain a complete visual odometry.

[0057] (2) Back-end nonlinear global optimization

[0058] Three-dimensional visual RGB space extraction uses a pose graph to represent the pose calculated by the visual odometry, including nodes and edges. Nodes represent the poses of each camera, and edges represent the transformation between camera poses. The pose graph not only intuitively describes the visual odometry but also facilitates understanding of changes in camera poses. Nonlinear global optimization is expressed as graph optimization. The same scene does not appear in many locations, making the pose graph sparse. A sparse BA algorithm is used to solve the pose graph to correct the camera pose.

[0059] Preferably, the drone hardware system includes the following four parts: an onboard computer; an onboard module assembly; a camera and a gimbal; an M100 quad-rotor drone, wherein the M100 quad-rotor drone provides a flight platform to realize basic flight functions; the camera and the gimbal are image acquisition components in the onboard hardware system; the onboard module assembly is a hardware assembly part located between the camera gimbal, the M100 quad-rotor drone and the onboard computer: 1) the video data of the camera is collected and sent to the onboard computer, 2) the onboard computer controls the gimbal through the onboard module assembly, 3) the onboard computer can realize flight control of the M100 quad-rotor drone through the onboard module assembly, 4) the video acquisition image data of the camera can be input into the image transmission system of the M100 quad-rotor drone through the onboard module assembly; 5. voltage conversion, the onboard module assembly converts the 24V voltage obtained from the M100 drone battery into 12V to power the gimbal and the onboard computer;

[0060] (1) Onboard computer

[0061] The onboard computer uses the NVIDIA JETSON TX2 RTS-ASG003 microcomputer, with a total weight of 170g;

[0062] (2) Airborne module assembly

[0063] The airborne module assembly is an intermediate execution processing unit in the entire airborne hardware system. Its internal composition and implementation are as follows: 1) The video data output by the high-definition camera is divided into two paths through the HDMI distributor. One path goes to the video collector and is output by the video collector to the airborne computer; the other path is output to the wireless image transmission system of the M100 UAV through the N1 encoder; 2) The USB to UART and PWM modules are used by the airborne computer to control the flight control and gimbal control of the M100; 3) The visual sensor realizes autonomous obstacle avoidance of the M100 UAV; 4) The RC receiver receives the control signal of the ground remote control to control the movement of the gimbal; 5) The wireless data transmission module provides a low-bandwidth data link between the airborne hardware system and the ground system; 6) It supplies power to the airborne computer, gimbal, and various components within the airborne module assembly;

[0064] The video acquisition part of the airborne module assembly consists of three parts: an HDMI distributor, a video collector, and an N1 encoder. The HDMI distributor splits the video stream from the HD camera into two channels, which then enter the video collector and N1 encoder respectively through the HDMI interface.

[0065] In the airborne module assembly, the wireless data transmission module and the ground-side wireless data transmission module establish a low-bandwidth wireless data transmission link, providing a bidirectional data path for data transmission between the air and ground terminals. The Guidance establishes a relatively independent visual sensing system equipped with five sets of visual and ultrasonic combination sensors to monitor the environment in multiple directions in real time and detect obstacles. In conjunction with the UAV flight controller, it can timely avoid potential collisions during high-speed flight. The RC wireless receiver R7008SB receives control signals from the ground remote control, processes them internally, and outputs PWM waveforms to control the gimbal's motion. The receiver has 16 receive channels. The USB to UART and PWM module performs two functions: completing USB to UART (TTL level) conversion, connecting the module's UART interface to the M100's UART interface, enabling the onboard computer to control the M100's flight, and achieving autonomous flight. The USB to PWM converter enables the onboard computer to output PWM control signals through the module. The module's PWM control signals are then connected to the gimbal's heading and pitch control signals, enabling the onboard computer to control the gimbal's motion.

[0066] (3) Camera and gimbal

[0067] A GoPro Hero4 HD camera is used to collect video data, and a MiNi3DPro gimbal is used with the GoPro Hero4 HD camera to control the camera's viewing angle. The gimbal is a three-axis gimbal that can achieve motion control in three directions: pitch, roll, and heading. There are two control methods for the gimbal: one is to use the onboard computer to control the gimbal through the USB to UART and PWM module to output PWM waveforms to achieve control, and the other is to use the ground remote control to control the onboard gimbal receiver RS7008SB to output PWM waveforms to control the gimbal.

[0068] Set the control signals for the pan and tilt axes of the gimbal. The control signals for both the pan and tilt axes are PWM waveforms with a period of 50 Hz. Position control of the pan and tilt axes is achieved by adjusting the duty cycle of the control signals. A duty cycle of 5.1% corresponds to the minimum position, 7.6% corresponds to the equilibrium position, and 10.1% corresponds to the maximum position.

[0069] Set the gimbal's mode control signal to achieve control in three modes: lock mode, heading and pitch follow mode, and heading follow mode. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 5% and 6%, the gimbal enters lock mode, in which the heading, pitch, and roll are all locked, and the heading and pitch are controlled by the remote control or onboard computer. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 6% and 9%, the gimbal enters heading and pitch follow mode, in which the roll is locked, the heading rotates smoothly in the direction of the head, and the pitch rotates with the aircraft's pitch angle. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 9% and 100%, the MiNi3D Pro gimbal enters heading follow mode, in which the heading, pitch, and roll are all locked, the heading rotates smoothly in the direction of the head, and the pitch is controlled by the remote control or onboard computer.

[0070] Preferably, the drone hardware system is integrated: the unit modules of the drone onboard module assembly are set, and the gimbal power supply system is designed. The 25V voltage output by the drone battery is output as 19V, 12V and 5V voltages respectively after passing through the three-way DCDC voltage conversion module of the onboard module assembly, of which the 19V voltage is used to power the onboard computer; the 12V voltage is used to power the gimbal; the 5V voltage is used to power the RC receiver and the HDMI distributor. The internal power supply system of the drone platform powers the Guidence visual sensor, N1 encoder, onboard wireless image transmission, and wireless data transmission respectively. The HDMI video collector and the onboard wireless data transmission are powered by the USB interface of the onboard computer, and the camera is powered by the camera's own battery.

[0071] Preferably, the integration of the drone software system is completed using the ROS system platform, and the ROS message mechanism is used for communication between modules, while the coupling between the two modules is very loose;

[0072] 1) Target Detection Process

[0073] The object detection system uses the Jetson Inference system, which includes classification, detection, and segmentation. The detection modules include image acquisition detection, video detection, and real-time camera detection. The video stream is ultimately decomposed into acquisition frames. The essence of these detections is image acquisition detection.

[0074] The target detection system converts the video stream into captured image frames and then uses the trained network model to detect the captured images. It ultimately obtains the target category and pixel coordinates of the target frame. The output of the target detection system is then passed to the visual real-time spatial perception modeling system for positioning and navigation.

[0075] 2) 3D visual RGB space extraction process

[0076] The 3D visual RGB space extraction system receives RGB acquisition images and depth maps, first matches the RGB acquisition images to obtain key frames, then builds a point cloud image with the depth acquisition image, then corrects the point cloud image through nonlinear global optimization and local detection, and finally receives the target detection results to complete the target positioning in 3D space.

[0077] 3) Integration of UAV onboard processing methods

[0078] The onboard processing process of the software system is that after TX2 receives the captured image data from the camera, the target detection module first detects the content of the captured image to obtain the target position coordinates, and then transmits it to the visual real-time spatial perception modeling system in real time. At this time, the visual real-time spatial perception modeling system is composing the key frames and reconstructing the target detection results corresponding to the key frames in three-dimensional space to realize spatial content extraction.

[0079] Compared with the existing technology, the innovation and advantages of this application are:

[0080] (1) This application improves the spatial perception module of the UAV by integrating the three-dimensional visual RGB space extraction method, optimizes the target detection method to increase the real-time target precise positioning and target precise identification capabilities of the UAV, and improves the content extraction module. Finally, by integrating the two methods, the UAV is provided with a navigation and path planning basis, and an intelligent UAV system including software and hardware is established. The content extraction method based on convolutional neural network target detection is used to provide the UAV with a real-time three-dimensional map to help the UAV accurately locate, while the content extraction method based on target detection can identify the target in the scene and give the target's position. The combination of the two provides meaningful reference data for the obstacle avoidance and path planning of the UAV, realizing the non-intervention or semi-intervention flight of the UAV. The spatial scene perception method embedded in the microcomputer on the UAV and the content extraction method based on target detection are used to establish an intelligent processing system on the UAV. The two methods achieve real-time processing. The implementation of the real-time spatial perception and target identification-based content extraction method on the UAV is not limited to solving the obstacle avoidance and path planning problems, but can also make efficient use of the data collected. In the case of an unknown environment where the UAV needs to explore while flying, the obstacle avoidance and path planning can be successfully completed. It can be extended to other application directions, which has a great application value for the in-depth exploration and expansion of UAV applications.

[0081] (2) This application constructs a fast deep convolutional neural network feature extractor to learn and extract the features of all targets of interest for use in specific target detection and recognition; uses single-purpose ORB-visual real-time spatial perception modeling to achieve rapid composition of specific scenes; integrates the target detection module to mark the target of interest in the visual real-time spatial perception modeling diagram and gives specific location information; the target detection method of this application has real-time performance while ensuring a certain accuracy rate, applies the machine vision target detection method to content-based scene perception, integrates the space-based scene perception method, and establishes an intelligent drone system. Using drones and JETSON TX2 as the carrier and development platform, it realizes the combination of visual real-time spatial perception modeling and target detection methods of drones in specific scenarios. It is used for drones whose flight speed is restricted by the scene, has high spatial perception efficiency, fast target detection speed, and high accuracy, and achieves the successful completion of obstacle avoidance and path planning when the environment is unknown and the drone needs to explore while flying.

[0082] (3) This application establishes a sample library of indoor targets of interest and a feature library of targets of interest, and uses an end-to-end fast neural network to achieve real-time target detection and matching. While detecting the target, it uses a three-dimensional visual RGB space extraction system for real-time composition, annotates the identified targets of interest in the simulation map, and gives specific location information, especially the core method; 1) Based on visual real-time spatial perception modeling method: the depth camera is mounted on the drone to provide RGB acquisition map and depth acquisition map for real-time spatial perception modeling, and finally provides a three-dimensional point cloud map for the drone, gives the spatial positioning of the drone, and provides data support for drone navigation; 2) Target detection content extraction based on deep learning: based on drone flight With high speed, the target detection network optimizes the real-time processing capability, and combines the conversion relationship obtained by scene perception to transform the target detected into three-dimensional space, provide the spatial position of various targets for the UAV, and use the spatial relationship for obstacle avoidance and path planning of the UAV; 3) Establish a real-time perception and target detection system for UAV: Based on the optimization of real-time scene perception and real-time target detection, the implementation of the two modules is embedded in the microcomputer, and an intelligent system for real-time perception and target detection suitable for UAVs including software and hardware is established. The system has good stability, high spatial perception efficiency, fast target detection speed, high accuracy, and better solves the problem of obstacle avoidance and path planning of UAVs, realizing non-intervention or semi-intervention flight of UAVs. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 It is the overall block diagram of the onboard hardware of the UAV hardware system.

[0084] Figure 2 It is a system block diagram of the UAV hardware onboard module assembly.

[0085] Figure 3 This is a connection diagram of the USB to UART and PWM module output pins and the pan / tilt control signal line.

[0086] Figure 4 This is a schematic diagram of the control signals for setting the pan and tilt axes of the gimbal.

[0087] Figure 5 This is a diagram of the control signal for setting the gimbal mode.

[0088] Figure 6 It is a schematic diagram of setting up each unit module of the UAV airborne module assembly.

[0089] Figure 7 It is a schematic diagram of the overall framework of the UAV onboard processing method.

[0090] Figure 8 This is a schematic diagram of the outdoor application case of the drone in this application.

[0091] Figure 9 This is a schematic diagram of the indoor application case of the drone in this application. DETAILED DESCRIPTION

[0092] The following, in conjunction with the accompanying drawings, further describes the technical solution of the intelligent drone real-time spatial perception and target precision detection system provided by the present application, so that those skilled in the art can better understand the present application and implement it.

[0093] With the significant improvement in software and hardware performance, machine vision has also rapidly improved in terms of accuracy and real-time performance. In recent years, content-based scene perception methods and space-based scene perception methods have increased. Target detection methods have achieved real-time performance while ensuring a certain accuracy rate. By applying machine vision target detection methods to content-based scene perception and integrating them with space-based scene perception methods, an intelligent UAV system is established. Using UAVs and JETSON TX2 as the carrier and development platform, this system combines visual real-time spatial perception modeling and target detection methods for UAVs in specific scenarios. This system is suitable for UAVs whose flight speed is restricted by the scenario.

[0094] (1) Content-based scene perception method: A deep learning-based target detection method is used to complete the content-based scene perception task. UAVs as a carrier platform are different from ordinary target detection. First, the UAV itself has a certain flight speed, which puts forward real-time requirements for target detection. Secondly, the UAV approaches the object from far to near during the flight and passing by the object. At the same time, the shooting angle will also have frontal, oblique, and downward viewing due to the relative position of the UAV and the object. In this way, the target detection method requires the characteristics of rotation and scale invariance. The neural network is optimized to complete the task of real-time content extraction of the UAV. When the amount of data is sufficient, the target in the specific scene can be accurately identified through training.

[0095] (2) Space-based scene perception method: The fast ORB corner detection method is used to complete matching and achieve fast composition, and the visual real-time spatial perception modeling method is integrated into the real-time spatial perception of specific UAV scenes.

[0096] (3) Intelligent UAV System: The spatial perception module based on the visual real-time spatial perception modeling method and the target detection content extraction method based on the convolutional neural network are embedded in the microcomputer and combined with the UAV system to form an intelligent UAV system. Through experiments, the problems encountered during actual flight with the two perception modules are solved.

[0097] By integrating a 3D visual RGB spatial extraction method to improve the UAV's spatial perception module, and optimizing the target detection method to enhance the UAV's real-time target positioning and recognition capabilities, the content extraction module is improved. Ultimately, by integrating the two methods, the UAV is provided with a basis for navigation and path planning, and an intelligent UAV system, including software and hardware, is established. First, a fast deep convolutional neural network feature extractor is constructed to learn and extract the features of all targets of interest for specific target detection and recognition. Then, a single-purpose ORB-based visual real-time spatial perception model is used to achieve rapid composition of specific scenes. Finally, the integrated target detection module marks the targets of interest in the visual real-time spatial perception modeling diagram and provides specific location information.

[0098] This application establishes a sample library and feature library of indoor targets of interest, uses an end-to-end fast neural network to achieve real-time target detection and matching, and uses a three-dimensional visual RGB space extraction system to perform real-time composition while detecting targets. The identified targets of interest are marked in the simulation map and their specific location information is given. The core methods include:

[0099] (1) Based on the visual real-time spatial perception modeling method: the depth camera is mounted on the UAV to provide RGB acquisition images and depth acquisition images for real-time spatial perception modeling. The real-time spatial perception modeling, i.e., 3D visual RGB space extraction, is used to complete the perception of the spatial scene. Finally, a 3D point cloud map is provided for the UAV, which gives the spatial positioning of the UAV and provides data support for the UAV navigation.

[0100] (2) Object detection content extraction based on deep learning: Based on the speed of UAV flight, the network optimizes the real-time processing capability in object detection, combines the conversion relationship obtained by scene perception, transforms the object detected into three-dimensional space, provides the spatial position of various objects for UAV, and uses the spatial relationship to perform obstacle avoidance and path planning.

[0101] (3) Establish a real-time perception and target detection system for UAVs: Based on the optimization of real-time scene perception and real-time target detection, the implementation of the two modules is embedded in a microcomputer to establish an intelligent system including software and hardware for real-time perception and target detection suitable for UAVs.

[0102] 1. Content Extraction Method for Real-time Precision Target Detection

[0103] (1) Target real-time precision detection network data format

[0104] The training data samples include large collection images of multiple objects. For each object in the collection image, the training label includes not only the class of the object, but also the coordinates of the corner points of the bounding box. The number of objects varies between different training collection images. By introducing a fixed three-dimensional label format, the problem that the selection of label formats of different lengths and dimensions makes it difficult to define the loss function is solved. The defined format can input collection images of any size including any number of objects.

[0105] The collected image is divided into regular grids with a grid size slightly smaller than the smallest object expected to be detected. Each grid has two key pieces of information, including the category of the object and the coordinates of the corner points of the grid that contains the object. In addition, when there is no object in the grid, a special custom class, the "dontcare" class, is used to uniformly maintain a fixed size in data representation, and an object coverage value represented by 0 or 1 is also set to indicate whether there is an object in the grid. For the case where many objects are in the same grid, the object that occupies the most pixels in the grid is selected. In the case of overlapping objects, the object with the bounding box with the smallest Y value is used.

[0106] (2) Target Real-time Precision Detection Network Framework

[0107] The target real-time precision detection network training is divided into three steps:

[0108] Step 1: The data layer obtains training collection images and labels, and the conversion layer performs online data enhancement;

[0109] Step 2: The fully convolutional network extracts and predicts the object class and bounding box of each grid;

[0110] Step 3: Predict the object category and target bounding box for each grid separately, and then use the loss function to calculate the error of the two prediction tasks simultaneously;

[0111] The prediction process consists of two steps: first, during the validation process, a clustering function is used to generate the final bounding box set; second, the model performance is measured by a simplified mAP calculation value on the validation dataset;

[0112] The network accepts input images of varying sizes and effectively applies a CNN in a sliding window fashion with a stride, outputting a multidimensional array that is superimposed on the image. Using GoogLeNet with the final pooling layer removed, the CNN uses a sliding window of up to 555×555 pixels with a stride of 16 pixels.

[0113] The final optimized loss function is generated using a linear combination of two independent loss functions, which include the sum of squares of the difference between the true and predicted object coverage of all grids in the training data samples, and the mean absolute difference loss between the true and predicted corner points of the bounding box of the object covered by each grid.

[0114] 2. Visual Real-time Spatial Perception Modeling Method

[0115] 1. Real-time spatial perception modeling framework

[0116] The process of real-time spatial perception modeling includes:

[0117] Process 1: Reading sensor image data: This involves reading and preprocessing the image information collected by the drone camera in real-time spatial perception modeling. The depth camera data includes the RGB image and its corresponding depth map.

[0118] Process 2: Visual odometry modeling: Visual odometry estimates the rotation and translation relationship between two adjacent captured images to calculate the camera's pose change and local map. The key to this step is feature point extraction and captured image matching.

[0119] Process 3: Back-end global optimization: The back-end uses a nonlinear global optimization algorithm to optimize the camera position and posture from the front-end and the past-and-back detection results obtained by another thread, correcting them to a globally unified trajectory map and point cloud map.

[0120] Process 4: Past-location detection: This process determines whether the sensor or the drone carrying the sensor has passed through a certain scene before. If it is detected that the sensor has been there before, the information is provided to the backend to correct the position and posture.

[0121] Process 4, composition: Based on the estimated camera trajectory, a drone cruise map that meets the mission requirements is created.

[0122] 1. Visual Odometry Modeling

[0123] The front-end receives the camera's video stream, or captured image sequences, and estimates the camera's motion between adjacent frames using feature matching methods. This allows for the initial acquisition of mileage information with a certain degree of error accumulation. Visual odometry modeling consists of four parts:

[0124] First, the acquisition frame: The information carried includes the drone camera pose, RGB acquisition image, and depth map when the frame was captured;

[0125] Second, the camera model: corresponds to the camera used in the actual shooting, and only includes internal parameters;

[0126] Third, local map: including keyframes and landmark information points. Keyframes and landmark information points that meet the matching rules will be added to the map. The map is only a local map and not a global map. It only includes landmark information points near the current location, and farther landmark information points are deleted.

[0127] Fourth, landmark information points: These are map points with known information in the map. The known information included in the landmark information points is the feature description corresponding to them, and the acquisition method is to use feature matching algorithms for batch extraction.

[0128] 2. Backend global optimization

[0129] The global process analyzes the data and handles noise problems, including linear global optimization and nonlinear global optimization algorithms;

[0130] Linear global optimization assumes that there is a linear relationship between each frame of the captured image during the shooting process, and uses the Kalman filter algorithm for state estimation. If it is assumed that only the previous frame and the next frame have a linear relationship, the state estimation is completed by extending the Kalman filter; the difference between the observed value and the algorithm estimated value is calculated, that is, the error value between the pixel coordinates and the pixel coordinates of the corresponding 3D point projected onto the two-dimensional plane through the camera position. Linear global optimization assumes that the error between the camera position and the spatial point has a causal relationship. The camera position and posture are first calculated, and then the position of the spatial point is further calculated based on the camera position; nonlinear global optimization directly puts all data into the same model for optimization and solution, downplaying the relationship between the data.

[0131] 3. Testing of travel to and from the old place

[0132] The key to correcting the accumulated errors of the visual odometry through back-and-forth detection is to establish a bag-of-words model. The bag-of-words model abstracts features into individual words. The detection process is to match the words that appear in the two images to determine whether the two images describe the same scene. To classify features into words, it is necessary to train a dictionary that includes all possible word sets. Building a dictionary requires massive data for training. The establishment of the dictionary is a clustering process. Assuming that a total of 100 million features are extracted from all images, they are clustered into 100,000 words using the K-means clustering method. A tree with k branches and a depth of d is constructed for the dictionary during training. The upper nodes of the tree provide coarse classification, and the lower nodes provide fine classification, extending all the way to the leaf nodes. Using this tree, the time complexity is reduced to the logarithmic level, which speeds up the feature matching.

[0133] 4. Composition

[0134] The data collected by the camera is optimized and the camera posture is corrected to convert the two-dimensional plane points into three-dimensional space, forming three-dimensional space point cloud information. In addition to the point cloud map, the camera posture optimization process is represented in the g2o tool to form a posture graph. Depending on the specific situation, you can also define your own map for description.

[0135] (2) 3D visual RGB space extraction method

[0136] The sensor uses a monocular camera to obtain depth information, and the data source includes RGB acquisition image and depth map.

[0137] Use the depth camera to obtain the color acquisition image and the depth acquisition image, and use the geometric model to convert the 2D plane data into a 3D stereo space. From the pixel coordinates to the acquisition image coordinates: the center points of the coordinate systems are different, and there is only an offset relationship between the two. From the acquisition image coordinate system to the camera coordinate system, the coordinate axes are parallel, and there is only a scaling relationship. The conversion relationship from the pixel coordinate system to the camera coordinate system is expressed as:

[0138]

[0139] Where u and v are the offsets between the origins of the coordinate system and the center of the acquisition plane, d x , d y is the scaling ratio between pixel coordinates and the actual imaging plane, and the scaling ratio is d x =z c / f x , d y =z c / f y , f x , f y is the focal length of the camera on the x and y axes, written in matrix form:

[0140]

[0141] When the camera moves, the camera coordinate system and the world coordinate system are not parallel, and there is a rotation and translation relationship. In the subsequent visual odometry calculation, the relationship between the previous and next frames is the same as above, and the matrix relationship is given as:

[0142]

[0143] Convert the points on the two-dimensional plane to three-dimensional space, and finally obtain a series of point cloud data, which are assigned RGB color attributes to initially obtain a color three-dimensional map.

[0144] 1. Implementation of 3D visual RGB space extraction

[0145] (1) Front-end visual odometry

[0146] Initialization is based on the first frame of the acquisition image, and the search for key frames begins. The ORB algorithm is used to extract key points between the acquisition images, and then the BRIEF descriptor is calculated for each key point. Finally, the fast approximate nearest neighbor algorithm is used for fast matching. The ORB corner point extraction algorithm adds scale and rotation descriptions to the FAST corner point extraction algorithm, adding feature information, richer feature descriptions, higher matching accuracy, and more accurate and reliable composition. The BRIEF descriptor uses a binary descriptor and uses randomly selected points for comparison.

[0147] After the matching is completed, the 2D points are projected into 3D space according to the depth acquisition map to obtain the 2D coordinates and corresponding 3D coordinates of a series of points. The camera's position is estimated by solving the PnP problem. The actual calculation result is the rotation and translation matrix between the two frames of acquisition. All data are matched in sequence and the camera pose is calculated to obtain a complete visual odometry.

[0148] (2) Back-end nonlinear global optimization

[0149] During the establishment of the front-end visual odometry, only two adjacent frames of captured images are continuously matched and the corresponding camera posture is solved. This will inevitably lead to error accumulation. At this time, it is necessary to consider correcting the camera posture.

[0150] Three-dimensional visual RGB space extraction uses a pose graph to represent the pose calculated by the visual odometry, including nodes and edges. Nodes represent the poses of each camera, and edges represent the transformation between camera poses. The pose graph not only intuitively describes the visual odometry but also facilitates understanding of changes in camera poses. Nonlinear global optimization is expressed as graph optimization. The same scene does not appear in many locations, making the pose graph sparse. A sparse BA algorithm is used to solve the pose graph to correct the camera pose.

[0151] (3) Testing of travel to and from the old place

[0152] Although the back-end global optimization has been done and the camera posture has been corrected to a certain extent, there is still a problem. That is, after a period of time, when the drone returns to the origin or the same place, can the system identify that it is the origin or has been there before? This is the problem to be solved by the old place travel detection. If the same scene is identified by collecting map matching, the back-end can obtain another optimization information to adjust the trajectory and map to meet the old place travel detection results and complete the global optimization. To detect whether the same place has been visited, it is necessary to compare the similarity between the current frame and all previous frames. However, the longer the time, the larger the amount of data, which will greatly reduce the real-time performance. In order to achieve better results, the 3D visual RGB space extraction uses both close-range loopback and random loopback to replace the traversal method. Close-range loopback is to match the current frame with the previous n frames, and n is selected according to the situation; random loopback is to match the current frame with n frames randomly selected from all previous frames, and n is also selected according to the situation. After the old place travel detection, the camera position is further corrected.

[0153] 3. UAV scene perception and target detection system

[0154] (1) UAV hardware system

[0155] The overall block diagram of the onboard hardware is as follows: Figure 1 As shown, it includes the following four parts: an onboard computer; an onboard module assembly; a camera and a gimbal; an M100 quad-rotor drone, wherein the M100 quad-rotor drone provides a flight platform to realize basic flight functions; the camera and the gimbal are image acquisition components in the onboard hardware system; the onboard module assembly is a hardware assembly part located between the camera gimbal, the M100 quad-rotor drone and the onboard computer: 1) the video data of the camera is collected and sent to the onboard computer, 2) the onboard computer controls the gimbal through the onboard module assembly, 3) the onboard computer can realize flight control of the M100 quad-rotor drone through the onboard module assembly, 4) the video acquisition image data of the camera can be input into the image transmission system of the M100 quad-rotor drone through the onboard module assembly; 5. voltage conversion, the onboard module assembly converts the 24V voltage obtained from the M100 drone battery into 12V to power the gimbal and the onboard computer.

[0156] 1. UAV hardware system composition

[0157] (1) Onboard computer

[0158] The onboard computer uses the NVIDIA JETSON TX2 RTS-ASG003 microcomputer, with a total weight of 170g and is only the size of a bank card, making it very light.

[0159] (2) Airborne module assembly

[0160] The airborne module assembly is an intermediate execution processing unit in the entire airborne hardware system, such as Figure 2 The figure shows a system block diagram of the onboard module assembly. Its internal components and implementation are as follows: 1) The video data output by the HD camera is split into two paths through an HDMI splitter. One path is fed into the video capture device, which then outputs it to the onboard computer; the other path is output via the N1 encoder to the M100 drone's wireless image transmission system. 2) A USB-to-UART and PWM module is used by the onboard computer to control the M100's flight control and gimbal. 3) A visual sensor enables autonomous obstacle avoidance for the M100 drone. 4) An RC receiver receives control signals from a ground remote control to control the gimbal. 5) The wireless data transmission module provides a low-bandwidth data link between the onboard hardware system and the ground system. 6) Power is provided to the onboard computer, gimbal, and various components within the onboard module assembly.

[0161] The video acquisition part of the airborne module assembly consists of three parts: HDMI distributor, video collector and N1 encoder. The HDMI distributor divides the video stream from the high-definition camera into two channels, and the two video streams enter the video collector and N1 encoder respectively through the HDMI interface.

[0162] In the airborne module assembly, the wireless data transmission module and the ground-side wireless data transmission module establish a low-bandwidth wireless data transmission link, providing a bidirectional data path for data transmission between the air and ground terminals. The Guidance establishes a relatively independent visual sensing system equipped with five sets of visual and ultrasonic combination sensors to monitor the environment in multiple directions in real time and detect obstacles. In conjunction with the UAV flight controller, it can timely avoid potential collisions during high-speed flight. The RC wireless receiver R7008SB receives control signals from the ground remote control, processes them internally, and outputs PWM waveforms to control the gimbal's motion. The receiver has 16 receive channels. The USB to UART and PWM module performs two functions: completing USB to UART (TTL level) conversion and connecting the module's UART interface to the M100's UART interface, enabling the onboard computer to control the M100's flight and achieve autonomous flight. The USB to PWM converter enables the onboard computer to output PWM control signals through the module. The module's PWM control signals are then connected to the gimbal's heading and pitch control signals, enabling the onboard computer to control the gimbal's motion. Figure 3 Describes the connection between the USB to UART and PWM module output pins and the pan / tilt control signal lines.

[0163] (3) Camera and gimbal

[0164] A GoPro Hero4 HD camera is used to collect video data, and a MiNi3DPro gimbal is used with the GoPro Hero4 HD camera to control the camera's viewing angle. The gimbal is a three-axis gimbal that can achieve motion control in three directions: pitch, roll, and heading. There are two control methods for the gimbal: one is that the onboard computer outputs PWM waveforms through a USB to UART converter and a PWM module to control the gimbal; the other is that the ground remote control controls the onboard gimbal receiver RS7008SB to output PWM waveforms to control the gimbal.

[0165] Figure 4 Set the control signals for the pan and tilt axes of the gimbal. The control signals for the pan and tilt axes are both PWM waveforms with a period of 50 Hz. The position of the pan and tilt axes is controlled by adjusting the duty cycle of the control signal. A duty cycle of 5.1% corresponds to the minimum position, 7.6% corresponds to the balance position, and 10.1% corresponds to the maximum position.

[0166] Figure 5Set the gimbal's mode control signal to achieve control in three modes: lock mode, heading and pitch follow mode, and heading follow mode. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 5% and 6%, the gimbal enters lock mode, in which the heading, pitch, and roll are all locked, and the heading and pitch are controlled by the remote control or onboard computer. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 6% and 9%, the gimbal enters heading and pitch follow mode, in which the roll is locked, the heading rotates smoothly in the direction of the head, and the pitch rotates with the aircraft's pitch angle. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 9% and 100%, the MiNi3D Pro gimbal enters heading follow mode, in which the heading, pitch, and roll are all locked, the heading rotates smoothly in the direction of the head, and the pitch is controlled by the remote control or onboard computer.

[0167] (4) UAV interface

[0168] The drone's battery power output is connected to the onboard module assembly power input, powering the onboard equipment. The visual obstacle avoidance system's output is connected to the drone's CAN-Bus, working with the drone's flight controller to achieve autonomous obstacle avoidance. The N1 encoder's power and video interfaces are both connected to dedicated interfaces on the drone.

[0169] 2. UAV hardware system integration

[0170] Figure 6 Set up the various unit modules of the drone's onboard module assembly and design the gimbal power supply system. The 25V voltage output by the drone battery is converted into 19V, 12V and 5V voltages after passing through the three-way DCDC voltage conversion module of the onboard module assembly. The 19V voltage is used to power the onboard computer; the 12V voltage is used to power the gimbal; the 5V voltage is used to power the RC receiver and the HDMI distributor. The internal power supply system of the drone platform powers the Guidence visual sensor, N1 encoder, onboard wireless image transmission, and wireless data transmission respectively. The HDMI video capturer and onboard wireless data transmission are powered by the USB port of the onboard computer, and the camera is powered by its own battery.

[0171] (2) UAV software system design

[0172] The integration of the drone software system is completed using the ROS system platform, and the ROS message mechanism is used for communication between modules. At the same time, the coupling between the two modules is very loose, which facilitates further improvement of the effect in the future.

[0173] 1. Object Detection Process

[0174] The target detection system uses the Jetson-Inference system, which includes classification, detection, and segmentation. The detection module includes acquisition image detection, video detection, and real-time camera detection. The video stream is ultimately decomposed into acquisition image frames. The essence of these detections is acquisition image detection.

[0175] The object detection system converts the video stream into captured image frames and then uses a trained network model to detect the captured images, ultimately obtaining the target category and pixel coordinates of the target frame. The output of the object detection system is then passed to the visual real-time spatial perception modeling system for positioning and navigation.

[0176] 2. 3D visual RGB space extraction process

[0177] The three-dimensional visual RGB space extraction system needs to receive RGB acquisition images and depth maps, first match the RGB acquisition images to obtain key frames, then build a point cloud image with the depth acquisition image, and then perform nonlinear global optimization and local detection to correct the point cloud image. Finally, the target detection results are received to complete the target positioning in the three-dimensional space.

[0178] 3. Integration of UAV onboard processing methods

[0179] The overall framework of the UAV onboard processing method is as follows Figure 7 As shown in the figure, the onboard processing process of the software system is as follows: after the TX2 receives the acquisition image data from the camera, the target detection module first detects the content of the acquisition image to obtain the target position coordinates, and then transmits them in real time to the visual real-time spatial perception modeling system. At this time, the visual real-time spatial perception modeling system is composing images based on key frames. Experimental results show that since the continuous acquisition images after the video stream is decomposed into acquisition frames include key frames, and key frames inevitably also include target detection results, the target detection results corresponding to the key frames are directly reconstructed in 3D space to achieve spatial content extraction.

[0180] 4. UAV Outdoor Application Cases

[0181] The experiment was conducted in a playground with many pedestrians, which fully tested the speed and accuracy of target detection on a microcomputer, as well as the adaptability of target detection in flight.

[0182] like Figure 8 Experimental results show that the drone-mounted microcomputer using DetectNet performs very well in detecting dense crowds, with virtually no missed detections. The performance is comparable at slow speeds, and it also performs well at high speeds. The only issue is that the detection frame may briefly shift when the camera rotates, but it quickly returns to its correct position. This demonstrates that this application can perfectly meet the goal of real-time and accurate target detection in practical applications.

[0183] 5. UAV indoor application cases

[0184] The experiment was conducted indoors with a relatively complex scene. The indoor environment is characterized by a small space and many obstacles. If the GPS navigation system fails when the drone is flying indoors, the navigation of the drone must be completed by fully utilizing the fusion of inertial navigation and visual perception. First, inertial navigation has a relatively high accuracy when the carrier changes direction instantly, but the error is large when running for a long time. In this case, the scene perception based on real-time spatial perception modeling can be used to locate the drone. The attitude of the carrier when it changes direction instantly, measured by inertial navigation, is then added to the back-end global optimization stage of the real-time spatial perception modeling to correct the attitude of the drone. At the same time, target detection is used to perceive the specific spatial position of the obstacle, providing a basis for drone navigation and path planning. Indoor composition and detection effects are as follows. Figure 9 As shown:

[0185] The system's output data is the spatial position of the detected target and the drone's posture. This demonstrates the relatively accurate spatial perception effect, and the normalized target frame provides the drone with a precise target position. This demonstrates that precise indoor positioning of drones can also be achieved through a combination of real-time spatial perception modeling and target detection.

[0186] VI. Summary

[0187] After integrating real-time spatial perception modeling, target detection, and the UAV system, the drone itself also carries the UAV system for control, positioning navigation, path planning, and other functions. Therefore, when integrating other applications, there is still much work to be done to ensure the coordination between these modules and improve the effectiveness. Using the drone software platform ROS, real-time perception and real-time target detection methods are integrated into the microcomputer, which together with the UAV system form an intelligent UAV system.

[0188] (1) The UAV’s spatial scene perception based on SLAM is realized, and real-time spatial perception modeling is used to establish a three-dimensional point cloud map and positioning to provide spatial position information for the UAV.

[0189] (2) The content scene perception based on the target detection method of DetectNet is realized, which provides the UAV with the position of the specific target in three-dimensional space, which can be used for intelligent navigation and path planning of the UAV.

[0190] (3) The drone itself does not have devices such as cameras and microcomputers. This study also realized the hardware structure design of the intelligent drone and completed the framework design for the coordinated work of drones equipped with specific equipment.

Claims

1. Intelligent UAV real-time spatial perception and target precision detection system, characterized by: By integrating a 3D visual RGB spatial extraction method to improve the UAV's spatial perception module, and optimizing the target detection method to enhance the UAV's real-time target positioning and recognition capabilities, the content extraction module is improved. Ultimately, by integrating the two methods, the UAV is provided with a basis for navigation and path planning, and an intelligent UAV system, including software and hardware, is established. First, a fast deep convolutional neural network feature extractor is constructed to learn and extract the features of all targets of interest for specific target detection and recognition. Then, a single-purpose ORB-based visual real-time spatial perception model is used to achieve rapid composition of specific scenes. Finally, the integrated target detection module marks the targets of interest in the visual real-time spatial perception modeling diagram and provides specific location information. This application establishes a sample library and a feature library of indoor targets of interest, uses an end-to-end fast neural network to achieve real-time target detection and matching, and uses a three-dimensional visual RGB space extraction system to perform real-time composition while detecting targets. The identified targets of interest are marked in the simulation map and their specific location information is given. The core methods include: (1) Based on the visual real-time spatial perception modeling method: the depth camera is mounted on the UAV to provide RGB acquisition images and depth acquisition images for real-time spatial perception modeling. The real-time spatial perception modeling, i.e., 3D visual RGB space extraction, is used to complete the perception of the spatial scene. Finally, a 3D point cloud map is provided for the UAV, which gives the spatial positioning of the UAV and provides data support for the UAV navigation. (2) Object detection content extraction based on deep learning: Based on the speed of UAV flight, the object detection network optimizes real-time processing capabilities, combines the conversion relationship obtained by scene perception, transforms the target detected into three-dimensional space, provides the spatial position of various targets for the UAV, and uses the spatial relationship to perform obstacle avoidance and path planning. (3) Establish a real-time perception and target detection system for UAVs: Based on the optimization of real-time scene perception and real-time target detection, the implementation of the two modules is embedded in a microcomputer to establish an intelligent system including software and hardware for real-time perception and target detection suitable for UAVs.

2. The intelligent unmanned aerial vehicle real-time spatial perception and target precision detection system according to claim 1 is characterized in that: Data format for real-time precision object detection networks: Training data samples include large collection images of multiple objects. For each object in the collection image, the training label includes not only the object class but also the coordinates of the corner points of the bounding box. The number of objects varies between different training collection images. By introducing a fixed three-dimensional label format, the problem of difficulty in defining the loss function due to the selection of label formats of different lengths and dimensions is solved. The defined format can input collection images of any size containing any number of objects. The collection image is divided into regular grids with a grid size slightly smaller than the smallest object expected to be detected. Each grid has two key pieces of information, including the category of the object and the coordinates of the corner points of the grid that contains the object. In addition, when there is no object in the grid, a special custom class, the "dontcare" class, is used to uniformly maintain a fixed size in data representation, and an object coverage value represented by 0 or 1 is also set to indicate whether there is an object in the grid. For the case where many objects are in the same grid, the object that occupies the most pixels in the grid is selected. In the case of overlapping objects, the object with the bounding box with the smallest Y value is used.

3. The intelligent unmanned aerial vehicle real-time spatial perception and target precision detection system according to claim 1 is characterized in that: The target real-time precision detection network training is divided into three steps: Step 1: The data layer obtains training collection images and labels, and the conversion layer performs online data enhancement; Step 2: The fully convolutional network extracts and predicts the object class and bounding box of each grid; Step 3: Predict the object category and target bounding box for each grid separately, and then use the loss function to calculate the error of the two prediction tasks simultaneously; The prediction process consists of two steps: first, during the validation process, a clustering function is used to generate the final bounding box set; second, the model performance is measured by a simplified mAP calculation value on the validation dataset; The network accepts input images of varying sizes and effectively applies a CNN in a sliding window fashion with a stride, outputting a multidimensional array that is superimposed on the image. Using GoogLeNet with the final pooling layer removed, the CNN uses a sliding window of up to 555×555 pixels with a stride of 16 pixels. The final optimized loss function is generated using a linear combination of two independent loss functions, which include the sum of squares of the difference between the true and predicted object coverage of all grids in the training data samples, and the mean absolute difference loss between the true and predicted corner points of the bounding box of the object covered by each grid.

4. The intelligent unmanned aerial vehicle real-time spatial perception and target precision detection system according to claim 1 is characterized in that: The process of real-time spatial perception modeling includes: Process 1: Reading sensor image data: This involves reading and preprocessing the image information collected by the drone camera in real-time spatial perception modeling. The depth camera data includes the RGB image and its corresponding depth map. Process 2: Visual odometry modeling: Visual odometry estimates the rotation and translation relationship between two adjacent captured images to calculate the camera's pose change and local map. The key to this step is feature point extraction and captured image matching. Process 3: Back-end global optimization: The back-end uses a nonlinear global optimization algorithm to optimize the camera position and posture from the front-end and the past-and-back detection results obtained by another thread, correcting them to a globally unified trajectory map and point cloud map. Process 4: Past-location detection: This process determines whether the sensor or the drone carrying the sensor has passed through a certain scene before. If it is detected that the sensor has been there before, the information is provided to the backend to correct the position and posture. Process 4, composition: Based on the estimated camera trajectory, a drone cruise map that meets the mission requirements is created.

5. The intelligent unmanned aerial vehicle real-time spatial perception and target precision detection system according to claim 1 is characterized in that: Real-time spatial perception modeling framework: 1) Visual odometry modeling The front end is responsible for receiving the camera's video stream, i.e., the captured image sequence, and estimating the camera's motion between adjacent frames through feature matching methods, and initially obtaining mileage information with a certain amount of error accumulation. The visual odometry model consists of four parts: composition: First, the acquisition frame: The information carried includes the drone camera pose, RGB acquisition image, and depth map when the frame was captured; Second, the camera model: corresponds to the camera used in the actual shooting, and only includes internal parameters; Third, local map: including keyframes and landmark information points. Keyframes and landmark information points that meet the matching rules will be added to the map. The map is only a local map, not a global map. It only includes landmark information points near the current location, and landmark information points farther away are deleted. Fourth, landmark information points: These are map points with known information. The known information included in the landmark information points is the feature description corresponding to them. The acquisition method is to use feature matching algorithms to extract them in batches. 2) Backend global optimization The global process analyzes the data and handles noise problems, including linear global optimization and nonlinear global optimization algorithms; Linear global optimization assumes that there is a linear relationship between each frame of the captured image during the shooting process, and uses the Kalman filter algorithm to estimate the state. If it is assumed that only the previous frame and the next frame have a linear relationship, then the state estimation is completed through the extended Kalman filter. The difference between the observed value and the algorithm's estimated value is calculated, that is, the error value between the pixel coordinates and the pixel coordinates of the corresponding 3D point projected onto the two-dimensional plane through the camera position. Linear global optimization assumes that the error between the camera position and the spatial point has a causal relationship. The camera position and attitude are first calculated, and then the position of the spatial point is further calculated based on the camera position. Nonlinear global optimization directly puts all data into the same model for optimization and solution, downplaying the relationship between the data; 3) Testing of travel to and from the old place The key to correcting the accumulated errors of the visual odometry for past-time detection lies in building a bag-of-words model. This model abstracts features into individual words. The detection process involves matching the words that appear in two images to determine whether they depict the same scene. To classify features into words, it is necessary to train a dictionary that includes all possible word sets. Building this dictionary requires massive amounts of data for training. Dictionary building is a clustering process. Assuming that a total of 100 million features are extracted from all images, they are clustered into 100,000 words using the K-means clustering method. During the dictionary training process, a tree with k branches and a depth of d is constructed. The upper nodes of the tree provide coarse classification, and the lower nodes provide fine classification, extending all the way to the leaf nodes. Using this tree, the time complexity is reduced to logarithmic level, accelerating feature matching. 4) Composition The data collected by the camera is optimized and the camera posture is corrected to convert the two-dimensional plane points into three-dimensional space, forming three-dimensional space point cloud information. In addition to the point cloud map, the camera posture optimization process is represented in the g2o tool to form a posture graph. Depending on the specific situation, you can also define your own map for description.

6. The intelligent UAV real-time spatial perception and target precision detection system according to claim 1 is characterized in that: Three-dimensional visual RGB space extraction method: The sensor uses a monocular camera to obtain depth information, and the data source includes RGB acquisition image and depth map. Use the depth camera to obtain the color acquisition image and the depth acquisition image, and use the geometric model to convert the 2D plane data into a 3D stereo space. From the pixel coordinates to the acquisition image coordinates: the center points of the coordinate systems are different, and there is only an offset relationship between the two. From the acquisition image coordinate system to the camera coordinate system, the coordinate axes are parallel, and there is only a scaling relationship. The conversion relationship from the pixel coordinate system to the camera coordinate system is expressed as: Where u and v are the offsets between the origins of the coordinate system and the center of the acquisition plane, d x , d y is the scaling ratio between pixel coordinates and the actual imaging plane, and the scaling ratio is d x =z c / f x , d y =z c / f y , f x , f y is the focal length of the camera on the x and y axes, written in matrix form: When the camera moves, the camera coordinate system and the world coordinate system are not parallel, and there is a rotation and translation relationship. In the subsequent visual odometry calculation, the relationship between the previous and next frames is the same as above, and the matrix relationship is given as: Convert the points on the two-dimensional plane to three-dimensional space, and finally obtain a series of point cloud data, which are assigned RGB color attributes to initially obtain a color three-dimensional map.

7. The intelligent UAV real-time spatial perception and target precision detection system according to claim 1 is characterized in that: Implementation of 3D visual RGB space extraction: (1) Front-end visual odometry Initialization is based on the first frame of the acquisition image, and the search for key frames begins. The ORB algorithm is used to extract key points between the acquisition images, and then the BRIEF descriptor is calculated for each key point. Finally, the fast approximate nearest neighbor algorithm is used for fast matching. The ORB corner point extraction algorithm adds scale and rotation descriptions to the FAST corner point extraction algorithm, adding feature information, richer feature descriptions, higher matching accuracy, and more accurate and reliable composition. The BRIEF descriptor uses a binary descriptor and uses randomly selected points for comparison. After matching is completed, the 2D points are projected into 3D space based on the depth acquisition image to obtain the 2D coordinates and corresponding 3D coordinates of a series of points. The camera's position is estimated by solving the PnP problem. The actual calculation result is the rotation and translation matrix between the two frames of acquisition. All data are matched pairwise in sequence and the camera pose is calculated to obtain a complete visual odometry. (2) Back-end nonlinear global optimization Three-dimensional visual RGB space extraction uses a pose graph to represent the pose calculated by the visual odometry, including nodes and edges. Nodes represent the poses of each camera, and edges represent the transformation between camera poses. The pose graph not only intuitively describes the visual odometry but also facilitates understanding of changes in camera poses. Nonlinear global optimization is expressed as graph optimization. The same scene does not appear in many locations, making the pose graph sparse. A sparse BA algorithm is used to solve the pose graph to correct the camera pose.

8. The intelligent unmanned aerial vehicle real-time spatial perception and target precision detection system according to claim 1 is characterized in that: UAV hardware system: includes the following four parts: onboard computer; onboard module assembly; camera and gimbal; M100 quad-rotor UAV, among which the M100 quad-rotor UAV provides a flight platform to realize basic flight functions; the camera and gimbal are image acquisition components in the onboard hardware system; the onboard module assembly is a hardware assembly part located between the camera gimbal, M100 quad-rotor UAV and the onboard computer: 1) the video data of the camera is collected and sent to the onboard computer, 2) the onboard computer controls the gimbal through the onboard module assembly, 3) the onboard computer can realize the flight control of the M100 quad-rotor UAV through the onboard module assembly, 4) the video acquisition image data of the camera can be input into the image transmission system of the M100 quad-rotor UAV through the onboard module assembly; 5. voltage conversion, the onboard module assembly converts the 24V voltage obtained from the M100 UAV battery into 12V to power the gimbal and onboard computer; (1) Onboard computer The onboard computer uses the NVIDIA JETSON TX2 RTS-ASG003 microcomputer, with a total weight of 170g; (2) Airborne module assembly The airborne module assembly is an intermediate execution processing unit in the entire airborne hardware system. Its internal composition and implementation are as follows: 1) The video data output by the high-definition camera is divided into two paths through the HDMI distributor. One path goes to the video collector and is output by the video collector to the airborne computer; the other path is output to the wireless image transmission system of the M100 UAV through the N1 encoder; 2) The USB to UART and PWM modules are used by the airborne computer to control the flight control and gimbal control of the M100; 3) The visual sensor realizes autonomous obstacle avoidance of the M100 UAV; 4) The RC receiver receives the control signal of the ground remote control to control the movement of the gimbal; 5) The wireless data transmission module provides a low-bandwidth data link between the airborne hardware system and the ground system; 6) It supplies power to the airborne computer, gimbal, and various components within the airborne module assembly; The video acquisition part of the airborne module assembly consists of three parts: an HDMI distributor, a video collector, and an N1 encoder. The HDMI distributor splits the video stream from the HD camera into two channels, which then enter the video collector and N1 encoder respectively through the HDMI interface. In the airborne module assembly, the wireless data transmission module and the ground-side wireless data transmission module establish a low-bandwidth wireless data transmission link, providing a data path for bidirectional data transmission between the air and ground terminals. The Guidance establishes a relatively independent visual sensing system equipped with five sets of visual ultrasonic combination sensors, which monitor environmental information in multiple directions in real time and sense obstacles. In conjunction with the UAV flight controller, it can enable the aircraft to avoid possible collisions during high-speed flight. The RC wireless receiver R7008SB receives control signals from the ground remote control, outputs PWM waveforms after internal processing to control the gimbal motion state, and has 16 receiving channels. The USB to UART and PWM module implements two functions: completing USB to UART (TTL level) conversion, connecting the module's UART interface to the M100 drone's UART interface, enabling the onboard computer to control the M100 drone's flight and achieve autonomous flight; the USB to PWM module enables the onboard computer to output PWM control signals through the module, connecting the module's output PWM control signals to the gimbal's heading and pitch control signals, enabling the onboard computer to control the gimbal's motion state; (3) Camera and gimbal A GoPro Hero4 HD camera is used to collect video data, and a MiNi3DPro gimbal is used with the GoPro Hero4 HD camera to control the camera's viewing angle. The gimbal is a three-axis gimbal that can achieve motion control in three directions: pitch, roll, and heading. There are two control methods for the gimbal: one is to use the onboard computer to control the gimbal through the USB to UART and PWM module to output PWM waveforms to achieve control, and the other is to use the ground remote control to control the onboard gimbal receiver RS7008SB to output PWM waveforms to control the gimbal. Set the control signals for the pan and tilt axes of the gimbal. The control signals for both the pan and tilt axes are PWM waveforms with a period of 50 Hz. Position control of the pan and tilt axes is achieved by adjusting the duty cycle of the control signals. A duty cycle of 5.1% corresponds to the minimum position, 7.6% corresponds to the equilibrium position, and 10.1% corresponds to the maximum position. Set the gimbal's mode control signal to achieve control in three modes: lock mode, heading and pitch follow mode, and heading follow mode. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 5% and 6%, the gimbal enters lock mode, in which the heading, pitch, and roll are all locked, and the heading and pitch are controlled by the remote control or onboard computer. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 6% and 9%, the gimbal enters heading and pitch follow mode, in which the roll is locked, the heading rotates smoothly in the direction of the head, and the pitch rotates with the aircraft's pitch angle. When the signal input to the mode control pin is a PWM signal with a period of 50 Hz and a duty cycle between 9% and 100%, the MiNi3D Pro gimbal enters heading follow mode, in which the heading, pitch, and roll are all locked, the heading rotates smoothly in the direction of the head, and the pitch is controlled by the remote control or onboard computer.

9. The intelligent UAV real-time spatial perception and target precision detection system according to claim 1 is characterized in that: UAV hardware system integration: Set up each unit module of the UAV airborne module assembly, design the gimbal power supply system, the 25V voltage output by the UAV battery is converted into 19V, 12V and 5V voltages after passing through the three-way DCDC voltage conversion module of the airborne module assembly. The 19V voltage is used to power the onboard computer; the 12V voltage is used to power the gimbal; the 5V voltage is used to power the RC receiver and the HDMI distributor. The internal power supply system of the UAV platform powers the Guidence visual sensor, N1 encoder, airborne wireless image transmission, and wireless data transmission respectively. The HDMI video capturer and airborne wireless data transmission are powered by the USB port of the onboard computer, and the camera is powered by its own battery.

10. The intelligent UAV real-time spatial perception and target precision detection system according to claim 1, characterized in that: The integration of the drone software system is completed using the ROS system platform, and the ROS message mechanism is used for communication between modules. At the same time, the coupling between the two modules is very loose; 1) Target Detection Process The object detection system uses the Jetson Inference system, which includes classification, detection, and segmentation. The detection modules include image acquisition detection, video detection, and real-time camera detection. The video stream is ultimately decomposed into acquisition frames. The essence of these detections is image acquisition detection. The target detection system converts the video stream into captured image frames and then uses the trained network model to detect the captured images. It ultimately obtains the target category and pixel coordinates of the target frame. The output of the target detection system is then passed to the visual real-time spatial perception modeling system for positioning and navigation. 2) 3D visual RGB space extraction process The 3D visual RGB space extraction system receives RGB acquisition images and depth maps, first matches the RGB acquisition images to obtain key frames, then builds a point cloud image with the depth acquisition image, then corrects the point cloud image through nonlinear global optimization and local detection, and finally receives the target detection results to complete the target positioning in 3D space. 3) Integration of UAV onboard processing methods The onboard processing process of the software system is that after TX2 receives the captured image data from the camera, the target detection module first detects the content of the captured image to obtain the target position coordinates, and then transmits it to the visual real-time spatial perception modeling system in real time. At this time, the visual real-time spatial perception modeling system is composing the key frames and reconstructing the target detection results corresponding to the key frames in three-dimensional space to realize spatial content extraction.