Autonomous unmanned aerial vehicle with deep learning-based audio and image processing technologies
The UAV integrates deep learning-based audio and image processing to enhance autonomy and effectiveness by enabling simultaneous data processing for real-time object detection and classification, addressing limitations in existing systems.
Patent Information
- Application Number
- PCT/TR2025/050947
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2026-04-16
AI Technical Summary
Existing UAV systems lack the capability to simultaneously process audio and image data for real-time object detection and classification, limiting their autonomy and effectiveness in adverse weather conditions or noisy environments.
An unmanned aerial vehicle equipped with deep learning-based convolutional neural networks for image and sound processing, enabling simultaneous audio and image data processing, allowing real-time object detection and classification, and autonomous decision-making.
Enhances UAV autonomy and operational effectiveness by enabling real-time object detection and classification in various conditions, including adverse weather and noisy environments, through integrated audio and image processing capabilities.
Smart Images

Figure TR2025050947_16042026_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] AUTONOMOUS UNMANNED AERIAL VEHICLE WITH DEEP LEARNING-BASED AUDIO AND IMAGE PROCESSING TECHNOLOGIES
[0003] Technical Field of the Invention
[0004] The invention relates to an Unmanned Aerial Vehicle (UAV) that can perform real-time object detection and classification operations by combining deep learning based environmental sound classification technology and computer vision techniques, and can make autonomous decisions thanks to these features.
[0005] State of the Art
[0006] Deep learning-based audio and image processing technologies significantly enhance the capabilities of autonomous UAVs and have great potential for industrial applications. In the state of the art, in object detection and classification, image and sound processing technology is used separately for different purposes in different fields. However, there is no system with the technology to process audio and image together. An embedded artificial intelligence module is required to process both data to enable autonomous decision-making and real-time detection as well as object detection and classification from recorded images and sounds. The ability of a platform to use both audio and image processing technologies simultaneously through embedded systems greatly increases its object detection and classification capacity and gives it a significant advantage. For these reasons, it is important to develop a system capable of using both audio and image processing technologies simultaneously, which can be applied to UAVs or other platforms.
[0007] The application numbered TR2024 / 000173 in the state of the art enables the creation of a more advanced object recognition system by analyzing the images taken from the cameras of unmanned vehicles with the help of artificial intelligence-supported algorithms and introducing the results obtained to the development board.
[0008] When the systems in the state of the art are examined, image and sound processing techniques are used separately in various platform applications, for example, in unmanned aerial vehicles. There is not yet a platform with the ability to simultaneously use audio and image processing capabilities and move autonomously. For this reason, it is of great importance to develop a system that enables sound-based object recognition and classification by detecting environmental sounds and performs object recognition and classification by visually detecting objects in the environment during flight with image processing.
[0009] Summary and Objects of the Invention
[0010] The invention relates to an Unmanned Aerial Vehicle (UAV) that can perform real-time object detection and classification operations by combining deep learning based environmental sound classification technology and computer vision techniques, and can make autonomous decisions thanks to these features.
[0011] An object of the invention is to provide a platform with the ability to simultaneously use audio and image processing capabilities and move autonomously.
[0012] Another object of the invention is to enable performing detection and classification operations in real time or close to real time.
[0013] Another object of the invention is to enable object detection and classification tasks to be performed with audio data in conditions where image data cannot be obtained properly (foggy weather, darkness at night, too much sunlight affecting the camera).
[0014] Another object of the invention is to enable tasks to be performed with image data in extremely noisy environments where audio data cannot be obtained properly.
[0015] Description of the Drawings
[0016] Fig. 1. Drawing showing a schematic view of the unmanned aerial vehicle of the invention.
[0017] Description of the References in the Drawings
[0018] 100. Unmanned aerial vehicle 110. Control board
[0019] 111. Camera
[0020] 112. Microphone
[0021] 113. Processor
[0022] 114. Decision-making module
[0023] 120. Flight control board
[0024] Detailed Description of the Invention
[0025] The invention relates to an unmanned aerial vehicle (100) (UAV) that can perform realtime object detection and classification operations by combining deep learning based environmental sound classification technology and computer vision techniques, and can make autonomous decisions thanks to these features.
[0026] The autonomous decision-making UAV (100) performs image processing in an embedded manner. This requires a control board (110) to be placed on the UAV, so that the video and image information obtained from the camera (111) is fed directly into this control board (110) and image processing is performed on the processor (113) located in the control board (110).
[0027] The deep learning-based convolutional neural network architecture implemented in the processor enables the UAV to transform object information in its immediate vicinity into abstract information that can be interpreted by machines without human intervention. Based on the information available, machines are able to perform real-time decision making. Integrating image processing technology including a deep learning-based convolutional neural network into a UAV's flight control system significantly improves the UAV's autonomous decision-making capability and flight safety. Compared to other machine learning methods, the main advantage of convolutional neural network algorithms is their ability to detect and classify objects in real time with computationally faster and superior performance. The convolutional neural network algorithm used in this study is based on a combination of deep learning algorithms and advanced TPU technology.
[0028] The UAV is equipped with at least one microphone (112) to receive sounds from the environment and at least one camera (111) to receive images from the environment. In this way, object detection and classification tasks can be performed with audio data in conditions where image data cannot be obtained properly (foggy weather, darkness at night, too much sunlight affecting the camera (111)). In extremely noisy environments where audio data cannot be received properly, the task can be performed with image data. This will enable cross-validation with both technologies. Deep learning will be able to detect each object introduced as sound and / or image. In a preferred embodiment of the invention, in military applications, object detection is performed on helicopters, tanks, aircraft, and other unmanned aerial vehicles, and in civilian applications, for the detection and tracking of endangered birds (by sound and image) as well as animals distinguishable by their sounds, and, based on both images and human voices, in natural disasters such as fires, earthquakes, and floods.
[0029] Since the processor in the unmanned aerial vehicle has image and sound processing capabilities, it can process the image and sound data captured by the camera (111 ) and microphone (112) without the need for a central computing engine, detect the object based on this data and make autonomous decisions based on the information previously transmitted to it. There is a processor that performs real-time object detection and classification from the image and audio data provided by the UAV camera (111) and microphone (112). According to the data obtained as a result of these processes, there is a decision-making module that provides the UAV with autonomous decision-making and movement capability. The decision-making module (114) receives task information to be executed through an interface from the user and executes the tasks on the objects detected by the processor (113). The decisionmaking module (114) makes the decision to perform on the objects detected by the processor (113) the tasks of transferring the location information of the object defined by the user, tracking the relevant object, destroying the relevant object, and transferring all information about the relevant object to the ground control system. The decisionmaking module (114) allows users to perform operations of adding new tasks and canceling tasks through the interface. In the decision-making module (114), both condition and decision commands can be updated, e.g. execution of the Decision 1 command in case of Condition 1 , execution of the Decision 2 command in case of Condition 2. For example, in Condition 1 , the UAV is asked to detect an object and only image data can be received. If there is a detection with 90% or more accuracy with image data, it can be said that decision 1 should take place. In the condition, both what data will be (audio or image) and the accuracy value in the detection are dynamic. Decision 1 is also dynamic according to this condition. Condition 2: if both audio and image data can be obtained and the accuracy values in detecting both types of data using artificial intelligence are as follows, Decision 2 should be executed. The control module in the UAV also has 4G internet with an integrated modem and these conditions and control structures can be changed remotely. Instant task descriptions can be provided.
[0030] Platform noises that would negatively affect the prediction accuracy of the sound classification model (e.g. sounds produced by UAV engines and propellers) were modeled separately according to platform movements and designed in a variable structure with digital filtering analysis. Communication was provided between the control board (110) with audio and image processing technology and the flight control board (120) of the unmanned aerial vehicle (100) via the MAVLink communication protocol. The control board (110) manages the flight control board (120) in parallel with the audio and image processing tasks. For moving platforms, the word autonomous is more commonly used for planned flight, which is uploaded to the motion control board via a remote control station. The present invention is capable of executing different operations in different scenarios according to the task execution.
[0031] This model is customized by being training from scratch separately for two different tasks. The deep learning model, trained in two separate planes with the transfer learning method, was then converted into "TensorFlow Lite" format to run the Edge TPU co-processor. The aim is the detection of the military helicopter as an object by the UAV platform with image and sound data.
[0032] The innovative aspect of this invention is to introduce an Unmanned Aerial Vehicle (UAV) platform that can perform real-time object detection and classification operations by combining deep learning-based computer vision and techniques and deep learningbased environmental sound classification technology. This platform has advanced capabilities in the processing and interpretation of environmental sounds and images, allowing it to perform defined tasks automatically. In this context, the developed system provides an important contribution to increase the functionality of UAVs in both defense and civilian applications. The use of deep learning-based audio and image processing technologies can offer significant benefits for various industries by enabling UAVs to operate more effectively, efficiently, and autonomously.
Claims
CLAIMS1 . Autonomous unmanned aerial vehicle (100), characterized in that it comprises:- at least one microphone (112) for receiving sounds around the unmanned aerial vehicle (100); at least one camera (111) for capturing images around the unmanned aerial vehicle (100); at least one processor (113) that performs object detection and classification by performing image processing on images from the camera (111 ) using a deep learning-based convolutional neural network architecture, object detection and classification on sounds from the microphone (112) using deep learning-based audio processing, and cross- validation by comparing the object information obtained by image processing and audio processing; at least one decision-making module (114) that receives task information to be executed through an interface from the user and executes the tasks on the objects detected by the processor (113); at least one control board (110) for transmitting the decisions made by the included decision-making module (114) to the flight control board (120),- at least one flight control board (120) that enables the command of the unmanned aerial vehicle (100) in line with the tasks given by the decisionmaking module (114).
2. Unmanned aerial vehicle (100) according to claim 1 , characterized in that it comprises a processor (113) that performs digital filtering by modeling the platform noises separately according to the platform movements, which would negatively affect the prediction accuracy of the sound classification model.
3. Unmanned aerial vehicle (100) according to claim 1 , characterized in that it comprises a processor (113) that detects each object that is presented visually and / or audibly to the deep learning-based convolutional neural network.
4. Unmanned aerial vehicle (100) according to claim 1 , characterized in that it comprises at least one decision-making module (114) for deciding to perform ondetected objects user-defined tasks of taking photographs, recording audio and image, transmitting object location, tracking the object, and destroying the object.
5. Unmanned aerial vehicle (100) according to claim 1 or claim 4, characterized in that it comprises a decision-making module (114) that enables users to add new tasks and cancel tasks via an interface.
Citation Information
Patent Citations
UAV detection
EP3371619B1
Object detection and analysis via unmanned aerial vehicle
US20170053169A1
Neural processing unit
US20230168921A1