Autonomous aerial display control system using biometric-driven activation
Patent Information
- Authority / Receiving Office
- IN · IN
- Patent Type
- Patents
- Current Assignee / Owner
- VELLORE INSITUTE OF TECH
- Filing Date
- 2025-12-24
- Publication Date
- 2026-07-16
AI Technical Summary
Existing UAV-based display systems lack autonomous, context-aware actuation, rely on manual or remote operation, and lack dual-factor biometric authentication, leading to operational delays, increased risk of unintended activation, and imprecise deployment.
An autonomous aerial display control system integrating real-time facial detection, deep learning-based facial recognition, and offline automatic speech recognition, with a closed-loop encoder-based actuation module for precise banner deployment, ensuring secure and reliable operation without external connectivity.
Enables secure, precise, and reliable deployment of display elements on UAVs through dual-factor biometric authentication, enhancing operational autonomy and reducing human intervention, even under dynamic conditions.
Abstract
Description
FIELD OF INVENTIONThe present invention relates to the domain of unmanned aerial systems (UAS) and, morespecifically, to an autonomous aerial display control system that integrates biometric-basedactivation mechanisms and feedback-regulated actuation for controlled deployment of displayelements. The invention lies at the intersection of human-drone interaction, computer vision-based facial recognition, automatic speech recognition, and precision electromechanical control,and is directed towards applications in aerial advertising, event management, and automated visualdisplay operations.BACKGROUND OF THE INVENTIONUnmanned Aerial Vehicles (UAVs) are increasingly employed in surveillance, logistics, publiccommunication, and aerial advertising. Existing UAV-based display or banner systems are,however, still predominantly dependent on manual operation through handheld controllers ormobile interfaces. This dependency introduces operational delays, increases pilot workload, andlimits the scalability of drone-based visual communication systems, particularly in dynamic eventenvironments where rapid, precise, and secure actuation is required. Conventional advertisingdrones lack the capability to autonomously identify authorized personnel or initiate deploymentactions based on contextual cues, resulting in reduced efficiency and increased risk of unintendedactivation.Several prior arts disclose components of this technology space, but none provide a unified,autonomous, biometric-triggered deployment mechanism. For example, US10930182B2 describesa modular lightweight banner display system for UAVs, focusing primarily on mechanicalscaffolding and rigging configurations to support banner towing. While effective for physicalmounting, this system does not incorporate autonomous control, biometric verification, or adaptiveactuation mechanisms.Similarly, US20210034843A1 discloses methods for improving facial recognition in dronecaptured imagery by adapting the drone's position based on real-time image quality assessments.This teaches improved face capture but does not address voice-based activation, mechanicaldeployment, or closed-loop control of any physical payload.Further, US10538329B2 highlights the use of biometric or sensor-based verification on drones forsecurity and package monitoring. Although this reference demonstrates recognition-drivendecision-making, it does not provide any mechanism for controlled physical actuation of a displaysystem, nor does it integrate dual-factor authentication or precision encoder-based motion control.Across the existing solutions, several critical gaps remain. Current UAV advertising and displaysystems lack autonomous, context-aware actuation, as they cannot independently initiate banneror display deployment based on user identity or voice-based triggers. Moreover, prior technologiesdo not incorporate dual-factor biometric authentication, since facial recognition or voicerecognition may exist individually, but none of the earlier approaches employ a sequentialcombination of face verification followed by voice-command activation to ensure secure,intentional operation. Existing systems also suffer from the absence of precision-controlleddeployment mechanisms, relying on simple mechanical rigs without closed-loop encoderfeedback, resulting in inconsistent banner unrolling, risk of overextension, and poor adaptabilityto environmental variations. Additionally, most UAV-based display systems remain dependent onmanual or remote operation, requiring human pilots or external communication links, which limitsautonomy, increases intervention requirements, and elevates the risk of human error.The present invention overcomes these limitations by integrating real-time facial detection, deeplearning-based facial recognition, and offline automatic speech recognition into a unified controlpipeline executed entirely on the UAV. Following biometric verification, the system triggers aclosed-loop encoder-based actuation module that governs banner deployment with high precision,ensuring accurate roll-out based on banner length and rod circumference. This combinationeliminates the need for manual intervention, enhances operational security through dual-factorauthentication, and enables reliable deployment even under dynamic flight or environmentalconditions.By merging perceptual intelligence, biometric security, and precision electromechanical control,the invention fills the long-standing gaps in existing UAV advertising systems and introduces arobust autonomous platform for aerial display operations.SUMMARY OF THE INVENTIONThe following summary is provided to facilitate a clear understanding of the new features in thedisclosed embodiment and it is not intended to be a full, detailed description. A detailed descriptionof all the aspects of the disclosed invention can be understood by reviewing the full specification,the drawing and the claims and the abstract, as a whole.In one aspect, the present invention provides an autonomous aerial display control systemconfigured to deploy a visual display element from an unmanned aerial vehicle (UAV) based onbiometric activation. The system includes a facial detection module, a facial recognition module,an automatic speech recognition module, and an actuation module that operates through feedbackregulated motor control to manage deployment of a banner or equivalent display element.In another aspect, the invention provides a method for secure and controlled activation of an aerialdisplay mechanism, the method comprising capturing image data of a user, performing real-timefacial detection and recognition to verify identity, capturing an audio command from the verifieduser, processing the audio input to detect a predefined activation phrase, and initiating deploymentof a display element using a closed-loop encoder-based control system.In yet another aspect, the invention provides a closed-loop actuation system for UAV-mounteddisplay mechanisms, wherein an encoder continuously monitors angular position and rotationalspeed of a motor to compute and regulate the required deployment length of the display elementbased on its predefined physical characteristics, ensuring smooth, precise, and repeatable motion.In a further aspect, the invention provides an integrated on-board processing framework thatenables all biometric verification and actuation decisions to be performed directly on the UAVwithout dependence on remote servers or external connectivity, thereby enhancing operationalautonomy, reliability, and deployment accuracy.BRIEF DESCRIPTION OF THE DRAWINGSThe manner in which the present invention is formulated is given a more particular descriptionbelow, briefly summarized above, may be had by reference to the components, some of which isillustrated in the appended drawing It is to be noted; however, that the appended drawing illustratesonly typical embodiments of this invention and are therefore should not be considered limiting ofits scope, for the system may admit to other equally effective embodiments.Throughout the drawings, the same drawing reference numerals will be understood to refer to thesame elements and features.The features and advantages of the present invention will become more apparent from thefollowing detailed description a long with the accompanying figures, which forms a part of thisapplication and in which:Figure 1 illustrates the overall workflow of the autonomous aerial display control system, depictingthe sequence from image acquisition, biometric verification, voice-command processing, toclosed-loop banner deployment.Figure 2 illustrates a drone equipped with the facial detection, facial recognition, and voicerecognition modules that collectively enable biometric-driven activation of the displaymechanism.Figure 3 illustrates the process of guest face detection and recognition, showing how the systemidentifies an authorized individual based on captured facial features.Figure 4 illustrates the voice-recognition stage, representing the detection of a predefined voiceactivation command that triggers the deployment of the display element.Figure 5 illustrates the distribution of cosine similarity scores generated during facialauthentication processing, showing how genuine facial-embedding pairs and impostor embeddingpairs are separated to enable reliable identity verification.Figure 6 illustrates the relationship between the false-acceptance rate and false-rejection rateacross different similarity thresholds, highlighting the threshold region where balancedauthentication performance is achieved.Figure 7 illustrates the receiver operating characteristic (ROC) curve of the facial-authenticationmodule, demonstrating the discriminative capability of the model across varying decisionthresholds.REFERENCE NUMERALS100 - Workflow Diagram101 - Start node102 - System power-on and camera initialization module103 - Facial detection module (YOLOv8n-face)104 - Facial detection decision node105 - Facial recognition module (ResNet-34 embedding generator)106 - Facial verification completion node107 - Voice-recognition system initialization module108 - Voice-command listening module109 - Voice-command detection decision node110 - Encoder-based actuation initialization module for banner deployment111 - Target-rotation decision node (encoder feedback)112 - Banner-movement termination module113 - End node200 - Drone with Facial and Voice Detection System201 - On-board processing unit (Jetson Orin)202 - Voice-detection and audio-processing subsystem203 - Facial-recognition imaging subsystem (Intel RealSense camera)300 - Guest Face Detection and Recognition400 - Voice Recognition500 - Distribution of cosine similarity scores600 - FAR-FRR threshold behavior700 - ROC curve for facial-authentication performanceDETAILED DESCRIPTION OF THE INVENTIONThe present invention will now be described in detail with reference to the accompanyingdrawings, wherein like reference numerals refer to corresponding components throughout thefigures. The embodiments described herein are illustrative and are not intended to limit the scopeof the invention.Referring to Figure 1 (100), the operation of the autonomous aerial display control system beginsat the start node (101). The UAV initiates power-up and activates its onboard imaging hardwarethrough the system power-on and camera initialization module (102). Once the visual subsystembecomes active, the captured video stream is processed through the facial detection module (103),which employs a lightweight deep-learning architecture such as YOLOv8n-face for real-timeidentification of facial regions within the incoming frames.The detection output is evaluated at the facial detection decision node (104). In the event that novalid face is detected, the system continues to loop back for continuous monitoring. When a faceis successfully detected, the corresponding image frame is supplied to the facial recognitionmodule (105), which utilizes a convolutional neural network-based embedding generator, such asa ResNet-34 encoder, to compute stable biometric feature vectors. These embeddings arecompared against a pre-stored list of authorized profiles to verify identity. Upon a positive match,the system proceeds to the facial verification completion node (106).Following successful facial verification, control shifts to the voice-recognition systeminitialization module (107), wherein the UAV activates an offline automatic speech-recognitionengine configured to capture and interpret user speech. The system continues to actively listen fora predefined trigger phrase via the voice-command listening module (108). Incoming audio isanalyzed at the voice-command detection decision node (109), and in the case of mismatch ornoise, the listening loop persists until the correct activation command is detected.Once the required voice command is recognized, the system transitions to the actuation phasethrough the encoder-based actuation initialization module (110). Here, the system computes therequired number of rotational cycles needed to deploy the banner or visual display elementmounted on the UAV. These computations are based on stored banner parameters, includingbanner length and roller-rod circumference.During deployment, the system continuously monitors the encoder feedback at the target-rotationdecision node (111). Encoder signals provide real-time data on angular displacement, rotationalspeed, and motion direction, enabling fine-grained closed-loop control. When the target rotationthreshold corresponding to full or partial banner deployment is reached, the system executes thebanner-movement termination module (112), thereby halting the motor and stabilizing thedeployed display. The workflow concludes at the end node (113).Referring to Figure 2 (200), the UAV platform incorporates an embedded high-performanceprocessing unit, illustrated as the on-board processing unit (Jetson Orin) (201). This processorexecutes all computer-vision inference, speech recognition, and actuation logic locally on thedrone without reliance on external networks or cloud resources.A directional or array-based microphone system forms part of the voice-detection and audioprocessing subsystem (202), enabling the capture and decoding of speech commands throughoffline ASR models. Likewise, the UAV integrates a depth-enabled imaging module, shown as thefacial-recognition imaging subsystem (Intel RealSense camera) (203), which supplies RGB-Dvideo input for face detection, recognition, and environmental assessment under diverse lightingand distance conditions.The collective functioning of the processing unit, voice-recognition subsystem, and facialrecognition subsystem enables reliable, secure, and real-time biometric activation of the displaymechanism.Referring to Figure 3 (300), the system captures the front-facing region of a guest or authorizedoperator. The depth-enhanced imaging frames undergo noise reduction, normalization, andcontrast enhancement prior to being submitted to the detection module. The module isolates faciallandmarks and passes the cropped regions to the recognition model to compute embeddings. Theseembeddings are matched against stored profiles to validate credentials. Only upon achieving asimilarity score higher than a predetermined threshold is the user permitted to proceed to the voiceactivation stage.Referring to Figure 4 (400), once facial authentication is complete, the UAV activates itsembedded ASR engine. The microphone subsystem continuously captures audio from theenvironment. The ASR engine processes the input using acoustic and language models tuned foroffline operation, ensuring reliability even in low-connectivity environments. When the predefinedtrigger phrase is uttered, the system registers a successful command event and initiates the bannerdeployment sequence.Referring now to Figure 5 (500), the system further includes a biometric-performance evaluationcomponent used during calibration of the facial-recognition module. The figure illustrates thedistribution of cosine similarity scores generated from facial-embedding pairs obtained during theenrollment and testing phases. Genuine user pairs form a distinct cluster separated from impostorpairs, enabling the facial-recognition module to determine an effective similarity threshold foridentity verification. This separation guides the configuration of the facial-verification decisionnode such that only embedding pairs exhibiting sufficient similarity confidence progress to thesubsequent voice-activation stage. The distribution supports the robustness of the recognitionmodel and ensures stable identity-verification performance under varying lighting and poseconditions.Referring to Figure 6 (600), the system incorporates threshold-tuning logic informed by falseacceptance and false-rejection rate behaviour. The figure illustrates how the false-acceptance rate(FAR) and false-rejection rate (FRR) vary as the similarity threshold is adjusted. By analysing thisbehaviour, the system identifies an operating threshold that achieves balanced authenticationperformance. This threshold is stored within the facial-recognition module and influences theverification node responsible for gating access to the voice-processing sequence. The use of FAR-FRR behaviour ensures that the UAV maintains a secure activation pipeline by minimisingunintended acceptance of impostor faces while maintaining convenient access for authorised users.Referring to Figure 7 (700), the system further evaluates the discriminative capability of the facialauthentication pipeline through a receiver operating characteristic (ROC) analysis. The ROC curveillustrates the trade-off between true-positive and false-positive rates across all possible similaritythresholds. A curve that remains close to the upper-left region of the plot indicates high separabilitybetween authorised and unauthorised user embeddings, confirming the effectiveness of the deeplearning-based encoding model. This ROC-based assessment is utilised during systemconfiguration to validate that the facial-recognition module meets reliability requirementsnecessary for UAV-mounted actuation, where erroneous activation may lead to unsafe deploymentevents. The ROC evaluation therefore forms an integral part of the system-initialisation andcalibration process.These performance-characterisation drawings collectively demonstrate the biometric reliabilitythat underpins the dual-factor authentication mechanism of the system. The cosine-similaritydistribution (500) provides embedding-space separability, the FAR-FRR analysis (600)establishes threshold selection for balanced operation, and the ROC curve (700) validates theoverall discriminative capability of the facial-authentication module. Together, these elementsstrengthen the security and operational robustness of the autonomous aerial display control systemand support its integration into a UAV-based deployment pipeline that requires high-confidenceidentity verification prior to actuation.Experimental Validation ResultsThe autonomous aerial display control system was experimentally validated in a controlled indoorenvironment to assess the accuracy of biometric activation and the precision of encoder-regulatedbanner deployment. The UAV was equipped with the Intel RealSense camera, Jetson-basedonboard processor, an offline speech-recognition subsystem, and a motor-encoder assemblyconsistent with the configuration described in the Detailed Description. The results demonstratethat the system performs reliably and consistently across repeated trials.1. Facial Detection and Recognition PerformanceThe system was evaluated using ten individuals, including authorized and unauthorized subjects,under varying illumination and distances ranging from 1.5 m to 4 m. The YOLOv8n-face detectionmodule achieved a mean detection confidence of 0.89, with an average detection time of 23 msper frame, suitable for real-time onboard inference. The ResNet-34 embedding-based facialrecognition module exhibited a recognition accuracy of 96% for authorized profiles, with a falseacceptance rate below 2% when using a cosine-similarity threshold of 0.52. The systemconsistently rejected unauthorized faces, demonstrating robust identity discrimination.2. Voice-Command Activation ReliabilityFollowing successful facial verification, the system activated the offline PocketSphinx speechrecognition engine. Across 50 trials using a predefined activation phrase, the ASR moduleachieved a command-recognition accuracy of 94%, with a mean response latency of 210 ms.Background noise from UAV propellers did not significantly degrade recognition accuracy, owingto the noise-tolerant acoustic models used in PocketSphinx. No activation was triggered withoutprior face verification, confirming reliable sequential dual-factor authentication.3. Encoder-Based Banner Deployment PrecisionThe banner deployment mechanism was tested with banners of three different lengths: 0.5 m, 1.0m, and 1.5 m. The encoder was pre-configured with the roller-rod circumference to compute therequired rotational displacement. Across 30 deployment attempts, the system achieved a meanpositional error of 1.8%, with smooth motor operation and no observed over-extension. Real-timeencoder feedback allowed the system to dynamically compensate for minor variations in spooltension and airflow, maintaining consistent motion quality. Deployment was successfully haltedwithin ±2 rotations of the target threshold, confirming effective closed-loop control.4. Integrated System Response TimeThe complete biometric-to-deployment pipeline was measured end-to-end. From initial facialdetection to final banner deployment, the system demonstrated an average total activation time of1.46 seconds, indicating that the onboard processor successfully handled the combined inferenceand actuation workload without latency bottlenecks. The integration of facial recognition, voiceactivation, and encoder-feedback allowed the system to operate autonomously without externalnetwork dependence.5. Biometric-Performance Analysis Using Embedding-Space Evaluation5.1 Cosine Similarity DistributionReferring to Figure 500, the system generated cosine-similarity scores from facial embeddingscollected under varied lighting and pose conditions. Genuine embedding pairs and impostorembedding pairs formed distinct score distributions, demonstrating clear separability within theembedding space. This distribution supports the system's ability to select an effective similaritythreshold for reliable user verification during flight.5.2 FAR-FRR Threshold BehaviourReferring to Figure 600, the false-acceptance rate and false-rejection rate were analysed across arange of similarity thresholds. The opposing trends of these rates indicate the trade-off betweenusability and security inherent to biometric systems. The system identifies a balanced operatingthreshold based on this behaviour, ensuring that impostor activations are minimised withouthindering access for authorised users.5.3 ROC-Based Authentication AssessmentReferring to Figure 700, the receiver operating characteristic curve was generated to evaluate thediscriminative capability of the facial-authentication model. The curve demonstrates strongseparation between authorised and unauthorised comparisons across all thresholds. This ROCbehaviour verifies that the recognition engine provides the reliability required for UAV-basedactuation, where false activation poses operational risks.5.4 Evaluation MethodologyAll biometric-performance plots were generated from a dataset comprising multiple facial imagesof enrolled users captured under diverse environmental conditions. Facial regions were firstdetected using a deep-learning-based face detector, and embeddings were generated using aconvolutional neural-network encoder. Genuine and impostor similarity scores were computed bycomparing embedding pairs from the same user and across different users, respectively. Thesesimilarity distributions formed the basis for the threshold analysis and ROC evaluationincorporated into the system's calibration and authentication logic.6. Overall Validation OutcomeThe combined results confirm that the proposed system:- Performs reliable dual-factor biometric verification suitable for autonomous UAVactivation,- Maintains robust authentication performance across environmental variations,- Deploys banners with high positional precision using encoder feedback, and- Operates end-to-end in real time using onboard processing without external computation.These findings validate the technical feasibility, reliability, and operational safety of theautonomous aerial display control system for aerial advertising and automated displayapplications.This detailed description provides a comprehensive overview of the autonomous aerial displaycontrol system and its key biometric and electromechanical features. The invention delivers apractical and highly effective solution to the challenges associated with conventional drone-basedadvertising, such as manual operation, unreliable activation, and imprecise banner deployment. Byintegrating dual-factor biometric verification with closed-loop encoder-regulated actuation, thesystem enhances operational efficiency, deployment accuracy, and overall reliability in aerialdisplay applications.
Claims
1. A system (200) for autonomous deployment of an aerial display element (500) from an unmanned aerial vehicle (UAV), comprising: a UAV carrying an on-board processing unit (201), a facial-recognition imaging subsystem (203), a voice-detection and audio-processing subsystem (202), a facial detection module (103), a facial recognition module (105), a voice-recognition module (107,108,109), and an actuation mechanism including a motor and an encoder forming part of a bannerdeployment system (500), wherein the system is configured to detect a face, recognize an identity, receive a voice command, and deploy the display element; characterised by the on-board processing unit (201) being configured to: perform real-time facial detection using the imaging subsystem (203) through the facial detection module (103), and verify an authorized user via the facial recognition module (105) corresponding to the facial verification completion node (106); activate the voice-recognition module (107,108,109) only upon successful facial verification, and detect a predefined activation command through the voicecommand listening module (108) and voice-command detection node (109); and initiate deployment of the aerial display element (500) through an encoderregulated actuation module (110), wherein an encoder provides continuous rotational feedback to a target-rotation decision node (111) to regulate motor operation and automatically terminate deployment via the banner-movement termination module (112) upon reaching a predefined rotation threshold.
2. The system as claimed in claim 1, wherein the imaging subsystem (203) comprises a depthenabled RGB-D camera configured to enhance facial detection accuracy under varying illumination and distance conditions.
3. The system as claimed in claim 1, wherein the facial recognition module (105) generates facial embeddings and performs identity verification using similarity comparison against pre-stored profiles.
4. The system as claimed in claim 1, wherein the voice-recognition module (107-109) comprises an offline automatic speech-recognition engine configured to operate without an external network connection.
5. The system as claimed in claim 1, wherein the encoder-based actuation module (110) computes rotational displacement based on banner length and roller-rod circumference for precision deployment of the display element (500).
6. A method (100) for autonomous deployment of an aerial display element (500) from an unmanned aerial vehicle (200), comprising the steps of: capturing image data using an imaging subsystem (203), detecting a face using a facial detection module (103), recognizing an identity using a facial recognition module (105), capturing audio input using a voice-detection subsystem (202), detecting a voice activation command, and deploying a display element (500), wherein the method enables biometric-driven activation of a banner-deployment mechanism; characterised by verifying an authorized user through facial recognition at the verification node (106) prior to activating any voice-processing step; initiating a voice-recognition module (107) only after successful facial verification, detecting a predefined activation command via the listening module (108) and command-detection node (109); and deploying the display element (500) using an encoder-regulated actuation step (110), wherein encoder feedback is continuously evaluated at the targetrotation decision node (111) to control motor operation and terminate deployment through the termination module (112) upon reaching a predefined rotation count.
7. The method as claimed in claim 6, wherein facial detection is performed using a lightweight convolutional neural network optimized for real-time inference on embedded hardware (201).
8. The method as claimed in claim 6, wherein the predefined voice activation command is processed using an offline automatic speech-recognition engine forming part of the subsystem (202).
9. The method as claimed in claim 6, wherein encoder feedback includes angular position, rotational velocity, and rotational direction for maintaining smooth and stable deployment.
10. The method as claimed in claim 6, wherein deployment is halted automatically when the encoder indicates completion of a required number of rotations stored as a predefined banner-deployment parameter in the actuation module (110).