Multi-Modal Deep Learning for UAV Video Aesthetic Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for evaluating the aesthetic quality of UAV videos do not effectively utilize aerial-specific features and lack deep learning approaches, making it difficult to distinguish between professional and amateur footage, and are not robust against changes in lighting and image quality.
Innovation Solution
A multi-modal deep learning method that uses a neural network to analyze UAV videos by extracting features from image, motion, and structure branches, and concatenates these features to evaluate aesthetic quality, incorporating SLAM technology for camera pose and scene reconstruction, and scene type classification to improve accuracy and robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional manual feature design is used for aesthetic evaluation, then the method is simple and easy to implement, but the accuracy of distinguishing professional and amateur videos is poor
Solution Approach 1:
The patent replaces manual feature design (mechanical approach) with deep learning-based automatic feature extraction. The convolutional neural network automatically learns aesthetic features from video frames, eliminating the need for manual feature engineering while significantly improving evaluation accuracy.
Solution Approach 2:
The system enables self-service by allowing the neural network to automatically extract and learn aesthetic features from raw video data without human intervention. The model trains itself on labeled datasets and continuously improves its evaluation capabilities through self-learning.
2Measurement precision
If deep learning methods are applied to video aesthetic evaluation, then the classification accuracy is greatly improved, but the requirement for data sets and computational resources increases
Solution Approach 1:
The patent segments the video evaluation task into multiple components: frame-level aesthetic quality assessment, camera motion analysis, and composition evaluation. This segmentation allows the system to process videos efficiently by analyzing individual frames and their temporal relationships, reducing the computational burden compared to processing entire videos as single units.
3Adaptability or versatility
If single-modal image features are used for evaluation, then the evaluation process is simple, but the ability to capture video-specific aesthetic characteristics is insufficient
Solution Approach 1:
The patent merges multiple modalities including image quality assessment, camera motion analysis, and composition evaluation into a unified multi-modal deep learning framework. This integration allows the system to comprehensively evaluate video aesthetics by combining information from different sources, achieving superior performance compared to single-modal approaches.
4Reliability
If conventional evaluation methods are used, then the computational cost is low, but the robustness against lighting changes and image quality variations is poor
Solution Approach 1:
The patent employs data augmentation techniques that transform training data through various parameter changes including lighting adjustments, color adjustments, and geometric transformations. This exposes the model to diverse conditions during training, enabling it to maintain robust performance across varying lighting and image quality conditions in deployment.
Data Source
AI summary
The present disclosure provides a UAV video aesthetic quality evaluation method based on multi-modal deep learning, which establishes a UAV video aesthetic evaluation data set, analyzes the UAV video through a multi-modal neural network, extracts high-dimensional features, and concatenates the extracted features, thereby achieving aesthetic quality evaluation of the UAV video. There are four steps, step one to: establish a UAV video aesthetic evaluation data set, which is divided into positive samples and negative samples according to the video shooting quality; step two to: use SLAM technology to restore the UAV's flight trajectory and to reconstruct a sparse 3D structure of the scene; step three to: through a multi-modal neural network, extract features of the input UAV video on the image branch, motion branch, and structure branch respectively; and step four to: concatenate the features on multiple branches to obtain the final video aesthetic label and video scene type.


