Multi-Modal Deep Learning for UAV Video Aesthetic Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for evaluating the aesthetic quality of UAV videos do not effectively utilize aerial-specific features and lack deep learning approaches, making it difficult to distinguish between professional and amateur footage, and are not robust against changes in lighting and image quality.

Innovation Solution

A multi-modal deep learning method that uses a neural network to analyze UAV videos by extracting features from image, motion, and structure branches, and concatenates these features to evaluate aesthetic quality, incorporating SLAM technology for camera pose and scene reconstruction, and scene type classification to improve accuracy and robustness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional manual feature design is used for aesthetic evaluation, then the method is simple and easy to implement, but the accuracy of distinguishing professional and amateur videos is poor

Engineering Contradiction:
Improveaesthetic evaluation accuracyVSAvoidevaluation system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces manual feature design (mechanical approach) with deep learning-based automatic feature extraction. The convolutional neural network automatically learns aesthetic features from video frames, eliminating the need for manual feature engineering while significantly improving evaluation accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the neural network to automatically extract and learn aesthetic features from raw video data without human intervention. The model trains itself on labeled datasets and continuously improves its evaluation capabilities through self-learning.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If deep learning methods are applied to video aesthetic evaluation, then the classification accuracy is greatly improved, but the requirement for data sets and computational resources increases

Engineering Contradiction:
Improveclassification accuracyVSAvoiddata set requirement
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the video evaluation task into multiple components: frame-level aesthetic quality assessment, camera motion analysis, and composition evaluation. This segmentation allows the system to process videos efficiently by analyzing individual frames and their temporal relationships, reducing the computational burden compared to processing entire videos as single units.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If single-modal image features are used for evaluation, then the evaluation process is simple, but the ability to capture video-specific aesthetic characteristics is insufficient

Engineering Contradiction:
Improvevideo aesthetic evaluation capabilityVSAvoidmulti-modal network structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple modalities including image quality assessment, camera motion analysis, and composition evaluation into a unified multi-modal deep learning framework. This integration allows the system to comprehensively evaluate video aesthetics by combining information from different sources, achieving superior performance compared to single-modal approaches.

Inventive Principle:
Principle #5Merging (Combining)

4Reliability

If conventional evaluation methods are used, then the computational cost is low, but the robustness against lighting changes and image quality variations is poor

Engineering Contradiction:
Improverobustness to lighting and quality variationsVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent employs data augmentation techniques that transform training data through various parameter changes including lighting adjustments, color adjustments, and geometric transformations. This exposes the model to diverse conditions during training, enabling it to maintain robust performance across varying lighting and image quality conditions in deployment.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11568637B2UAV video aesthetic quality evaluation method based on multi-modal deep learning
Publication Date: 2023.01.31 BEIHANG UNIV
  • US11568637B2 patent drawing
  • US11568637B2 patent drawing
  • US11568637B2 patent drawing

AI summary

The present disclosure provides a UAV video aesthetic quality evaluation method based on multi-modal deep learning, which establishes a UAV video aesthetic evaluation data set, analyzes the UAV video through a multi-modal neural network, extracts high-dimensional features, and concatenates the extracted features, thereby achieving aesthetic quality evaluation of the UAV video. There are four steps, step one to: establish a UAV video aesthetic evaluation data set, which is divided into positive samples and negative samples according to the video shooting quality; step two to: use SLAM technology to restore the UAV's flight trajectory and to reconstruct a sparse 3D structure of the scene; step three to: through a multi-modal neural network, extract features of the input UAV video on the image branch, motion branch, and structure branch respectively; and step four to: concatenate the features on multiple branches to obtain the final video aesthetic label and video scene type.