Multitasking Deep Learning Model for Machine Vision Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding schemes are optimized for human vision and do not adequately address the needs of machine vision applications, which require improved coding efficiency and cost reduction in multitasking systems, such as self-driving systems, where multiple tasks need to be performed with limited latency and computational resources.
Innovation Solution
A VCM coding apparatus and method that generates and compresses a common feature map for multiple tasks, allowing for task-specific feature maps to be generated when needed, using deep learning-based models for encoding and decoding, ensuring efficient transmission and performance for both machine and human vision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple single-tasking deep learning models are trained for each task in a multitasking system, then task-specific performance is improved, but the number of models and information to be transmitted increases proportionally, leading to increased system complexity and transmission costs
Solution Approach 1:
The patent merges multiple single-tasking deep learning models into a single multitasking deep learning model that shares common feature extraction layers. This consolidation reduces the number of separate models from N (for N tasks) to one unified model, thereby decreasing system complexity while maintaining task-specific performance through task-specific output layers.
Solution Approach 2:
The patent creates a universal deep learning model that performs multiple functions simultaneously - it can execute multiple different tasks (object detection, segmentation, tracking, etc.) using a single model structure. The model extracts common features that are useful across all tasks while maintaining the ability to perform each specific task through appropriate output layers.
2Adaptability or versatility
If multiple single-tasking deep learning models are trained for each task, then comprehensive task coverage is achieved, but the amount of information to be transmitted and processed increases, leading to higher transmission costs and latency
Solution Approach 1:
By combining multiple task-specific models into one multitasking model, the patent reduces the total amount of information that needs to be transmitted. Instead of transmitting N separate model structures and their respective feature maps, the system transmits a single unified model structure that processes all tasks, thereby reducing transmission costs and energy consumption.
Solution Approach 2:
The patent segments the deep learning model into common feature extraction layers that are shared across all tasks and task-specific output layers. This segmentation allows the system to transmit and process only the essential common features once, then branch out to handle different tasks efficiently, reducing redundant transmission of identical information.
3Manufacturing precision
If existing video coding schemes optimized for human vision are used, then video quality is maximized, but coding efficiency for machine vision applications is insufficient, failing to meet strict latency and computational resource limits
Solution Approach 1:
The patent applies local quality by optimizing the video coding process specifically for machine vision requirements rather than general-purpose human vision. The coding scheme is tailored to preserve features that are critical for machine analysis (such as edge detection, motion vectors, and semantic information) while potentially reducing redundancy for human-perceived quality, thereby improving coding efficiency for machine applications.
Solution Approach 2:
The patent changes the optimization parameters of the video coding scheme from human vision metrics (such as PSNR and SSIM optimized for human perception) to machine vision metrics that prioritize feature preservation for automated analysis. This parameter change enables the coding system to achieve better efficiency and lower latency for machine vision tasks while maintaining adequate quality for the intended application.
Data Source
AI summary
A VCM coding apparatus and a VCM coding method, related to a deep learning-based feature map coding apparatus in a multitasking system for machine vision, are provided for performing default procedures of generating and compressing a common feature map related to multiple tasks implied by an original video. The VCM coding apparatus and the VCM coding method can further generate and compress a task-specific feature map whenever needed for higher performance than obtainable with the common feature map to ensure relatively acceptable performance for both machine vision and human vision.


