Dynamic Residual Network for Adaptive Video Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video classification methods face challenges in efficiently processing videos with different image complexities due to the use of uniform feature maps, leading to increased computational costs and reduced recognition accuracy.

Innovation Solution

The method involves extracting feature maps with resolutions corresponding to the image complexity of each video, using a dynamic residual network to adjust the resolution based on the difficulty of recognizing image features, thereby improving the efficiency of feature recognition and classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If uniform feature maps with the same resolution are used for all videos, then the network structure can be simplified, but the computational cost increases and classification efficiency decreases for complex videos

Engineering Contradiction:
Improvenetwork structure complexityVSAvoidclassification efficiency
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent applies dynamics by making the feature map resolution adaptive rather than fixed. The network dynamically adjusts the resolution of feature maps based on the complexity of each video, allowing simple videos to use lower resolution (reducing computation) and complex videos to use higher resolution (maintaining accuracy). This resolves the contradiction by making the system flexible rather than static.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of feature map resolution from a fixed uniform value to a variable that adapts to video complexity. By adjusting this parameter based on input characteristics, the system achieves both simplified processing for easy cases and enhanced processing for difficult cases, resolving the efficiency-accuracy tradeoff.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If high-resolution feature maps are used for all videos to maintain accuracy, then classification accuracy is preserved, but computational cost increases significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent dynamically changes the resolution parameter of feature maps based on video complexity assessment. For simple videos, lower resolution is used reducing computational cost while maintaining sufficient accuracy. For complex videos, higher resolution is automatically selected to preserve classification accuracy. This adaptive parameter adjustment resolves the contradiction between accuracy and computational cost.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies partial action by using only the necessary resolution for each video rather than uniformly applying high resolution to all. Simple videos receive partial processing with lower resolution, while complex videos receive full processing with higher resolution, optimizing the balance between accuracy and computational resources.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If low-resolution feature maps are used to reduce computational cost, then processing speed improves, but classification accuracy deteriorates for complex videos

Engineering Contradiction:
Improveprocessing speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts feature map resolution based on real-time assessment of video complexity. This dynamic adaptation ensures that processing speed is optimized for simple videos while accuracy is preserved for complex videos, resolving the contradiction between speed and precision through adaptive behavior.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the resolution parameter adaptively based on video characteristics. The system assesses video complexity and adjusts the feature map resolution parameter accordingly, ensuring optimal balance between processing speed and classification accuracy for each individual video input.

Inventive Principle:
Principle #35Parameter changes

4Ease of manufacture

If the network processes all videos with the same feature map resolution, then the system is simple to implement, but it cannot adapt to different video complexities

Engineering Contradiction:
Improvesystem implementation simplicityVSAvoidadaptability to video complexity
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamics into the system by implementing adaptive resolution selection based on video complexity assessment. The network automatically adjusts its processing parameters without requiring complex manual configuration, achieving adaptability while maintaining implementation simplicity through automated decision-making.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs self-service by automatically assessing video complexity and selecting appropriate feature map resolutions without external intervention. The network adapts to different video types autonomously, maintaining implementation simplicity while achieving high adaptability through self-driven parameter adjustment.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP3852007B1Method, apparatus, electronic device, readable storage medium and program for classifying video
Publication Date: 2024.03.13 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • EP3852007B1 patent drawingFigure 1~2
  • EP3852007B1 patent drawingFigure 3
  • EP3852007B1 patent drawingFigure 4~6

AI summary

Embodiments of the present disclosure provide a method, apparatus, electronic device and computer readable storage medium for classifying a video, relate to the field of artificial intelligence, and in particular, relate to technical fields of computer version and deep learning. A specific implementation of the method includes: acquiring a to-be-classified video and an image complexity of the to-be-classified video; extracting a feature map with a resolution corresponding to the image complexity from the to-be-classified video; and determining a target video class to which the to-be-classified video belongs based on the feature map. This implementation aims at the improvement direction of the efficiency difference of the deep learning network in processing the to-be-classified videos with different image complexities. According to the difficulty degree of recognizing the image features of the to-be-classified videos with different image complexities, the feature maps with the resolutions corresponding to the image complexity of the to-be-classified videos are extracted from the to-be-classified videos, and based on this, the feature maps with low resolutions significantly improve the feature recognition and classification efficiency, thereby improving the classification efficient of the whole videos.