Neural Network Hand Shake Correction for Video Stabilization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional hand shake correction methods for video images in electronic devices face performance degradation when using images acquired before or after the current frame, and require significant memory usage and computation time due to storage and processing requirements.
Innovation Solution
An electronic device equipped with a camera, motion sensor, and processor uses a weight learned through an artificial neural network to detect movement and correct the camera position in real-time, reducing memory usage and computation time by processing images more efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If images acquired before or after the current frame are used for hand shake correction, then correction performance may be maintained, but memory usage increases and computation time increases
Solution Approach 1:
The patent extracts only the essential correction information from multiple image frames rather than storing and processing all frames. By using a neural network to learn correction patterns from limited frames (current frame and optionally one previous frame), the system extracts the necessary correction data without retaining all intermediate frames in memory, thus reducing memory consumption while maintaining correction effectiveness.
Solution Approach 2:
The patent changes the parameter of frame utilization by not requiring a fixed number of future frames for correction. Instead, the neural network adapts to determine the optimal number of frames to use based on the specific shaking conditions, allowing the system to reduce memory usage when fewer frames are sufficient while maintaining correction performance when more frames are needed.
2Reliability
If images acquired before or after the current frame are used for hand shake correction, then correction performance may be maintained, but computation time increases
Solution Approach 1:
The patent applies preliminary action by pre-training the neural network offline to learn hand shake correction patterns from large datasets. During actual video processing, the pre-trained network rapidly applies learned patterns to correct frames in real-time, avoiding the need for complex online computations and reducing computation time while maintaining high correction performance.
Solution Approach 2:
The patent replaces traditional mechanical image processing methods (which involve extensive frame-by-frame analysis and comparison) with a neural network-based computational approach. The neural network efficiently processes frame data and generates correction information, substituting complex mechanical processing algorithms with a more efficient learned model that reduces computation time.
3Loss of time
If the current image is corrected using images acquired before the current image, then computation time is reduced, but correction performance is degraded
Solution Approach 1:
The patent applies dynamics by making the neural network adaptive to different shaking conditions in real-time. The network dynamically adjusts its processing based on the characteristics of the current and previous frames, allowing it to achieve high correction performance even when using limited frame data, thus resolving the trade-off between computation time and correction quality.
Data Source
AI summary
An electronic device according to various embodiments of the present invention, may include a camera, a motion sensor, a memory, and at least one processor, wherein the at least one processor may be configured to, by using the camera, acquire a first image frame, a plurality of second image frames successive to the first image frame, and a third image frame immediately before the first image frame, while the camera acquires the third image frame, the first image frame, and the plurality of the second image frames, detect a movement of the electronic device using the motion sensor, determine a first position of the camera corresponding to the first image frame and a plurality of second positions of the camera corresponding to the plurality of the second image frames respectively, based at least in part on the movement of the electronic device, the first image frame, the plurality of the second image frames, and the third image frame, correct the first position, by conducting computations using a weight learned through an artificial neural network, the first position, the plurality of the second positions, and a post-correction position of a third position of the camera corresponding to the third image frame, and correct the first image frame, based at least in part on the corrected first position.


