End-to-End De-Smoking Model for Video Artifact Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image processing-based de-smoking methods are ineffective for non-homogenous smoke in videos, particularly in time-critical surgeries, due to assumptions of smoke as homogenous media and introduction of color distortions, and data-driven approaches lack dynamic smoke properties.
Innovation Solution
A processor-implemented method using an encoder-decoder model and a non-smoky frame critic model, trained with synthesized smoky video frames and binary smoke segmentation maps, to generate a de-smoking model that removes smoke by extracting frame-level spatial and long-term spatio-temporal features, optimizing objective functions for de-smoking and segmentation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional image processing-based de-smoking approaches are used, then the processing method is simple, but the de-smoking effectiveness is poor due to assuming smoke as homogenous media and introducing color distortions
Solution Approach 1:
The patent replaces conventional image processing methods with a deep learning-based encoder-decoder neural network model. This substitution enables the system to automatically learn complex smoke removal patterns without manual feature engineering, thereby improving de-smoking effectiveness while handling non-homogenous smoke without introducing color distortions.
Solution Approach 2:
The patent changes the approach from assuming smoke as homogenous media to modeling it as non-homogenous by incorporating spatially varying parameters in the neural network. This allows different regions of the video frame to be processed with appropriate parameters, improving accuracy for complex smoke patterns while maintaining computational feasibility.
2Ease of operation
If conventional dehazing techniques are applied, then the processing is straightforward, but color distortions are introduced which are not desirable for critical applications
Solution Approach 1:
The patent replaces conventional dehazing techniques with a trained neural network model that simultaneously performs smoke removal and color correction. This substitution eliminates the need for separate processing steps and prevents color distortions by learning the mapping from smoky to clear frames with accurate color preservation.
Solution Approach 2:
The encoder-decoder model is designed to perform multiple functions simultaneously: smoke removal, color correction, and detail preservation. This multi-functionality is achieved within a single unified model, making the processing both straightforward and effective without introducing harmful color distortions.
3Reliability
If data-driven based de-smoking approaches are used, then training data is required, but the approaches are limited and perform de-smoking only at video frame level without harnessing dynamic properties of smoke
Solution Approach 1:
The patent transitions from frame-level processing to spatio-temporal processing by incorporating temporal dimensions into the neural network model. This allows the model to capture dynamic smoke properties across multiple frames, improving reliability for time-critical applications while effectively utilizing available training data through data augmentation techniques.
Solution Approach 2:
The patent introduces dynamic modeling by capturing temporal variations in smoke patterns across video frames. The model learns to track and remove moving smoke while preserving the dynamic content of the underlying scene, overcoming the limitation of static frame-level approaches and effectively utilizing training data for temporal patterns.
4Quantity of substance
If synthesized smoky video frames are generated for training, then the training dataset can be expanded, but the synthesis process adds complexity to the model generation
Solution Approach 1:
The patent performs preliminary data preparation by generating synthesized smoky video frames before training the main de-smoking model. This pre-processing step creates a comprehensive training dataset that includes diverse smoke patterns, enabling the model to be trained more effectively without adding complexity to the actual training process.
Solution Approach 2:
The patent introduces a smoke generation model as an intermediary component that creates synthetic smoky frames from clear frames. This intermediary serves as a data augmentation tool that expands the training dataset without requiring additional complex processing in the main de-smoking model, effectively separating data preparation from model training.
Data Source
AI summary
The disclosure herein relates to methods and systems for generating an end-to-end de-smoking model for removing smoke present in a video. Conventional data-driven based de-smoking approaches are limited mainly due to lack of suitable training data. Further, the conventional data-driven based de-smoking approaches are not end-to-end for removing the smoke present in the video. The de-smoking model of the present disclosure is trained end-to-end with the use of synthesized smoky video frames that are obtained by source aware smoke synthesis approach. The end-to-end de-smoking model localize and remove the smoke present in the video, using dynamic properties of the smoke. Hence the end-to-end de-smoking model simultaneously identifies the regions affected with the smoke and performs the de-smoking with minimal artifacts. localized smoke removal and color restoration of a real-time video.


