An AI-powered
system for automated, genre-specific, and context-dependent video summarization, consisting of: an input module configured to receive and store healthcare video data comprising a plurality of frames, the input module comprising a storage module configured to store the input video; a centralized
processing unit comprising a plurality of modules implemented by an AI-assisted processor, a memory, and a
graphics processing unit, wherein the memory stores instructions executed by the processor and the
graphics processing unit, the centralized processing module comprising: a data preprocessing module configured to receive an input video having a plurality of frames and normalize the plurality of frames to generate preprocessed frames; a genre-specific complexity calculation module configured to quantify the perceptual complexity of each preprocessed frame using genre-specific
metrics to generate frame-by-frame complexity values; a subshot-wise analysis module configured to aggregate the frame-wise complexity values using a sliding window approach to generate subshot complexity vectors; a
time complexity graph module configured to model temporal relationships between sub-shots by creating a graph in which sub-shots are represented as nodes and temporal similarities between sub-shots are represented as edges; an adaptive
thresholding module configured to dynamically adjust complexity thresholds based on local and global trends in the subshot complexity vectors; a complexity-based
model selection module configured to classify video frames as complex or non-complex based on dynamically adjusted complexity thresholds and pass each frame to an appropriate
neural network architecture; a spatial
feature extraction module connected to the complexity-based
model selection module and configured to extract spatial features from each frame using the corresponding
neural network architecture and apply spatial
pyramid pooling to the extracted features; a temporal
feature modeling module configured to process the spatial features to capture temporal dependencies in the video; and a video summary generation module configured to predict frame importance values, select frames using diversity-aware optimization, and generate a summarized video while maintaining procedural accuracy and relevance; and an output module connected to the
central processing unit and configured to display the aggregated video via a
user interface, wherein the
user interface also facilitates uploading of the input video.