An AI-powered system for automatic, genre-specific, and contextual video summarization
Patent Information
- Application Number
- DE202025102735
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2035-05-31
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
FIELD OF THE INVENTIONThe present disclosure relates to an AI-based system for automated, generic, and context-sensitive video summary. The present invention particularly provides an AI-based, generic and context-sensitive video summary system consisting mainly of an input module, a central processing unit and an output module. These work together and configure the system to intelligently analyze, extract, and compress key moments while maintaining process accuracy and relevance.BACKGROUND OF THE INVENTIONIn healthcare, an enormous amount of medical video data is generated, including surgical records, medical procedures, endoscopic, radiology, and robot assisted surgical exposures. Manually evaluating these videos is time consuming and inefficient, particularly when physicians, researchers, and medical students must filter out important moments.Existing generic video summary systems fail healthcare because of the complexity of medical videos. Generic summary models cannot process long, unworked sequences, high precision gestures, and generic variations such as laparoscopic, robotic assisted, and microscopic interventions with different visual and motion characteristics, respectively.In view of the foregoing discussion, it will be apparent that there is a need for an efficient video summary system that can overcome the existing disadvantages of existing generic video summary systems and perform a generic, specific, context aware video summary.SUMMARY OF THE INVENTIONThe present disclosure relates to an AI-based system for automated, generic, and context-sensitive video summary. The invention is an AI-based system for automated, generic and context-sensitive video summary, which has been specially developed for healthcare applications. The system intelligently analyzes medical video, extracts key moments, and compresses long surgical and procedure records while maintaining process accuracy and clinical relevance. It uses generic complexity metrics, adaptive processing paths, and advanced temporal modeling to identify and preserve critical operational steps, tool interactions, and tissue changes, and simultaneously filter out redundant or irrelevant content.The present disclosure aims to provide an AI-based system for automated, generic, and contextual video summary. The system comprises: an input module for receiving and storing multi-frame healthcare video data, wherein the input module comprises a storage module for storing the input video. The system further includes a central processing unit having a plurality of modules implemented by an AI-based processor, a memory, and a graphics processor. The memory stores instructions executed by the processor and the graphics processor. The central processing module comprises: a data preprocessing module for receiving input video having a plurality of frames and normalizing these frames to generate preprocessed frames; a generic complexity calculation module for quantifying the perceptual complexity of each preprocessed frame using generic metrics to generate frame-by-frame complexity scores; a snapshot-by-analysis module for aggregating the frame-by-frame complexity scores using a sliding window approach to generate snapshot complexity vectors; a temporal complexity graph module configured to model temporal relationships between sub-images by creating a graph in which sub-images are represented as nodes and temporal similarities between sub-images are represented as edges; an adaptive thresholding module configured to dynamically adapt complexity thresholds based on local and global trends in the complexity vectors of the sub-images; a complexity-based model selection module configured to classify video images as complex or non-complex based on the dynamically adapted complexity thresholds and to forward each image to a suitable neural network architecture; a spatial feature extraction module connected to the complexity-based model selection module and configured to extract spatial features from each image using the suitable neural network architecture and apply spatial pyramid pooling to the extracted features; a temporal feature modelling module configured to process the spatial features to detect temporal dependencies in the video; A module for generating video digests configured to predict frame importance values, select frames using diversity aware optimization, and generate composite video The system further comprises an output module connected to the central processing unit and configured to display the composite video via a user interface, the user interface also facilitating uploading of the input video.An object of the present disclosure is to provide an AI-based system for automated generic and contextual video summaryAnother object of the present disclosure is an AI-based system configured to intelligently analyze, extract, and compress key moments while maintaining process accuracy and relevance.Another object of the present disclosure is to provide an adaptive processing pipeline system that enables the system to efficiently process different stages of video complexity while maintaining high quality summary results.Another object of the present disclosure is to perform gene specific analyses for various types of medical videos (laparoscopic, endoscopic, cataract, robotic, and microscopic imaging) to ensure appropriate feature extraction and summary.Another object of the present disclosure is to reduce the time spent medical professionals checking long surgical and procedure videos by intelligently compressing content while maintaining clinical relevance and procedure integrity.In order to further clarify the advantages and features of the present disclosure, the invention will be explained in more detail with reference to specific embodiments that are illustrated in the accompanying drawings. These drawings illustrate only typical embodiments of the invention and are therefore not to be considered as limiting the scope thereof. The invention will be described and explained in more detail with reference to the accompanying drawings.BRIEF DESCRIPTION OF THE FIGURESThese and other features, aspects, and advantages of the present disclosure will become more fully understood when the following detailed description is read with reference to the accompanying drawings, in which like characters represent like parts throughout. The following applies here: FIG. 1 shows a block diagram of an AI-based automated generic contextual video summary system according to an embodiment of the present disclosure. FIG. 2 is a diagram showing the operation of the proposed system according to an embodiment of the present disclosure.Those skilled in the art will also appreciate that the elements in the drawings are shown for simplicity and are not necessarily to scale. For example, the flowcharts illustrate the method using the key steps to improve understanding of aspects of the present disclosure. In addition, regarding the construction of the apparatus, individual or multiple components of the apparatus may be represented by conventional symbols in the drawings. The drawings may only show the specific details relevant to understanding the embodiments of the present disclosure in order not to obscure the drawings with details readily apparent to those skilled in the art after the present description.DETAILED DESCRIPTION:In order to aid in the understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings and will be described in an comprehensible manner. However, the scope of the invention is not limited thereby. Changes and further modifications of the illustrated system, as well as further applications of the principles of the invention, are possible, as would normally occur to a person skilled in the art.It will be understood by those skilled in the art that the foregoing general description and the following detailed description are exemplary and explanatory of the invention and are not intended to be limiting thereof.References throughout this specification to "one aspect," "another aspect," or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, the phrases "in one embodiment," "in another embodiment," and similar phrases in this specification may or may not refer to the same embodiment.The terms "comprises," "comprising," or other variations thereof are intended to cover a non-exclusive inclusion, such that a process or method comprising a list of steps may include not only those steps, but also other steps not expressly listed or inherent in that process or method. Likewise, the phrase "comprises... for" one or more devices, subsystems, elements, structures, or components does not exclude, without further limitations, the existence of other devices, subsystems, elements, structures, components, or additional devices, subsystems, elements, structures, or components.Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by one of ordinary skill in the art. The systems, methods, and examples provided herein are for illustrative purposes only and are not to be considered limiting.Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.The functional units described in this specification are referred to as devices. A device may be implemented in programmable hardware devices such as processors, digital signal processors, central processing units, field programmable gate arrays, programmable array logic systems, programmable logic devices, cloud processing systems, or the like. The devices may also be implemented in software for execution by various types of processors. An identified device may include executable code and may consist, for example, of one or more physical or logical blocks of computer instructions, which may be organized, for example, as an object, procedure, function, or other construct. However, the executable of an identified device need not be physically stored at the same location, but may consist of different instructions stored at different locations that, logically linked, form the device and serve its purpose.Device or module executable code may consist of one or more instructions and even be distributed over multiple code segments, different applications, and multiple storage devices. Likewise, operational data may be identified and displayed within the device and presented in any form and data structure. The operational data may be acquired as a single data set or distributed across different storage devices and may be at least partially present as electronic signals in a system or network.References throughout this specification to "a selected embodiment," "an embodiment," or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosed subject matter. Therefore, the terms "a selected embodiment," "in one embodiment," or "in one embodiment" in various places throughout this specification do not necessarily refer to the same embodiment.Moreover, the described features, structures, or characteristics may be combined in any manner in one or more embodiments. The following description contains numerous specific details to provide a thorough understanding of the embodiments of the disclosed subject matter. However, those skilled in the art will appreciate that the disclosed subject matter may be practiced without one or more of the specific details or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail in order not to obscure aspects of the disclosed subject matter.According to the example embodiments, the disclosed computer programs or modules may be executed in a variety of ways, such as an application in a device's memory or a hosted application on a server that communicates with the device application or browser via various standard protocols such as TCP / IP, HTTP, XML, SOAP, REST, JSON, and other suitable protocols. The disclosed computer programs may be written in example programming languages that execute from the memory of the device or from a hosted server, such as BASIC, COBOL, C, C++, Java, Pascal, or scripting languages such as JavaScript, Python, Ruby, PHP, Perl, or other suitable programming languages.Some of the disclosed embodiments include or otherwise involve data transfer over a network, for example, the transfer of various inputs or files over the network. The network may include, for example, the Internet, wide area networks (WANs), local area networks (LANs), analog or digital wired and wireless telephone networks (e.g., PSTN, Integrated Services Digital Network (ISDN), cellular networks and digital subscriber line (xDS)), radio, television, cable, satellite, and / or other transmission or tunneling mechanisms for data transmission. The network may comprise multiple networks or sub-networks, each including, for example, a wired or wireless data path. The network may comprise a circuit switched voice network, a packet switched data network or other network for transferring electronic communication. For example, the network may comprise networks based on Internet Protocol (IP) or Asynchronous Transfer Mode (ATM) and support voice, for example, via VoIP, voice over ATM, or other comparable protocols for voice data communication. In one implementation, the network includes a cellular network configured to exchange text or SMS messages.Examples of the network include a personal area network (PAN), a storage area network (SAN), a home area network (HAN), a campus area network (CAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a virtual private network (VPN), an enterprise private network (EPN), the Internet, a global area network (GAN), etc.FIG. 1 shows a block diagram of an AI-based automated generic contextual video summary system according to an embodiment of the present disclosure.Referring to FIG. 1, the system (100) includes: an input module (102) configured to receive and store multi-frame healthcare video data, wherein the input module (102) includes a storage module (102a) for storing the input video. The system (100) further includes a central processing unit (104) comprising a plurality of modules implemented by an AI-based processor (104a), a memory (104b), and a graphics processor (104c), wherein the memory (104b) stores instructions executed by the processor (104a) and the graphics processor (104c). The central processing module (104) comprises: a data preprocessing module (106) configured to receive input video having multiple frames and normalize those frames to generate preprocessed frames; a generic complexity calculation module (108) configured to quantify the perceptual complexity of each preprocessed frame using generic metrics to generate complexity values on a frame-by-frame basis; a single sub-shot analysis module (110) that aggregates the complexity values for each frame using a sliding window to generate sub-shot complexity vectors. A temporal complexity graphing module (112) that models temporal relationships between sub-shot by creating a graph in which sub-shot are represented as nodes and temporal similarities between sub-shot are represented as edges. An adaptive thresholding module (114) that dynamically adjusts complexity thresholds based on local and global trends in the sub-shot complexity vectors. A complexity-based model selection module (116) that classifies video shot as complex or non-complex based on the dynamically adjusted complexity thresholds and forwards each frame to a suitable neural network architecture. A spatial feature extraction module (118) coupled to the complexity-based model selection module (116) and using the appropriate neural network architecture extracts spatial features from each frame and applies spatial pyramid pooling to the extracted features. A temporal feature modelling module (120) that processes the spatial features to detect temporal dependencies in the video and a video summary generation module (122) configured to predict frame importance values, select frames using diversity aware optimization, and generate a summary video while maintaining method accuracy and relevance. The system (100) further comprises an output module (124) connected to the central processing unit (104) and configured to display the aggregated video via a user interface (126), the user interface (126) also facilitating uploading the input video.In one embodiment, the generic complexity calculation module (108) is configured to calculate: edge density for measuring details in each frame; entropy for detecting randomness and variability in pixel intensities; motion intensity between adjacent frames using dense optical flow; and color distribution by measuring the HSV histogram divergence from a generic template.In one embodiment, the temporal complexity graph module (112) is configured to: calculate edge weights based on temporal similarities between sub-shots, detect scene transitions when the edge weight falls below a similarity threshold, and identify shot redundancies when the edge weight exceeds the similarity threshold.In one embodiment, the adaptive threshold module (114) is configured to set a generic base threshold derived from maximum complexity scores observed in video; apply temporal smoothing to video subshotes to determine adaptive thresholds; increase thresholds during chaotic segments; and decrease thresholds during quiet segments.In one embodiment, the complexity-based model selection module ( 116) is configured to route complex frames to a neural network architecture DensityNet161 and to route non-complex frames to a neural network architecture MobileNetV3.In one embodiment, the temporal feature modeling module (120) includes: a transformer encoder configured to detect long-range global dependencies; and a long short-term memory bidirectional network (Bi-LSTM) configured to detect short-term temporal relationships; wherein outputs of the transformer encoder and the Bi-LSTM network are selectively merged based on frame complexity.In one embodiment, the transformer encoder includes: a generic bias component configured to cut attention to genes of health videos; a multi-head self-attention module configured to apply multiple parallel attention levels and focus on different aspects of the video image; a position encoding component configured to maintain image order; and a feedforward network for feature transformation.In one embodiment, the video summary generation module (122) comprises: a regressive network configured to predict image importance scores; a diversity aware backpack optimization component configured to select images with importance and diversity being considered; and a temporal kernel segmentation component configured to determine scene boundaries and ensure coherent video segments, wherein the diversity aware backpack optimization component is configured to integrate a generic diversity metric into a backpack optimization target; maximize the overall importance score with a predefined summary length constraint; and adjust image selection based on generic redundancy or diversity requirements.In one embodiment, the system (100) is also configured to detect and record critical surgical steps, tool interactions, and tissue changes, analyze various types of medical video, including laparoscopic, endoscopic, cataract, robotic, and microscopic imaging, and maintain method coherence and clinical relevance in the generated video summary.In one embodiment, the user interface (126) of the output module (124) is configured to: display the merged video adjacent to the original video, allow interactive navigation between corresponding segments of the original and merged video, allow the user to adjust the summary parameters, and allow the merged video to be exported in multiple formats.The present invention relates to an AI-based, generic and context sensitive video summary system. It essentially consists of an input module, a central processing unit and an output module. These work together and configure the system so that key moments are intelligently analyzed, extracted, and condensed, maintaining process accuracy and relevance. The system comprises various modules which together enable context-sensitive video summary. The data input module receives and stores healthcare video data consisting of a plurality of frames. The data preprocessing module first normalizes the input video images to provide consistent processing conditions. The generic complexity calculation module then analyzes each frame from various metrics such as edge density, entropy, motion intensity, and color distribution compared to generic templates to quantify perceptual complexity. These pictorial complexity values are aggregated by the sub-shot analysis module using a sliding window to generate temporal patterns identifying significant segments. The temporal complexity graph module generates a graphical representation in which sub-shot nodes and temporal similarities form edges. This enables transitions and redundancies to be detected. The adaptive threshold module dynamically adjusts complexity thresholds to local and global trends by increasing and decreasing the thresholds during chaotic stretches and in quiet phases. Based on these thresholds, the complexity-based model selection module forwards frames to either lightweight processing (MobileNetV3) for simple frames or deeper feature extraction (DenseNet161) for complex frames. The spatial feature extraction module employs suitable neural network architectures followed by spatial pyramid pooling to acquire multi-scale spatial information. For temporal modeling, the system uses a sophisticated architecture that combines transformer encoders with genrespecific tendency to detect dependencies over long distances and bidirectional LSTM networks for short-term relationships, with selective fusion based on frame complexity. Finally, the video summary generation module predicts frame importance values using a regressive network and applies diversity aware backpack optimization to select frames while maintaining method coherence, wherein temporal kernel segmentation ensures smooth transitions between surgical phases. The generated composite video is then displayed by the output module via the user interface.FIG. 2 is a diagram showing the operation of the proposed system according to an embodiment of the present disclosure.As shown in FIG. 2, the proposed system intelligently analyzes, extracts and compresses key moments from medical video recordings, while maintaining the accuracy and relevance of the procedure. The system essentially comprises oneThe data preprocessing module receives the input video consisting of a plurality of frames. Each frame is scaled to a fixed resolution, for example 224 x 224 pixels, to ensure standardized dimensions for efficient processing. After resize, the pixel values of each frame are normalized to a range of [0,1] by division by 255. This normalization allows stable gradient updates during subsequent neural processing. The preprocessing step ensures uniform processing of both visually simple and complex individual images and thus enables consistent and reliable analysis of the entire video content.The generic complexity calculation module analyzes the preprocessed frames to quantify the perceptual complexity of each frame using generic metrics. These metrics include: (1) edge density to measure level of detail within each frame, which facilitates distinguishing between static backgrounds and high-activity areas; (2) entropy to assess randomness and pixel intensity variability, which facilitates filtering out monotone frames; (3) motion intensity between adjacent frames calculated using dense optical flow to emphasize frames with critical motion events; and (4) color distribution measured as the deviation of the HSV histogram of each frame from a generic reference template. The reference template reflects typical color characteristics for certain medical video genes - for example, reddish hues in endoscopic procedures and bluish green hues in laparoscopic or microscopic surgery due to clinical lighting conditions. The module combines these complexity metrics using generic weights to generate complexity values frame by frame, thus enabling precise and contextual video summary.The individual subshote analysis module uses a sliding window approach to process the imagewise complexity values generated by the generic complexity calculation module. Rather than handling individual frames independently, this module aggregates complexity metrics over a fixed time window to form sub-shot complexity vectors. These vectors represent multi-dimensional time-series data that captures local context trends and helps to recognize generic patterns such as gradual increases in motion intensity, abrupt variations in edge density, and subtil shifts in color distribution. For each sub-shot, the average complexity over the windowed frames is calculated, with the window size determining the number of frames contained per sub-shot.The "temporal complexity graph" module models temporal relationships between subshotes by creating a graph in which each node corresponds to a subshote and each edge represents the temporal similarity between the subshotes. These edges are calculated using a similarity measure that takes both appearance and motion features into account and is controlled by a parameter for the trade-off between appearance and motion and a Gaussian scaling factor. The graph facilitates the detection of scene transitions when the similarity between subshotes falls below a defined threshold and identifies setting redundancies when the similarity exceeds the threshold.The adaptive threshold module receives sub-shot level complexity values and dynamically adjusts the complexity thresholds to local and global trends in video. First, a generic base threshold is established which serves as a reference point for further adaptations. This base value is determined by scaling the maximum observed complexity value in the video using a generic factor to ensure matching with content-specific features. Temporal smoothing is then applied to the video subshotes to calculate adaptive thresholds. These thresholds are increased in visually chaotic sections to avoid over selection of images and reduced in calmer sections to capture critical but visually less complex moments. This dynamic adjustment ensures stability in image selection and attenuates abrupt fluctuations in the importance values by the inclusion of a temporal smoothing factor.The complexity-based model selection module uses preprocessed frames and corresponding adaptive thresholds as input. Based on the calculated complexity, the module classifies video recordings into two categories: complex and non-complex. Complex frames are passed to a neural network architecture DensityNet161, which allows for deep feature extraction suitable for high-detail content. Non-complex frames, on the other hand, are processed using a neural network architecture MobileNetV3, which is optimized for the simple processing of simpler contents. Each frame is assigned the corresponding sub-picture and spatial features are extracted accordingly.The spatial feature extraction module, coupled to the complexity-based model selection module, extracts spatial features from each frame using the selected neural network architecture. A spatial pyramid pooling (SPP) layer is used to detect multi-scale spatial features. The SPP layer performs pooling operations over fixed size regions at multiple scales, thus obtaining spatial information across different resolutions. This multi-level pooling ensures a robust feature representation regardless of the dimensions of the input frame.The temporal feature modeling module processes spatial pyramid pooling features received from the spatial feature extraction module. This module consists of a transformer encoder and a bidirectional long-short-term memory network (Bi-LSTM), which jointly records both long-term and short-term time dependencies in medical video data. The transformer encoder component is configured to detect long term global dependencies such as tool changes, anatomical changes, and stepwise surgical procedures. It utilizes a multiple-head self-attention mechanism that allows the model to focus multiple aspects of video images in parallel, such as motion, color distribution, and object presence. The outputs of the multiple attention heads are linked and linearly transformed to obtain the embedding dimension. In order to obtain the sequence of the images, the transformer encoder has a position coding component which ensures that the chronological sequence is maintained during the processing. In addition, a generic bias component is integrated into the attention mechanism to adjust focus according to the video generic. In medical videos, this allows the model to prioritize medically relevant features such as patterns of motion and subtilous tissue changes. The gene-specific distortion serves as regularization, prevents the over emphasis of individual features and at the same time supports gene adaptation. Each transformer layer includes a feedforward network to introduce nonlinearity and transform the features under consideration. This enables more expressive representations for subsequent tasks. The Bi-LSTM component is configured to capture short term time dependencies by processing the spatial features both forward and backward. This is particularly useful in slower surgical phases such as preliminary steps. The final Bi-LSTM representation for each image is obtained by the concatenation of the forward and backward hidden states, thereby creating contextual feature encoding. The outputs of the transformer encoder and the Bi-LSTM network are selectively merged based on the frame complexity determined by the complexity-based model selection module. For complex images, the fused feature representation combining both transformer and Bi-LSTM ausgaben is used to improve temporal modeling. For non-complex images, the Bi-LSTM module is bypassed to reduce computational effort and temporal features are obtained directly from the output of the transformer encoder. These temporal features are then passed to the video summary generation module for prediction and summary of importance.The video summary generation module is configured to generate a contiguous summarized video using the temporal features obtained from the temporal feature modeling module. These temporal features are processed by a regression network within the video summary generation module that predicts the importance values of each individual image in the video. Based on a predefined summary length constraint, a diversity aware backpack optimization component selects a subset of images based on the predicted importance values. This selection mechanism includes a generic diversity metric to ensure that the selected images are both highly informative and sufficiently versatile. In the context of health videos, this approach enables the effective representation of various method steps, such as transitions in surgical instruments and important operation steps, while minimizing excessive redundancy. After the selection of the images, the kernel time segmentation component is applied to determine scene boundaries and ensure coherent video segments within the generated summary. This sequence of operations ensures that the resulting composite video maintains process accuracy and clinical relevance and matches the original content.In one embodiment, the AI-based system for automated generic contextual video summary is configured specifically for medical applications, but is not limited thereto. In this regard, the video summary generation module is configured to maintain process accuracy by detecting and maintaining critical surgical steps, tool interactions, and tissue changes while minimizing redundancy and excluding irrelevant content. The system is configured for analysis of various medical video types, including laparoscopic, endoscopic, cataract, robot assisted and microscopic photographs. The generic complexity calculation module quantitates perceptual complexity from metrics such as entropy, edge density, motion intensity, and color distribution. Based on these metrics, the system dynamically adjusts the image selection to optimize spatial and temporal feature extraction to obtain significant process details relevant to the particular medical context. The module for modeling temporal features comprises a transformer encoder with a multi-head self-attenuation mechanism and a generic bias component. This configuration allows the system to prioritize images with surgical instruments, tissue interactions, and motion patterns. This highlights substantial information, filters out non-essential images, and maintains method coherence and clinical relevance. The temporal complexity graph module models temporal relationships between sub-images by calculating edge weights and thus enables the recognition of image redundancies. This ensures the logical structuring of surgical sequences and maintains the process coherence. The video summary generation module includes a diversity aware backpack optimization component that uses mathematical optimization techniques to balance information content and redundancy. This configuration ensures that critical process steps are maintained and excessive repetitions are avoided. The video summary generation module also includes a kernel time segmentation component that determines scene boundaries and enables smooth and meaningful transitions between surgical phases. This ensures that the generated summary is comprehensible and easily comprehensible.The drawings and the foregoing description show examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be divided into multiple functional elements. Elements of one embodiment may be added to another embodiment. For example, the order of the processes described herein may be changed and is not limited to the manner described herein. Moreover, the actions of a flow chart need not be performed in the order shown; nor do all actions necessarily need to be performed. Also, actions that are not dependent on other actions may be performed in parallel with the other actions. The scope of the embodiments is by no means limited by these specific examples. Numerous variations, whether or not explicitly stated in the specification, such as differences in structure, dimensions, and material use, are possible. The scope of the embodiments is at least as broad as recited in the following claims.Advantages, other advantages and solutions to problems have been described above with reference to specific embodiments. However, the advantages, merits, solutions to problems and any components that may result in an advantage, merit or solution being introduced or enhanced are not to be understood as critical, required or essential features or components of individual or all claims.REFERENCES100 The system consists of: an input module. 102 Input module 102 a Speichermodul module 104 Central processing unit 104 a AI-based processor 104 b Speicher 104 c Grafikprozessor processor 106 Data preprocessing module 108 Genrespecific complexity calculation module 110 Subshotwise analysis 112 Temporal complexity graph module 114 Adaptive threshold module 116 Complexity-based model selection module 118 Spatial feature extraction module 122 Video summary generation module 124 Output module 126 User interface 202 Input video image 204 Data preprocessing for generation of scaled and normalized frames 206 Genrespecific complexity analysis for calculation of frame complexity score 208 Subshot level analysis for calculation of shot level complexity values 210 Temporal complexity graph for detection of scene transitions, and Setting redundancy 212 Adaptive thresholding to regulate frame selection 212 Diversity-aware Backpack optimization with KTS to generate a video summary 214 Diversity-aware Backpack optimization with KTS to generate a video summary 216 Regressive network to calculate importance scores 218 Bi-LSTM and Transform encoders to obtain fused spatio-temporal features 220 Spatial pyramid pooling for multi-scale spatial features 222 Model selection from MobileNet V3 and DenseNet 161 to extract spatial features 224 Output module to display a video summary
Claims
A AI-based system for automated, generic, and contextual video summary, comprising: an input module configured to receive and store healthcare video data comprising a plurality of frames, the input module comprising a storage module configured to store the input video; a centralized processing unit comprising a plurality of modules implemented by an AI-based processor, a memory, and a graphics processing unit, the memory storing instructions executed by the processor and the graphics processing unit, the centralized processing module comprising: a data preprocessing module configured to receive an input video comprising a plurality of frames and normalize the plurality of frames to generate preprocessed frames; a generic complexity calculation module configured to quantify the perceptual complexity of each preprocessed frame using generic metrics to generate imagewise complexity values; a sub-shot-like analysis module configured to aggregate the frame-like complexity values using a sliding window approach to generate sub-shot complexity vectors; a temporal complexity graph module configured to model temporal relationships between sub-shots by creating a graph in which sub-shots are represented as nodes and temporal similarities between sub-shots are represented as edges; an adaptive threshold module configured to dynamically adjust complexity thresholds based on local and global trends in the sub-shot complexity vectors; a complexity-based model selection module configured to classify video images as complex or non-complex based on the dynamically adjusted complexity thresholds and to forward each image to a suitable neural network architecture; a spatial feature extraction module connected to the complexity-based model selection module and configured to extract spatial features from each frame and apply spatial pyramid pooling to the extracted features using the corresponding neural network architecture; a temporal feature modeling module configured to process the spatial features to detect temporal dependencies in the video; A module for generating video summarys, configured to predict frame importance values, select frames using diversity aware optimization, and generate summary video while maintaining method accuracy and relevance; and an output module, coupled to the central processing unit, configured to display the summary video via a user interface, the user interface also facilitating uploading of the input video.The system of claim 1, wherein the generic complexity calculation module is configured to calculate: edge density to measure details in each frame; entropy to detect randomness and variability in pixel intensities; motion intensity between adjacent frames using dense optical flow; and color distribution by measuring the HSV histogram divergence from a generic template.The system of claim 1, wherein the temporal complexity graph module is configured to: calculate edge weights based on temporal similarities between sub-shots, detect scene transitions when the edge weight falls below a similarity threshold, and identify shot redundancies when the edge weight exceeds the similarity threshold.The system of claim 1, wherein the adaptive threshold module is configured to set a generic base threshold derived from maximum complexity scores observed in video; apply temporal smoothing to video subshotes to determine adaptive thresholds; increase thresholds during chaotic segments and decrease thresholds during quiet segments.The system of claim 1, wherein the complexity-based model selection module is configured to route complex frames to a neural network architecture DensityNet161 and to route non-complex frames to a neural network architecture MobileNetV3.The system of claim 1, wherein the temporal feature modeling module comprises: a transformer encoder configured to detect long-range global dependencies; and a long short-term memory (Bi-LSTM) bidirectional network configured to detect short-term temporal relationships; wherein outputs of the transformer encoder and the Bi-LSTM network are selectively merged based on frame complexity.The system of claim 6, wherein the transformer encoder comprises: a generic bias component configured to cut attention to generic health video; a multi-head self-attention module configured to apply multiple parallel attention levels and focus on different aspects of the video image; a position encoding component configured to maintain image order; and a feedforward network for feature transformation.The system of claim 1, wherein the video summary generation module comprises: a regressive network configured to predict image importance scores; a diversity aware backpack optimization component configured to select images with importance and diversity being considered; and a temporal kernel segmentation component configured to determine scene boundaries and ensure coherent video segments, wherein the diversity aware backpack optimization component is configured to integrate a generic diversity metric into a backpack optimization target; maximize the overall importance score with a predefined summary length constraint; and adjust the image selection based on generic redundancy or diversity requirements.The system of claim 1, wherein the system is further configured to detect and record critical surgical steps, tool interactions, and tissue changes; analyze different types of medical video including laparoscopic, endoscopic, cataract, robotic, and microscopic imaging; and maintain method coherence and clinical relevance in the generated video summary.The system of claim 1, wherein the user interface of the output module is configured to: display the merged video adjacent to the original video; enable interactive navigation between corresponding segments of the original and merged video; enable the user to adjust the merging parameters; and allow the merged video to be exported in multiple formats.
Citation Information
Cited By
Fan blade fault detection method and system, electronic equipment and storage medium
CN121561560A