Method and system to generate short videos

The method and system generate a playlist of short videos from immutable videos by detecting shots and classifying scenes, addressing content discoverability and resource efficiency challenges, and offering personalized video delivery.

WO2025172932A1PCT designated stage Publication Date: 2025-08-21TATA PLAY LTD

Patent Information

Application Number
PCT/IB2025/051618
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-15
Filing Date
2025-02-14
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing video databases face challenges in content discoverability and resource consumption due to the generation of short videos from long-form content, which often overlooks less popular content and requires significant computing and memory resources for clipping and editing.

Method used

A method and system that processes immutable videos to detect shots and classify scenes into genres without editing, generating a playlist of short videos using metadata, and dynamically adjusts the playlist based on user behavior to provide personalized content.

Benefits of technology

This approach enhances content discoverability by providing diverse videos based on user interests, reduces resource consumption by storing only metadata, and optimizes computational resources, resulting in personalized and efficient video delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025051618_21082025_PF_FP_ABST
    Figure IB2025051618_21082025_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to a method and a system for generating a playlist of short videos from an immutable video. The method comprises processing the immutable video to detect a plurality of shots within the immutable video and processing the immutable video to detect and classify a plurality of scenes into at least a plurality of genres. The method comprises identifying a plurality of short scenes by combining classification data associated with the plurality of scenes and the plurality of shots and storing a metadata related to each short scene of the plurality of short scenes in a database. The metadata comprises at least a timestamp of the short scene. The method also comprises generating the playlist of short videos using the metadata without editing the immutable video.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TITLE: METHOD AND SYSTEM TO GENERATE SHORT VIDEOS

[0002] TECHNICAL FIELD

[0003] The following disclosure relates to video processing and in particular to a method and a system to generate a playlist of short videos from immutable videos.

[0004] BACKGROUND OF THE DISCLOSURE

[0005] Video databases (for e.g., Over the Top (OTT) platforms) may include various long-form video contents such as, but not limited to, movies, shows, live telecasts, web-series, and gaming. Short videos, such as of a duration of 30 seconds to 120 seconds, may be generated from the long-form video contents and may be rendered to users. In some cases, the video databases provide an abundance of long-form content, resulting in challenges related to content discoverability. For instance, the video databases use black box content that only provides the most popular and trending short videos to the users for example, based on user’s interests. However, this may often overshadow other unknown and less popular short videos or video content from the long-form video contents that may be of similar interest to the user.

[0006] Moreover, generating short videos from the long-form content generally includes clipping and / or editing the long-form video content and storing the short videos in servers. Thus, the current methods consume significant computing resources for clipping and / or editing the long-form video content and memory resources to store the short videos.

[0007] The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the disclosure and should not be taken as acknowledgment or any form of suggestion that this information forms prior art already known to a person skilled in the art.

[0008] SUMMARY

[0009] Disclosed herein is a method of generating a playlist of short videos from an immutable video. The method comprises processing the immutable video to detect a plurality of shots within the immutable video and processing the immutable video to detect and classify a plurality of scenes into at least a plurality of genres. The method comprises identifying a plurality of short scenes by combining classification data associated with the plurality of scenes and the plurality of shots and storing a metadata related to each short scene of the plurality of short scenes in a database. The metadata comprises at least a timestamp of the short scene. The method also comprises generating the playlist of short videos using the metadata without editing the immutable video.

[0010] Also disclosed herein is an apparatus to generate a playlist of short videos from an immutable video. The apparatus comprises a memory and a processor communicatively coupled with the memory. The processor is configured to process the immutable video to detect a plurality of shots within the immutable video and process the immutable video to detect and classify a plurality of scenes into at least a plurality of genres. The processor is also configured to identify a plurality of short scenes by combining classification data associated with the plurality of scenes and the plurality of shots and store a metadata related to each short scene of the plurality of short scenes in a database. The metadata comprises at least a timestamp of the short scene. The processor is further configured to generate the playlist of short videos using the metadata without editing the immutable video.

[0011] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.

[0012] BRIEF DESCRIPTION OF DRAWINGS

[0013] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The same numbers are used throughout the figures to reference like features and components. Some embodiments of system and / or methods in accordance with embodiments of the present subject matter are now described, by way of example only, and regarding the accompanying figures, in which: FIG. 1 illustrates an exemplary architecture of a system to generate a playlist of short videos in accordance with an embodiment of the present disclosure;

[0014] FIG. 2 illustrates a detailed block diagram of a Short Video Generation System (SVGS) to generate the playlist of the short videos in accordance with an embodiment of the present disclosure;

[0015] FIG. 3 illustrates an exemplary flowchart of a method to generate the playlist of short videos from the immutable video in accordance with an embodiment of the present disclosure; and

[0016] FIG. 4 illustrates a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.

[0017] It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether such computer or processor is explicitly shown.

[0018] DETAILED DESCRIPTION

[0019] In the present document, the word “exemplary” is used herein to mean “serving as an example, instance, or illustration”. Any embodiment or implementation of the present subject matter described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments. While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however, that it is not intended to limit the disclosure to the specific forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure.

[0020] The terms “comprise”, “comprising”, “includes”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device, or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises... a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or method.

[0021] In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense.

[0022] Methods and systems of the present disclosure provide various techniques of generating short videos without editing the long-form video content, herein after referred to as original video, from the video databases. A method according to the present disclosure includes detecting shot boundaries of the original video, and classifying various scenes in the original video into a plurality of genres and / or sub-genres based on sub-titles or captions of the original video. Further, the methods include detecting short scenes of the original video by combining the classifications and the shots and generating a playlist of the short scenes. Further, the method includes rendering the generated playlist to a user and dynamically fine-tuning the playlist based on a viewing behavior of the user.

[0023] FIG. 1 illustrates an exemplary architecture of a system to generate short videos in accordance with an embodiment of the present disclosure.

[0024] As shown in FIG. 1, the system architecture 100 may comprise one or more components configured to generate short videos. In one embodiment, the system architecture 100 may include, but not limited to, a Short Video Generation System (SVGS) 102, a user device 104 of a user 105, one or more video servers 106, a training database, 108 and a metadata database 109. The SVGS 102 may be connected to the user device 104, the one or more video servers 106, the training database 108, and the metadata database 109 using a communication network such as, but not limited to, an Internet, a mobile communication network and the like. In some embodiments, the SVGS 102 may be configured within the user device 104.

[0025] The SVGS 102 may be configured to retrieve a plurality of long-form video contents, or immutable videos 110 from the one or more video servers 106 and may generate a playlist 112 of short videos without editing or clipping the plurality of long-form video contents 110 as may be explained below in detail. The plurality of long-form video contents 110 may be defined as original videos that may be provided only by authorized one or more video servers. For example, a long- video may be a full-length movie, or a web-series streamed by an Over The Top (OTT) platform. As the SVGS 102 generates the plurality of short videos from the long-form video content 110 without editing or clipping the long-form video content 110 as may be explained herein later and hence the long-form video content 110 may be referred to as the immutable video 110.

[0026] The playlist 112 of short videos may be defined as a list of small portions of videos extracted from one or more of the plurality of immutable videos 110. For example, a short video may be a 1- minute portion of the full-length movie. In some embodiments, one or more visual dimensions of the playlist 112 of short videos may be same as one or more display parameters of a user interface of the user device 104. For example, visual dimensions, such as pixel dimensions, of the short video may be same as that of a pixel dimension of the user interface 114 of a smart phone or a smart watch 104. The playlist 112 may be displayed in a swipe able or scrollable manner on the user interface 114 of the user device 104.

[0027] The SVGS 102 may process each immutable video 110 to detect a plurality of shots and detect and classify a plurality of scenes into a plurality of genres. The SVGS 102 may further identify a plurality of short scenes by combining classification data associated with the plurality of scenes and the plurality of shots. The SVGS 102 may comprise one or more Artificial Intelligence / Machine Learning (AI / ML) models configured to detect shots and to classify each scene into one or more genres and / or sub-genres. To train the AI / ML models, the SVGS 102 may retrieve one or more training datasets from the training database 108.

[0028] The SVGS 102 may store metadata related to each short scene in the metadata database 109. The metadata may include, but not limited to, at least an identifier of corresponding immutable video 110 from which the short scene has been identified, an identifier of the short scene, one or more genres associated with the short scene, and a resource link to the corresponding immutable video 110. Further, the SVGS 102 may retrieve the metadata from the metadata database 109 and generate the playlist 112. The SVGS 102 may provide the playlist 112 to the user device 104 to display to the user 105. Further, the SVGS 102 may dynamically modify one or more parameters of the playlist of short videos based on a viewing behavior of the user 105.

[0029] The user device 104 may be any device that is configured to display the playlist 112 of short videos, herein also may be referred to as the playlist 112, to the user 105 and receive one or more user inputs associated with the viewing behavior of the user 105 related to the playlist 112. The user device 104 may include, but not limited to, a laptop device, a desktop device, a smartphone device, a smart watch, a Head up display (HUD) device, a virtual display device, an immersive display device or any other device capable of performing the aforementioned functions. The user device

[0030] 104 may comprise a user interface 114 configured to display the playlist 112 and receive the one or more user inputs. The user interface 114 may comprise at least one output device configured to display the playlist 112 and at least one input device configured to receive the one or more user inputs. The at least one output device may be any display device such as, but not limited to, Light Emitting Diode (LED) display, Organic LED display, Liquid Crystal Display (LCD) and the like. The at least one input device may comprise, but not limited to, a touch screen, a remote control, and the like.

[0031] The user 105 may be any user viewing the playlist 112 on the user device 104 and providing the one or more user inputs. In an example, the user 105 may swipe through the playlist 112. The user

[0032] 105 may skip few short videos of the playlist 112 or may view few short videos completely or partially. Thus, the one or more user inputs may include skipping, partial viewing or complete viewing of the user 105 or a time period for which the user 105 views each short video. The SVGS 102 may assess the viewing behavior of the user 105 based on the one or more user inputs.

[0033] The plurality of video servers 106 may include a plurality of server computing devices storing the plurality of immutable videos 110 and authorized to render or stream the plurality of immutable videos 110. The plurality of video servers 106 may be associated with a plurality of authorized entities hosting, delivering and / or streaming the plurality of immutable videos 110 to a plurality of subscribed users. Accordingly, the user 105 may be required to subscribe with the plurality of authorized entities to access the plurality of immutable videos 110. The video server 106 may also be referred to herein as an immutable video database in the present disclosure.

[0034] The training database 108 may be a memory comprising a plurality of training datasets to train one or more AI / ML models of the SVGS 102 to generate the playlist 112. The training database 108 may be implemented either as a hardware storage system or a software memory module to store and retrieve the plurality of training datasets.

[0035] The metadata database 109 may be configured to store the metadata of the plurality of short scenes or any other portions of the immutable video 110 which may later be retrieved by the SVGS 102 to generate the playlist 112. The metadata database 109 may be implemented either as a hardware storage system or a software memory module to store and retrieve the metadata. The metadata may be stored in any type of data structure such as, but not limited to, tables.

[0036] Each of the training database 108 and the metadata database 109 may be located at a same location or at different locations. Further, the training database 108 and the metadata database 109 may be implemented as centralized databases or distributed databases.

[0037] In operation, the SVGS 102 may process each of the plurality of immutable videos 110 to detect a plurality of shots and corresponding plurality of shot boundaries within each immutable video 110. Further, the SVGS 102 may further process textual content, such as, but not limited to, subtitles, of each immutable video to detect and classify the plurality of scenes into the plurality of genres. Further, the SVGS 102 may identify a plurality of short scenes based on classification data associated with the plurality of scenes and the plurality of shots. For example, the plurality of short scenes may be attentive portions or interesting scenes from each immutable video. The SVGS 102 may store a metadata related to each short scene in the metadata database 109. Further, the SVGS 102 may retrieve the metadata from the metadata database 109, generate the playlist 112 based on the metadata and user preferences and past history of the user 105 and may display the playlist 112 on the user device 104. Furthermore, the SVGS 102 may also monitor the viewing behavior of the user 105 based on a time period for which the user 105 has watched each short video of the playlist and may dynamically fine-tune one or more parameters of the playlist 112.

[0038] Thus, the system 100 renders various short videos to the user 105 based on user preferences and user interests irrespective of trending or rating of the short videos, thereby providing wide diversity of videos in each genre which may not be displayed due to overrated videos. Further, the system 100 also reduces memory resources required to generate the playlist 112 by storing only the metadata of the short scenes in the metadata database 109. Thus, the system 100 generates the playlist 112 of short videos without editing or clipping the original video content or the immutable videos which may not be always accessible to content aggregators or video applications. Moreover, by dynamically fine-tuning the parameters of the playlist 112 based on the user inputs, the system 100 provides a personalized and optimized playlist to the user 105.

[0039] A detailed explanation of the SVGS 102 is provided with reference to Fig. 2.

[0040] FIG. 2 illustrates a detailed block diagram of the SVGS 102 in accordance with an embodiment of the present disclosure.

[0041] The SVGS 102 may comprise, without limiting to, a processor 202, a memory 204, a plurality of modules 206, and data 208. The processor 202 may be any hardware processing system such as a microprocessor, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or System on Chip (SOC), Electronic Control unit (ECU) or any other type of processing system. The memory 204 may be any type of hardware data storage unit such as a Read Only Memory (ROM), Random Access Memory (RAM), a temporary or a permanent storage system.

[0042] The plurality of modules 206 may include a shot boundary detection module 210, a genre classification module 212, a short-scene identification module 214, and a dynamic playlist generation module 216. In some embodiments, the plurality of modules 206 may also be configured within the processor 202. The data 208 may include playlist file 218, a plurality of shots 220, classification data 224, a plurality of short scenes 228, metadata 230, and the playlist 112. The shot detection module 210 may comprise at least a Neural Network (NN) model 222 and the genre classification module 212 may comprise at least an ML model 226.

[0043] Techniques herein describe processing of a single immutable video 110 among the plurality of immutable videos 110 to generate the playlist 112 of the short videos only for the sake of simplicity and not limiting in any manner. It may be appreciated that the techniques described herein may be applicable to the plurality of immutable videos 110 which can be processed in parallel to generate the playlist 112.

[0044] In an embodiment, the shot detection module 210 may be configured to retrieve the immutable video 110 from a video server 106 and detect one or more boundaries between shots within the immutable video 110. In general, a video may be composed of a series of images called frames. A sequence of frames captured from the same camera at the same time may be defined as a shot. A group of shots relating to a specific event or narrative may be called a scene. The shot detection module 210 may retrieve the immutable video 110 using a playlist file 218 comprising streaming data of the immutable video 110 such as, but not limited to, resolution, codec and bandwidth requirements to stream the immutable video. In one example, the playlist file 218 may include, but not limited to, M3U8 file.

[0045] The shot detection module 210 may retrieve a plurality of frames of the immutable video 110 from the playlist file 218 and process the plurality of frames to identify a plurality of shots 220, shot boundaries and scene boundaries within the immutable video 110. The shot boundaries may be defined as time instants at which a previous shot ends and a current shot begins in the immutable video 110. The shot boundaries may be represented in units of time or a number of the frame within the immutable video 110. For example, a shot boundary may be at 3611thsecond in the immutable video 110. In another example, the shot boundary may be at 40thimage frame in the immutable video 110. In another example, the plurality of shots within the immutable video 110 may be detected in time (seconds) as 0 to 4, 5 to 199, 200 to 215, 217 to 323, 324 to 437, 438 to 603, 604 to 2342, 2343 to 2586, 2587 to 2922, 2923 to 3026, 3027 to 3383, 3384 to 3611, 3612 to 3683, 3684 to 3709, 3710 to 3746, 3747 to 3778, and 3779 to 3827. The shot detection module 210 may be configured with a neural network (NN) model 222 such as, but not limited to, a Deep Neural Network (DNN) or a Convolutional NN (CNN), to process the plurality of frames to identify the shot boundaries. The shot detection module 210 may train the NN model 222 using the plurality of training datasets from the training database 108 to identify the shot boundaries within a video. The plurality of training datasets may comprise playlist files of a plurality of immutable videos 110 and corresponding shot boundaries. The NN model 222 may comprise a plurality of layers of interconnected nodes, also termed as neurons. The NN model 222 may also comprise an input layer, one or more hidden layers and an output layer. The input layer may receive input data, the one or more hidden layers may process the input data through weighted connections and activation functions; and the output layer may generate predictions based on the processing. Each connection between neurons may comprise an associated weight that may be modified as the NN model 222 learns from input data and / or training data.

[0046] During the training, the shot detection module 210 may continuously update one or more parameters of the NN model 222 to generate predictions with desired accuracy. The one or more parameters may include, but not limited to, a number of nodes, associated weights, types and number of activation functions, number of layers and the like. The desired accuracy may be automatically set by the shot detection module 210 or may be provided manually by a person. In some embodiments, the shot detection module 210 may be configured with an Optimal Sequential Grouping (OSG) for video scene detection or video shot detection combined with deep learning. In other embodiments, the shot detection module 210 may also use local invariant video feature descriptor such as, but not limited to, CSIFT to detect the shots.

[0047] The genre classification module 212 may process the textual content of the immutable video 110 to detect and classify the plurality of scenes into at least the plurality of genres. The genre classification module 212 may extract a textual content of the immutable video 110 from the playlist file 218 or from the corresponding video server 106. The textual content may include, but not limited to, subtitles, captions, and title of the immutable video 110. Further, the genre classification module 212 may analyze the textual content to detect classification data 224 comprising one or more genres, one or more sub-genres, one or more keywords and one or more tags associated with each of the plurality of scenes of the immutable video 110. The one or more genres may include, but not limited to, action, romance, comedy, tragedy, horror, devotional, drama, gaming, thriller, and the like.

[0048] The one or more sub-genres may be defined as a sub-classification of the genre to provide further features of the scene. The genre classification module 212 may assign a sub-genre based on the corresponding one or more genre. For example, emotional drama’ may be sub-genre of the genre ‘drama’.

[0049] The one or more keywords may be defined as one or more words or phrases that may be used to describe content within the scene. The one or more keywords may be identified based on various features of the immutable video 110. The various features may include, but not limited to, actors or actresses, a language, a location, a specific event within each scene of the immutable video 110 and the like. The one or more keywords may enable the SVGS 102 to identify and select specific scenes for the user 105 based on the user interests and provide personalized playlist 112 to the user 105. For example, few keywords may be ‘bull fight’, ‘animal fight’, ‘shot in France’, ‘actor speech’, ‘vehicle, ‘accident’, ‘pedestrian’ and the like.

[0050] The one or more tags may be defined as one or more textual identifiers required to identify a scene or the immutable video 110 as a whole. The genre classification module 212 may identify one or more tags based on various features of each scene or the immutable video 110. The various features may include, but not limited to, a cast and crew, a release time period, a language, a location of recording or launch, unique events of the immutable video 110 and the like. For example, few tags may be ‘ Kannada-language films’, ‘Films scored by ABC (music director)’, ‘2010s Kannada- language films’, ‘2016 films’ and the like. The one or more tags may enable the SVGS 102 to group similar scenes together and may facilitate efficient searches.

[0051] Further, the genre classification module 212 may be configured with the ML model 226 such as, but not limited to, a Natural Language Processing (NLP) model, to analyze the textual content and detect plurality of scenes based on the classification data 224. The ML model 226 may be pretrained using a plurality of training datasets from the training database 108 to classify a plurality of portions of the immutable video 110 into at least the plurality of genres based on the textual content and detect the scenes based on the classification. The plurality of training datasets may comprise input training data and output training data. The input training data may include, but not limited to, textual content of the one or more genres, the one or more sub-genres, the one or more keywords and the one or more tags and the output training data may include, but not limited to, corresponding one or more genres, one or more sub-genres, one or more keywords and one or more tags. During training, the ML model 226 may receive the input training data and generate predictions of classification data 224 of each scene. Further, the generated predictions of the ML model 226 may be compared with the output training data to evaluate a loss function. The ML model 226 may be trained with each dataset of the input training data until the loss function is reduced to a minimum value, such as, but not limited to, ‘O’.

[0052] The ML model 226 may receive textual content as input and process the textual content to extract one or more features from the textual content. The one or more features may represent hidden patterns of the textual content that may be useful to predict the classification data 224 of various portions of the immutable video 110 belonging to the plurality of genres. The ML model 226 may predict the classification data 224 of a plurality of portions of the immutable video 110 based on the one or more features. The ML model 226 may identify one or more portions of the immutable video 110 comprising same or similar classification data 224 as a scene. For example, the ML model 226 may identify a 5 minutes portion of the immutable video 110 classified as ‘drama’, ‘action’ genres, and ‘emotional drama’ sub-genre as a scene.

[0053] In other embodiments, the genre classification module 212 may extract the textual content of the immutable video 110 from the playlist file 218 or the video server 106 and may compare the textual content with pre-stored textual content of each of the plurality of genres. Further, the genre classification module 212 may evaluate a similarity between the textual content of the immutable video 110 with the pre-stored textual content based on the comparison. If the textual content of a portion, for example, 10 minutes, of the immutable video 110 has high similarity, for example, when compared to a threshold, with a genre or a sub-genre, the genre classification module 212 may detect the portion as a scene. Thus, the genre classification module 212 may detect the plurality of scenes based on the similarity. The classification data 224 may comprise the plurality of scenes and / or shots and / or portions of the immutable video 110 of a particular genre / sub-genre / key word / tag with the corresponding start times and end times. In one example, the classification data 224 may be stored in a JavaScript Object Notation (JSON) format as indicated below.

[0054] “Comedy”: [

[0055] “start_time”: “00:02:27,958”, “end_time”: “00:02:39,790”

[0056] “start_time”: “00:03:22,208”, “end_time”: “00:03:43,332”

[0057] “start_time”: “00:05:06,833”, “end_time”: “00:05:48,999”

[0058] “start_time”: “00:06:02,416”, “end_time”: “00:06:20,999” }

[0059] The short scene identification module 214 may analyze the classification data 224 for each of the scenes to identify a plurality of short scenes 228 from the immutable video 110. The short scene 228 may be defined as an attentive portion of the immutable video 110 that may be related to a particular genre / sub-genre / keyword / tag and may be of interest to a plurality of users, for example, the user 105. The short scene may be a portion of a scene and may be of a short time period. In one example, the short scene may span from a few seconds to a few minutes, such as, but not limited to, 30 seconds to 2 minutes.

[0060] The short scene identification module 214 may identify the plurality of short scenes 228 by analyzing one or more of an audio content, a video content and the classification data 224 of the plurality of shots and / or scenes of the immutable video 110. The audio content may include, but not limited to, background music, songs, and the like within the shot or the scene. The video content may include, but not limited to, the plurality of image frames. The short scene identification module 214 may be configured to extract key features from the audio content and the video content for example using an ML model. In an example, the short scene identification module 214 may use Mel Frequency Cepstral Coefficients (MFCC) to extract features from the audio content of sequential shots in a scene or a story. The short scene identification module 214 may combine the key features and the corresponding classification data 224 to identify the plurality of short scenes 228. The short scene identification module 214 may combine two consecutive scenes together when the two consecutive scenes are from same genre to bring continuity of the video. For example, a series of image frames with ‘dance’ combined with ‘Salsa’ sub-genre may be identified as a short scene. In an embodiment, the short scene identification module 214 may be configured with an Al model to combine the key features and the corresponding classification data 224.

[0061] In some embodiments, the short scene identification module 214 may identify the plurality of short scenes 228 based on only one of the audio content, the video content and the classification data 224. For example, a short scene 228 with firing of bullets may be identified based on the audio content. In another example, a short scene 228 visualizing celebration of a festival may be identified based on the video content. In a further example, a short scene 228 with ‘rainy weather’ and ‘heavy traffic’ may be identified based on the one or more keywords of the classification data 224.

[0062] Further, the short scene identification module 214 may determine metadata 230 of each short scene 228. The metadata 230 may be defined as additional data related to each short scene 228 that may be later used by the dynamic playlist generation module 216 to retrieve the audio and video content of each short scene 228. The metadata 230 of the short scene 228 may comprise at least an identifier of the immutable video 110 comprising the short scene 228, an identifier of the short scene 228, the classification data 224 associated with the short scene 228, and a resource link to the immutable video 110. The identifier of the immutable video 110 may be defined as a string of alphanumeric characters to uniquely identify the immutable video 110. The identifier of the short scene 228 may be defined as a string of alphanumeric characters to uniquely identify the short scene 228 from the plurality of short scenes 228 of the immutable video 110. The resource link may be defined as a path to access the immutable video 110. For example, the resource link may be a Uniform Resource Locator (URL) to a website of the video server 106 having rights to render the immutable video 110.

[0063] Further, the metadata 230 may also comprise a timestamp, a title of the immutable video 110 and an identifier or a name of an entity streaming the immutable video 110. The timestamp may comprise at least a start time and an end time of the short scene 228. The start time may be a time instant at which the short scene 228 starts in the immutable video 110 and the end time may be a time instant at which the short scene 228 ends in the immutable video 110. In some embodiments, the timestamp may comprise a start time and a duration of the short scene 228. The duration may be a difference between the start time and the end time indicating a duration of the short scene 228. In other embodiments, the timestamp may comprise the start time, the end time and the duration of the short scene 228. Similarly, the short scene identification module 214 may determine the metadata 230 of the plurality of short scenes 228 of the immutable video 110 and the plurality of short scenes 228 of the plurality of immutable videos 110.

[0064] Further, the short scene identification module 214 may store the metadata 230 of the plurality of short scenes 228 in the memory 204 and / or in the metadata database 109. The metadata 230 may be stored in any form of data such as, but not limited to, graph structure, a table or any other format. For example, the metadata 230 of the immutable video 110 may be stored in JSON format as indicated below. “start_time”: “00:07:30”, “end_time”: “00:07:40”

[0065] }

[0066] ],

[0067] “Love”: [

[0068] “start_time”: “00:09:38”,

[0069] “end_time”: “00:09:58”

[0070] }

[0071] ],

[0072] “Emotion_Drama”: [

[0073] “start_time”: “00: 12:32”,

[0074] “end_time”: “00: 12:52”

[0075] }

[0076] ],

[0077] “Debate”: [

[0078] “start_time”: “00:22:41”,

[0079] “end_time”: “00:22:60”

[0080] }

[0081] ]

[0082] }

[0083] In another example, the metadata 230 may be stored in a table format as shown in Table 1.

[0084] Table 1

[0085] As shown in Table 1, the metadata 230 may comprise a table of data including the short scene identifier, the immutable video identifier, the plurality of genres, a video streaming channel, title of the immutable video, the start time, the end time and the duration of each short scene.

[0086] The dynamic playlist generation module 216 may be configured to generate the playlist 112 from the metadata 230 based on one or more user preferences and user history. To generate the playlist 112, the dynamic playlist generation module 216 may analyze the one or more user preferences and user history. The one or more user preferences may comprise one or more genres, sub-genres or keywords that have been pre-selected by the user 105. The user history may be defined as a plurality of historical short videos or videos viewed, liked, shared or created by the user 105. The user history may also comprise the classification data 224 associated with the plurality of historical short videos or videos.

[0087] Further, the dynamic playlist generation module 216 may select a set of user-specific scenes from the plurality of short scenes 228 based at least on the one or more user preferences and the user history of the user 105. The set of user-specific scenes may comprise one or more short scenes from the plurality of short scenes 228 that may be of interest to the user 105 based on previously viewed short videos and / or user preferences. For example, the set of user-specific scenes may comprise few short scenes of ‘thriller’ genre from the year range 1970-1980.

[0088] Further, the dynamic playlist generation module 216 may retrieve the metadata 230 of each userspecific scene from the set of the user-specific scenes from the memory 204 or the metadata database 109. The dynamic playlist generation module 216 may identify the timestamp of the short scene 228 and the resource link of the immutable video 110 from the metadata 230. Further, the dynamic playlist generation module 216 may retrieve the video content of the immutable video 110 at the timestamp from the resource link. The dynamic playlist generation module 216 may generate a short video from the retrieved video content by modifying the retrieved video content based on one or more dimensions of the user interface 114 of the user device 104. The dimensions may vary for each type of the user device 104. For example, the short video may comprise a selected portion of pixels generating a vertical video version to render on a smart phone.

[0089] Further, the dynamic playlist generation module 216 may generate a plurality of short videos for each of the set of user- specific scenes of the immutable video 110 and may generate the playlist 112 based on the plurality of short videos. Similarly, the dynamic playlist generation module 216 may generate another plurality of short videos from the plurality of immutable videos 110 and may aggregate the another plurality of short videos to the playlist 112. The aggregated playlist 112 may comprise a random arrangement or an ordered arrangement of the short videos. The ordered arrangement may be based on one or more of the duration, the genre, the sub-genre, the keywords, the tags associated with each shot video and / or based on the user interests. Further, as the user views each short video from the aggregated playlist 112, the dynamic playlist generation module 216 may monitor the viewing behavior of the user 105. The viewing behaviour may include, but is not limited to, a viewing time for which the user 105 has watched each short video of the aggregated playlist 112, a number of short videos watched of a particular genre or sub-genre has watched, genres or sub-genres of short videos liked, shared and / or commented by the user 105.

[0090] Based on the viewing behavior, the dynamic playlist generation module 216 may dynamically modify one or more parameters of the aggregated playlist 112 in real-time. The one or more parameters may comprise any parameters of the aggregated playlist 112 that may be modified to render various short videos that may be aligned with the viewing behavior. The one or more parameters may comprise, but not limited to, a number of short videos of a particular genre or a sub-genre, a number of short videos associated with particular keywords or tags, a duration of each short video and the like, including new short videos and deleting existing short videos and the like. For example, if the user 105 skips ‘action’ short videos continuously without watching complete short video, other ‘action’ short videos from the aggregated playlist 112 may be deleted and short videos from other genres may be added. In an embodiment, the SVGS 102 may be configured with Long Term Short Memory (LSTM) model to perform one or more functions described herein to generate the playlist 112 of the short videos.

[0091] Thus, the SVGS 102 renders various short videos to the user 105 based on user preferences and user interests irrespective of trending or rating of the short videos, thereby providing wide diversity of videos in each genre which may not be retrieved due to overrated videos. Further, the SVGS 102 also optimizes memory resources required to generate the playlist 112 as only the metadata of the short scenes is stored in the metadata database 109 in contrast to storing the audio and video content of the short scenes. In addition, as the SVGS 102 retrieves the video or audio content of the short videos directly from the video server 106, the SVGS 102 avoids the requirement of editing or clipping the immutable videos 110 to generate the playlist 112, thereby optimizing usage of computational resources such as processing power. This may further enable various content aggregators who may not have access to the immutable videos 110 to generate versatile playlists 112 based on user interests. Moreover, by dynamically fine-tuning the parameters of the playlist 112 based on the viewing behavior, the SVGS 102 provides a personalized and optimized playlist to the user 105, boosting user relevancy of the short videos and minimizing buffering of the short videos, thereby enhancing user experience. Accordingly, the SVGS 102 may thereby leverage cost-effective solutions by increasing a number of users 105 due to personalized playlists 112.

[0092] Fig. 3 illustrates a flowchart of a method 300 to generate the playlist 112 of short videos from the immutable video 110, in accordance with an embodiment of the present disclosure.

[0093] At block 302, the SVGS 102 may process the immutable video 110 to detect a plurality of shots within the immutable video 110.

[0094] At block 304, the SVGS 102 may process the immutable video 110 to detect and classify a plurality of scenes into at least a plurality of genres.

[0095] At block 306, the SVGS 102 may identify a plurality of short scenes 228 by combining classification data 224 associated with the plurality of scenes and the plurality of shots. At block 308, the SVGS 102 may store a metadata related to each short scene of the plurality of short scenes 228 in a metadata database 109, wherein the metadata 230 comprises at least a timestamp of the short scene 228.

[0096] At block 310, the SVGS 102 may generate the playlist of short videos using the metadata 230 without editing the immutable video 110.

[0097] At block 312, the SVGS 102 may render the playlist 112 to the user 105 on the user device 104.

[0098] At block 314, the SVGS 102 monitors viewing behavior of the user 105 and may proceed to block 310 to modify the playlist 112 based on the viewing behavior of the user 105.

[0099] The blocks 312 and 314 of the method 300 may be optional based on the viewing behavior of the user 105.

[0100] The method 300 may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform specific functions or implement specific abstract data types. The order in which the method 300 is described is not intended to be construed as a limitation, and any number of the method blocks described can be combined in any order to implement the methods. Additionally, individual blocks may be deleted from the methods without departing from the scope of the subject matter described herein. Furthermore, the method 300 can be implemented in any suitable hardware, software, firmware, or combination thereof.

[0101] Fig. 4 illustrates a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.

[0102] In an embodiment, the computer system 400 may be the SVGS 102 for generating the playlist 112 of short videos of the plurality of immutable videos 110 without editing or clipping the plurality of immutable videos 110. The computer system 400 may include a central processing unit (“CPU” or “processor”) 408. The processor 408 may comprise at least one data processor for executing program components for executing user or system-generated business processes. The processor 408 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc.

[0103] The processor 408 may be disposed in communication with one or more input / output (I / O) devices 402 and 404 via I / O interface 406. The I / O interface 406 may employ communication protocols / methods such as, without limitation, audio, analog, digital, stereo, IEEE- 1594, serial bus, Universal Serial Bus (USB), infrared, PS / 2, BNC, coaxial, component, composite, Digital Visual Interface (DVI), high-definition multimedia interface (HDMI), Radio Frequency (RF) antennas, S-Video, Video Graphics Array (VGA), IEEE 8O2.n / b / g / n / x, Bluetooth, cellular (e.g., Code- Division Multiple Access (CDMA), High-Speed Packet Access (HSPA+), Global System For Mobile Communications (GSM), Long-Term Evolution (LTE) or the like), etc.

[0104] Using the I / O interface 406, the computer system 400 may communicate with one or more I / O devices 402 and 404. In some implementations, the processor 408 may be disposed in communication with a communication network 409 via a network interface 410. The network interface 410 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), Transmission Control Protocol / Internet Protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc. Using the network interface 410 and the communication network 409, the computer system 400 may be connected to the training database 108, the metadata database 109, the user device 104 and the plurality of video servers 106.

[0105] The communication network 409 can be implemented as one of several types of networks, such as intranet or any such wireless network interfaces. The communication network 409 may either be a dedicated network or a shared network, which represents an association of several types of networks that use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol / Internet Protocol (TCP / IP), Wireless Application Protocol (WAP), etc., to communicate with each other. Further, the communication network 409 may include a variety of network devices, including routers, bridges, servers, computing devices, storage devices, etc. In some embodiments, the communication network 409 may be one or more of a V2V network, a V2C network and the like. In some embodiments, the processor 408 may be disposed in communication with a memory 430 e.g., RAM 414, and ROM 416, etc. as shown in Fig. 4, via a storage interface 412. The storage interface 412 may connect to memory 430 including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as Serial Advanced Technology Attachment (SATA), Integrated Drive Electronics (IDE), IEEE- 1594, Universal Serial Bus (USB), fiber channel, Small Computer Systems Interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory devices, solid-state drives, etc.

[0106] The memory 430 may store a collection of program or database components, including, without limitation, user / application 418, an operating system 428, a web browser 424, a mail client 420, a mail server 422, a user interface 426, and the like. In some embodiments, computer system 400 may store user / application data 418, such as the data, variables, records, etc. as described in this invention. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase.

[0107] The operating system 428 may facilitate resource management and operation of the computer system 400. Examples of operating systems include, without limitation, Apple Macintosh ™ OS X ™, UNIX ™, Unix-like system distributions (e.g., Berkeley Software Distribution (BSD), FreeBSD ™, Net BSD ™, Open BSD ™, etc.), Linux distributions (e.g., Red Hat ™, Ubuntu ™, K-Ubuntu ™, etc.), International Business Machines (IBM ™) OS / 2 ™, Microsoft Windows ™ (XP™, Vista / 7 / 8, etc.), Apple iOS ™, Google Android™, Blackberry™ Operating System (OS), or the like. A user interface may facilitate display, execution, interaction, manipulation, or operation of program components through textual or graphical facilities. For example, user interfaces may provide computer interaction interface elements on a display system operatively connected to the computer system 400, such as cursors, icons, check boxes, menus, windows, widgets, etc. Graphical User Interfaces (GUIs) may be employed, including, without limitation, Apple ™ Macintosh ™ operating systems’ Aqua ™, IBM ™ OS / 2 ™, Microsoft ™ Windows ™ (e.g., Aero, Metro, etc.), Unix X-Windows ™, web interface libraries (e.g., ActiveX, Java, JavaScript, AJAX, HTML, Adobe Flash, etc.), or the like. The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope of the disclosed embodiments.

[0108] Also, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be equivalent in meaning and be open ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items or meant to be limited to only the listed item or items. It must also be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. Accordingly, the disclosure of the embodiments of the disclosure is intended to be illustrative, but not limiting, of the scope of the disclosure. With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.

Claims

We claim:

1. A method of generating a playlist of short videos from an immutable video, the method comprising: processing the immutable video to detect a plurality of shots within the immutable video; processing the immutable video to detect and classify a plurality of scenes into at least a plurality of genres; identifying a plurality of short scenes by combining classification data associated with the plurality of scenes and the plurality of shots; storing a metadata related to each short scene of the plurality of short scenes in a metadata database, wherein the metadata comprises at least a timestamp of the short scene; and generating the playlist of short videos using the metadata without editing the immutable video.

2. The method of claim 1 , wherein processing the immutable video to detect and classify the plurality of scenes into the plurality of genres comprises: extracting a textual content of the immutable video from an immutable video database; comparing the textual content with a textual content related to each of the plurality of genres; detecting a plurality of scenes with the textual content similar to the textual content related to one or more of the plurality of genres; and classifying each of the plurality of scenes into one or more of the plurality of genres based on similarity.

3. The method of claim 1, wherein processing the immutable video to detect and classify the plurality of scenes into the plurality of genres comprises:extracting a textual content of the immutable video from an immutable video database or a playlist file; analyzing the textual content using a Machine Learning (ML) model; and classifying, by the ML model, each of the plurality of scenes into one or more of the plurality of genres based on the analysis.

4. The method of claim 2 or claim 3, further comprises: classifying each scene of the plurality of scenes into one or more sub-genres, one or more keywords and one or more tags.

5. The method of claim 1, wherein identifying the plurality of short scenes comprises: analyzing one or more of an audio content and a video content of each of the plurality of shots; and identifying the plurality of short scenes based on the analysis and the corresponding classification data.

6. The method of claim 1 , wherein the metadata further comprises at least an identifier of the immutable video, an identifier of the short scene, the classification data associated with the short scene, and a resource link to the immutable video.

7. The method of claim 1, wherein generating the playlist of short videos using the metadata comprises: selecting a set of user-specific scenes from the plurality of short scenes based at least on one or more user preferences and user history of a user; retrieving a video content of each user-specific scene of the set of user-specific scenes from the immutable video using the metadata; generating a short video by modifying one or more visual dimensions of the video content based on one or more dimensions of a user interface of a user device, wherein the user device is associated with the user; andgenerating the playlist of the short videos as a plurality of short videos corresponding to the set of user-specific scenes.

8. The method of claim 7, wherein retrieving the video content of each user-specific scene of the set of user-specific scenes from the immutable video using the metadata comprises: retrieving the timestamp and a resource link to the immutable video associated with the user-specific scene from the database, wherein the timestamp comprises at least a start time and an end time of the user-specific scene within the immutable video; and retrieving the video content of the immutable video at the timestamp using the resource link.

9. The method of claim 1, further comprising: aggregating another plurality of short videos generated from a plurality of immutable videos to the playlist of short videos; rendering the aggregated playlist of short videos on a user device of a user; monitoring a viewing behavior of the user during rendering one or more short videos of the aggregated playlist, wherein the viewing behavior includes at least a viewing time of the one or more short videos; and dynamically modifying one or more parameters of the aggregated playlist based on the viewing behavior.

10. An apparatus to generate a playlist of short videos from an immutable video, the apparatus comprising: a memory; and a processor communicatively coupled with the memory, wherein the processor is configured to: process the immutable video to detect a plurality of shots within the immutable video; process the immutable video to detect and classify a plurality of scenes into at least a plurality of genres;identify a plurality of short scenes by combining classification data associated with the plurality of scenes and the plurality of shots; store a metadata related to each short scene of the plurality of short scenes in a metadata database, wherein the metadata comprises at least a timestamp of the short scene; and generate the playlist of short videos using the metadata without editing the immutable video.

11. The apparatus of claim 10, wherein to process the immutable video to detect and classify the plurality of scenes into the plurality of genres, the processor is configured to: extract a textual content of the immutable video from an immutable video database; compare the textual content with a textual content related to each of the plurality of genres; detect a plurality of scenes with the textual content similar to the textual content related to one or more of the plurality of genres; and classify each of the plurality of scenes into one or more of the plurality of genres based on similarity.

12. The apparatus of claim 10, wherein to process the immutable video to detect and classify the plurality of scenes into the plurality of genres, the processor is configured to: extract a textual content of the immutable video from an immutable video database or a playlist file; analyze the textual content using a Machine Learning (ML) model; and classify, by the ML model, each of the plurality of scenes into one or more of the plurality of genres based on the analysis.

13. The apparatus of claim 11 or claim 12, wherein the processor is further configured to: classify each scene of the plurality of scenes into one or more sub-genres based on one or more keywords and one or more tags of the scene.

14. The apparatus of claim 10, wherein to identify the plurality of short scenes by processing the plurality of scenes, the processor is further configured to: analyze one or more of an audio content and a video content of each of the plurality of shots; and identify the plurality of short scenes based on the analysis and the corresponding classification data.

15. The apparatus of claim 10, wherein the metadata further comprises at least an identifier of the immutable video, an identifier of the short scene, the classification data associated with the short scene, and a resource link to the immutable video.

16. The apparatus of claim 10, wherein to generate the playlist of short videos using the metadata, the processor is configured to: select a set of user-specific scenes from the plurality of short scenes based at least on one or more user preferences and user history of a user; retrieve a video content of each user-specific scene of the set of user-specific scenes from the immutable video using the metadata; generate a short video by modifying one or more visual dimensions of the video content based on one or more dimensions of a user interface of a user device, wherein the user device is associated with the user; and generate the playlist of the short videos as a plurality of short videos corresponding to the set of user-specific scenes.

17. The apparatus of claim 10, wherein to retrieve the video content of each user-specific scene of the set of user-specific scenes from the immutable video using the metadata, the processor is configured to: retrieve the timestamp and a resource link to the immutable video associated with the user-specific scene from the database, wherein the timestamp comprises at least a start time and an end time of the user-specific scene within the immutable video; andretrieve the video content of the immutable video at the timestamp using the resource link.

18. The apparatus of claim 10, wherein the processor is further configured to: aggregate another plurality of short videos generated from a plurality of immutable videos to the playlist of short videos; render the aggregated playlist of short videos on a user device of a user; monitor a viewing behavior of the user during rendering one or more short videos of the aggregated playlist, wherein the viewing behavior includes at least a viewing time of the one or more short videos; and dynamically modify one or more parameters of the aggregated playlist based on the viewing behavior.

19. A non-transitory computer-readable medium having program instructions stored thereon, when executed by a short video generation system, facilitate the short video generation system for generating a playlist of short videos from an immutable video by performing operations comprising: processing the immutable video to detect a plurality of shots within the immutable video; processing the immutable video to detect and classify a plurality of scenes into at least a plurality of genres; identifying a plurality of short scenes by combining classification data associated with the plurality of scenes and the plurality of shots; storing a metadata related to each short scene of the plurality of short scenes in a metadata database, wherein the metadata comprises at least a timestamp of the short scene; and generating the playlist of short videos using the metadata without editing the immutable video.

20. The non-transitory computer-readable medium of claim 19, wherein the program instructions configured to process the immutable video to detect and classify the plurality of scenes into the plurality of genres facilitate: extracting a textual content of the immutable video from an immutable video database; comparing the textual content with a textual content related to each of the plurality of genres; detecting a plurality of scenes with the textual content similar to the textual content related to one or more of the plurality of genres; and classifying each of the plurality of scenes into one or more of the plurality of genres based on similarity.

21. The non-transitory computer-readable medium of claim 19, wherein the program instructions configured to process the immutable video to detect and classify the plurality of scenes into the plurality of genres facilitate: extracting a textual content of the immutable video from an immutable video database or a playlist file; analyzing the textual content using a Machine Learning (ML) model; and classifying, by the ML model, each of the plurality of scenes into one or more of the plurality of genres based on the analysis.

22. The non-transitory computer-readable medium of claim 20 or claim 21, wherein the program instructions configured to further facilitate: classifying each scene of the plurality of scenes into one or more sub-genres, one or more keywords and one or more tags.

23. The non-transitory computer-readable medium of claim 19, wherein the program instructions configured to identify the plurality of short scenes facilitate: analyzing one or more of an audio content and a video content of each of the plurality of shots; andidentifying the plurality of short scenes based on the analysis and the corresponding classification data.

24. The non-transitory computer-readable medium of claim 19, wherein the metadata further comprises at least an identifier of the immutable video, an identifier of the short scene, the classification data associated with the short scene, and a resource link to the immutable video.

25. The non-transitory computer-readable medium of claim 19, wherein the program instructions configured to generate the playlist of short videos using the metadata facilitate: selecting a set of user-specific scenes from the plurality of short scenes based at least on one or more user preferences and user history of a user; retrieving a video content of each user-specific scene of the set of user-specific scenes from the immutable video using the metadata; generating a short video by modifying one or more visual dimensions of the video content based on one or more dimensions of a user interface of a user device, wherein the user device is associated with the user; and generating the playlist of the short videos as a plurality of short videos corresponding to the set of user-specific scenes.

26. The non-transitory computer-readable medium of claim 25, wherein the program instructions configured to retrieve the video content of each user-specific scene of the set of user- specific scenes from the immutable video using the metadata facilitate: retrieving the timestamp and a resource link to the immutable video associated with the user-specific scene from the database, wherein the timestamp comprises at least a start time and an end time of the user-specific scene within the immutable video; and retrieving the video content of the immutable video at the timestamp using the resource link.

27. The non-transitory computer-readable medium of claim 19, wherein the program instructions configured to facilitate: aggregating another plurality of short videos generated from a plurality of immutable videos to the playlist of short videos; rendering the aggregated playlist of short videos on a user device of a user; monitoring a viewing behavior of the user during rendering one or more short videos of the aggregated playlist, wherein the viewing behavior includes at least a viewing time of the one or more short videos; and dynamically modifying one or more parameters of the aggregated playlist based on the viewing behavior.

Citation Information

Patent Citations

  • Systems and methods for providing video segments

    US10645468B1

  • Detection of demarcating segments in video

    US20150363648A1

Cited By

  • Determining topic chapters for digital videos utilizing video segmentation machine learning models

    US12652444B2