System and method for generating custom video content
The AI-powered system addresses the challenges of generating custom short-form video content by automating the process, reducing time and costs, and enhancing scalability and quality, enabling efficient production of premium video content for digital platforms.
Patent Information
- Application Number
- PCT/IN2024/052285
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-11-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing methods for generating custom short-form video content are time-consuming, costly, and require skilled operators, making it difficult to scale and meet the demand for premium video content on digital platforms.
A system and method utilizing artificial intelligence and computer vision to automatically generate custom video content from a source video, including transcribing audio to text, detecting regions of interest, and processing with a large language model to create high-quality, short-form videos in various formats and resolutions.
Enables efficient and cost-effective generation of custom video content, reducing production time by up to one tenth compared to manual processes, and allowing for the creation of multiple formats from a single video with a single click, enhancing scalability and quality.
Smart Images

Figure IN2024052285_30052025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR GENERATING CUSTOM VIDEO CONTENTTECHNICAL FIELDThe present invention relates, generally, to generating content files and more specifically to an artificial-intelligence based system and method for efficiently generating short-form custom video content in multiple resolutions including but not limited to 9:16, 4:5 etc along with meta data and other assets from a video source, which can be filed based or live and publishing to multiple platforms with a single click.BACKGROUND OF THE INVENTION
[0001] The delivery of video content in today’s time through sources such as the Internet on various content consumption mediums including handheld devices, tablet computers etc. and the rapidly decreasing attention span of users requires the video content to be special created and curated for such digital consumption of video content. Most of the content created today is long and in 16:9 ratio while a majority of the consumption has moved to short-form video in 9: 16 due to the mobile phone becoming the platform of choice for content consumption. While there is a great demand for custom video content, the original source videos does not get monetises as much as it should due to the low engagement given the length of the video and digital-readiness of the video.
[0002] Generating such custom video content from long form video requires inter alia editing and designing of various components of the long form video content such as images, text, audio, animation etc. on a timeline and incorporating further visual and audio effects. Said editing and designing typically requires experts skilled in using video editing tools. The generation of custom video content in said conventional way not only requires skilled and highly expensive operators, but also consumes a lot of time, especially when the demand is for continuous stream of custom video content. That is, manual creation of custom videos may require a team including an editorial person, a video editor, a visual designer, a search engine optimization expert and a social media expert to generate a video as per the content and form requirement of another editorialperson. Requiring a team to generate custom video content is, therefore, time and cost intensive and may also further lack standardization of quality. And for these reasons it is difficult to scale the current process whereas the need is to generate premium short-form video content at volume to get which is more skewed towards digital platforms.
[0003] For instance, existing platforms to identify and generate clips of interest from a long form video are bulky and tedious for an ordinary user since it requires a high-level video processing knowledge for producing custom video content with higher quality along with the content editors / producer. There is a skill friction as the operators needed to use such enterprise grade platforms are expensive and in short supply.
[0004] In view of the above shortcomings, there arises a need for a system and method which enables generation of custom short-form video content in vertical format especially video content creation process that is easy to scale and allows generating video in a simple, quick, computationally efficient and user-friendly manner while eliminating the need for skilled operators and also simplifying the current complex workflowsSUMMMARY
[0005] This section is intended to introduce certain objects of the disclosed methods and systems in a simplified form and is not intended to identify the key advantages or features of the present disclosure.
[0006] In one embodiment, a method for generating custom video content is disclosed. The method comprises receiving a source video, wherein the source video contains a plurality of frames, transcribing an audio of the source video to text to generate transcribed text, wherein the transcribed text corresponds to the timestamp of the plurality of frames in the source video, detecting regions of interest like objects, labels, text on frames corresponding to the timestamp of the plurality of frames in the source video using computer vision, wherein regions of interest is detected to generate visual meta data, receiving a request to generate the custom video content and processing the source video. The processing of the source video comprises feeding the transcribed text to a large language model, feeding the visual meta data to the large language model,identifying the type of content in the source video based on an output from the large language model, extracting relevant frames corresponding based on received request, type of content and the output from the large language model, and generating the custom video content based on the extracted relevant frames. The method further comprises transforming the generated custom video content based on the detected region of interest in the source video.
[0007] In another embodiment, an apparatus for generating custom video content is disclosed. The apparatus comprises a reading module configured to receiving a source video, wherein the source video contains a plurality of frames and a processor configured to transcribe an audio of the source video to text to generate transcribed text, wherein the transcribed text corresponds to the timestamp of the plurality of frames in the source video, detect regions of interest like objects, labels, text on frames corresponding to the timestamp of the plurality of frames in the source video using computer vision, wherein regions of interest is detected to generate visual meta data, receive a request to generate the custom video content, process the source video. The processing of the source video comprises feed the transcribed text to a large language model, feed the visual meta data to a large language model, identify the type of content in the source video based on an output from the large language model, extract relevant frames corresponding based on received request, type of content and the output from the large language model, and generate the custom video content based on the extracted relevant frames. The processor is further configured to transform the generated custom video content based on the detected region of interest in the source video.
[0008] Other general and specific objectives of the invention will in part be obvious and will in part appear hereinafter.BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 illustrates an overview of the technology stack of a system, in accordance with an embodiment of the present disclosure.
[0010] Figure 2 illustrates an overview of an exemplary architecture or network environment in which the system in accordance with the present disclosure is operating.[Oil] Figure 3 illustrates an exemplary flowchart of a method for generating custom video form video content, in accordance with an exemplary embodiment of the present invention.
[0012] Figure 4a, 4b, 4c, 4d, 4e, 4f, 4g illustrate exemplary graphical user interfaces of a system in accordance with an exemplary embodiment of the present invention.
[0013] Figure 5 illustrates a flowchart of a method for generating custom video content, in accordance with an exemplary embodiment of the present invention.
[0014] Figure 6 is a block diagram illustrating an exemplary computing device in which one or more embodiments of the present invention may operate, according to an embodiment.DETAILED DESCRIPTION
[0015] In the following description, for the purposes of explanation, numerous specific details have been set forth in order to provide a description of the invention. It will be apparent, however, that the invention may be practiced without these specific details and features.
[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of the ordinary skill in the art to which this invention belongs. The materials, methods, and examples provided herein are illustrative only and not intended to be limiting.
[0017] The present embodiments comprise a system and method for automatic generation of high-quality custom video content in vertical format especially custom video content from long form video content with reduced time, cost and complexity. In one embodiment, the custom video content may include, but not limited to, reels, shorts,video content of duration shorter than the duration of the source video, video content having aspect ratio different than the source video, etc. This also includes the generation of customised rich meta-data for every content asset that is created for different social media platforms.
[0018] The custom video(s) may be generated by the present system with one click, and each generated custom video includes all the required information including detailed meta data, searchable transcription, multiple thumbnail options and tags . In an exemplary embodiment, the custom video can be dynamic and interactive. The present system further encompasses generating a custom video content by enabling editing the transcript of the input video directly. This allows any user to create short-form video without needing expensive editors. The other important aspect is the ability of the platform to generate the custom video content from the original source of video in different aspect ratios without needing any other software or expensive personnel. The current manual process is expensive and time consuming and this is solved by the platform. In this application, custom video content and short form video have been used interchangeably.
[0019] An objective of the present disclosure is to provide improved techniques for overcoming the drawbacks of the existing platforms as discussed in the background section. An objective of the present disclosure is to provide an end-to-end platform that creates premium custom video content in different aspect ratios from any existing video content efficiently and without requiring skilled video editing professionals and human editors. Another objective of the present disclosure is to generate a platform that can automatically convert a 16:9 video into 9:16 and other aspect ratios using computer vision technology without requiring specialized staff and high-end expensive software. A further objective of the present disclosure is to enable a user to automatically create multiple custom video content in various formats from a single video with a single click with high quality and reduced time and cost. It is a further objective of the present disclosure to provide an in-platform suite of editing tools which allow editing the video by editing the transcript which otherwise requires highly-skilled and expensive editors and time, e.g. the current manual process takes approximately 45-60 minutes to create ashort-form video from a 30-miute original source video. The time taken to generate a custom video content in accordance with present disclosure may be reduced by upto one tenth of the amount taken by existing manual processes . Yet another objective of the present disclosure is to provide a platform that is able to generate metadata and relevant information from a long video for generation of a custom video content without relying on external inputs on the type and kind of video such as an interview or a speech. This makes the video highly discoverable and hence increased engagement and revenue. Yet another objective of the present disclosure is to address the massive asymmetry between the demand for content and the supply of content, and the inconsistency in the standard of quality of the generated content further aiding in discoverability of the generated content.
[0020] Referring to FIG. 1, an overview of a sequence diagram 100 illustrating generation of the custom video content is shown, in accordance with one embodiment of the present invention. A publisher 102 may upload a source video into a system running in accordance with the present disclosure. The source video is made up of a plurality of frames combined together. The source video may be a stored video file or may be a live stream video in Real-Time Messaging Protocol (RTMP) or HTTP live streaming (HLS) format.
[0021] Once the source video has been uploaded, the next step involves analyzing 104 the input source video. The step of analyzing involves the steps of transcribing the audio of the source video, detecting regions of interest on frames, processing the source video to generate custom video content and transforming the generated custom video content. Each of these steps will now be explained below.
[0022] As explained above, the source video contains a plurality of frames. The audio present in the source video is transcribed to text to generate transcribed text. The transcribed text as generated corresponds to the timestamp of the plurality of frames present in the source video. In other words, the transcribed text is linked with the time at which the text appears in the source video. For example, if person A speaks from minute 1 to minute 2 and person B speaks from minute 2 to minute 3 then the transcribed text corresponding to person A will denote minute 1 to minute 2 and transcribed textcorresponding to person B will denote minute 2 to minute 3. We also keep time codes for every word so that we can create short clips which can be accurate word-level. These can also be used to remove filler words (like ahs and umms) or remove silences between words. This helps with reducing the length of the short video.
[0023] As a next step, visual metadata is generated. The metadata is generated based on the detection of the object in the source video. The visual metadata involves information related to the one or more detected objections. The information may include but not limited to objects, celebs, scene types, logos, shape, size, color, texture, labels, texts, anything which denotes the visual appearance of the detected object.
[0024] The detection of the object in the source video involves using computer vision technology. An object can be detected by visually inspecting the source video on frame by frame basis. Each frame in the source video is pre-processed to identify patterns, boundaries, edges, corners in them. The frames may also be resized for consistency.
[0025] Once the frames are pre-processed, one or more object detection algorithms are applied on the frames. One or more objection detection algorithm may involve but not limited to, Segment Anything, Florence, You Only Look Once (YOLO), Single Shot Detectors (SSD), or Faster Region-based Convolution Neural Network (R-CNN). One or more objection detection algorithms are used to identify objects in each frame which include brands, logos, text, facial expressions and other data-points which will help determine the context of the frame . These detected objects are also tracked across each frame and for all frames to maintain consistency. In one embodiment, post-processing is also performed on the detected objects. The post-processing involves, but is not limited to, filtering out redundant detections, assigning confidence levels to the detections and also co-relating the detections to existing training data etc.
[0026] The next step involves receiving a request to generate the custom video content. The request may be received from a user wanting to generate the custom video content. The user may specify that the user would like to generate shorter formats like key moments or a summary of the video or even breaking the video into theme -basedsegments called chapters or the user may want to transform the aspect ratio of the source video to generate Reels and Shorts for social media plaforms
[0027] Once the request is received, processing is performed on the source video. The processing of the video data involves feeding the transcribed text to the large language model, feeding the visual meta data to the large language model (LLM), identifying the type of content in the source video based on an output from the large language model, extracting relevant frames corresponding based on received request, type of content and the output from the large language model, and generating the custom video content based on the extracted relevant frames.
[0028] In other words, the large language model identifies the type of content in the source video based on the training that has been done using thousands of hours of clean training data. In one embodiment, the LLM model may be trained such as to recognize the words related to a particular type of content. For example, the LLM may recognize the words related to sports (e.g., cricket, player, ball, bat, four, etc) in the source video and may determine that the source video relates to sports related content. As other examples, the type of content of the video source may include, but not limited to, speech, interview, news / bulletin / debate show, and question answer session. The type of content is not limited to the one mentioned here and may include any other types as well.
[0029] The large language model (LLM) is responsible for generating meaningful text in the transcribed text. The LLM may perform unsupervised learning from various language patterns, grammar, facts, reasoning using dataset. The LLM may assign different weights to different words, texts present in the transcribed text by breaking down the text into smaller units. The LLM uses natural language processing techniques to determine semantic relationships between different words present in the transcribed text. Thus, the LLM can easily determine what type of content is present in the source video, the important topics been talked about, the summary of the transcribed text, etc.
[0030] Based on the output from the LLM, the received request and the type of content determined from the transcribed text, the relevant frames are extracted. For example, ifthe user wants to have a short video with news content, the LLM may identify the news related content from the transcribed text and extract the frames related to the news from the source video. In other words, if the content related to the news starts from the minute2 to minute 3 in the source video, then the frame corresponding to the minute 2 to minute3 is extracted (since the transcribed text is associated with the timestamped frames). The custom video content is then generated using the extracted relevant frames.
[0031] After generating the custom video content, the generated custom video content is transformed based on the detected region of interest in the video. The transforming of the custom video content involves, but not limited to, changing the aspect ratio of the source video from horizontal (e.g., 16:9) to vertical (e.g., 9:16, 4:5, 1: 1), adjusting the position of the detected objects, extracting the audio present in the source video, auto-framing of the speaker according to the custom video content, performing image popularity assessment to assess faces, persons, events, vehicles and objects featuring in the input video including determining identity of famous person(s) or a type of action performed by a person, text etc. Based on the transformation of the custom video content, final custom video content is thus generated.
[0032] Thus, step 106 comprises automatically generating multiple relevant shorts / clips / segments / short-form videos out of the input video along with the headlines, meta tags in context of the content featured in the input video, hero images for thumbnails, captions for the input video. Further, a long source video may be converted from a 16:9 video into 9: 16 using computer vision technology including adjusting the position of faces of persons and objects to fit the modified view.
[0033] Finally, the publisher is then enabled to publish the custom video content / clip with a single click on one or more media or platforms in one or more forms. The video content may be distributed by the publisher on multiple platforms including social media without having to access the publishing console of each platform, thus making content publishing faster and more efficient.
[0034] Referring to FIG. 2, an exemplary architecture of a system is illustrated, in accordance with the present disclosure. The architecture in accordance with present disclosure may include a client-server scheme where a system is implemented at a server implemented at a computing device, including various storage, communications, and data processing components. Alternatively, the system is implemented at a cloud-based server. As illustrated in FIG. 2, the system includes a firebase to authenticate a user attempting to accessing the system. Pursuant to authentication, the user is allowed access to the system where the user / publisher may upload a long form input video from which one or more custom video content may be derived. The uploaded videos are processed by the system to generate one or more custom video content in accordance with the user requirements.
[0035] The system to process the input video is an Artificial intelligence and Machine Learning based system using one or more machine learning model based engines including, but not limited to, a narrator, a text to speech or speech to text engine, image processing engine etc. The system thus comprises a model trained to transcribe audio of the input video to text, a model trained to derive meaningful information from the digital images in the video such as important objects and person featuring in the video to be retained when changing the mode of viewing, a model to create meta data from the input video and determine key moments (highlights) for creating summary (that only contains portions of the input video that are determined to be highlights) and a model trained to generate one or more thumbnails for a custom video content using intrinsic image popularity assessment. The training of models may include any data, such as videos and metadata, permitted for use for training, the trained models may comprise one or more model forms or structures that include a type of neural network. The user may then review, edit and publish the custom video content generated by the system for consumption at any video consumption device. The present disclosure applies to all protocols related to video delivery, broadcast and streaming and standard video formats applicable to standard viewing.
[0036] FIG. 2 also includes a backend server for storing and organizing data for various processes relating to summarizing videos, fetching metadata, editing videos etc. The various processes may rely upon one or more large language models (LLMs) hostedwithin the system or at a separate location. The architecture also includes a storage and a database to save all input files, intermittent files, output files and the associated metadata. The one or more machine learning models operated by the machine learning (ML) model based engines may be stored in the database and are accessible and executable by the processor. Said machine learning models can be implemented by one or more components of the system. Further, a load balancer may also be used to manage and / or distribute the traffic across a pool of resources.
[0037] In detail, the system 200 comprises a load balancer 202 for managing traffic across the pool of resources. The system 200 further comprises website 204 where publisher can upload source videos, review, edit and publish custom video content. The source videos may be uploaded from the storage 206. The website 204 may run in an apparatus (for example, in computer 600 as described below). The firebase 208 is used to authorize the valid user / publisher attempting to access the website 204. The firebase 208 is also used to authenticate transcription server to convert the audio / speech present in the source video in textual format. The transcribed text is then used to generate custom video content by extracting relevant frames from the source video as already explained in detail above. The transcribed text is stored in the database 210.
[0038] A backend server 212 is used to perform processes like summarizing videos, fetching meta data, editing on the generated custom video content, etc. As explained above, to generate summarized videos, meta data and editing of custom video, the backend server 212 performs API calls to external services 214 like large language model (LLM). The summarized video content, meta data, edited video may be stored in the storage 206. All assets like meta data is saved in database 216 which is in communication with the database 210.
[0039] Referring to FIG. 3 now, an exemplary flow chart 300 of a method to generate high quality custom video content, is illustrated in accordance with one embodiment of the present disclosure. The method 300 may be performed in accordance with a system implemented by an architecture shown in FIG. 2 or FIG. 6 (below) or any system substantially based on the architecture shown in figures 2 and 6. The method 300 startswith a step 302 of receiving a request for uploading of a source video. The request may comprise additional parameters associated with the type of custom video(s) content required. The additional parameters can, for example include a type of desired format for the custom video content, a category, a time duration limit thereof and the online platforms where the generated custom video content are to be published. On uploading / reading the source video, the source video is processed by the system in accordance with one or more machine learning model based engines to perform an automatic speech to text conversion (at step 304). Also, the method comprises detecting regions of interest to create visual meta data (step 306). Said method further comprises creating one or more thumbnails for custom video content generated from the input video (step 308). Next, the method comprises generating one or more custom video content depending on the type of content identified from the transcript in the source video (step 310) using large language model (LLM). For instance, in case the source video is a news / bulletin / debate show video, metadata may be generated and custom video content for chapters from the input may be generated based on the topics covered in the source video (step 312). Similarly, in case the source video is an interview / speech, a summary may be generated and key moments from the transcribed text of the source video and custom video content based on the generated summary may be generated (steps 314). In case the source video is a question answer session, it is ensured that the question is followed by an answer in the custom video content.
[0040] After the custom video content is generated, the user may set additional attributes such as setting a time limit. Said custom video content is then transformed from a 16:9 format to 9: 16 format (step 316). The transformed custom video content is then stored and presented to the publisher for review (step 318). The system further enables the user to further edit the custom video content before being published on one or more desired platforms. The one or more custom video content, thus, generated are rendered as a video clip on a user interface of any media / website (step 320). Said reading module, is in communication with at least one hardware processor and perform particular tasks or functions.
[0041] The method also encompasses receiving feedback from users on the published custom video content indicating a level of interest in the respective versions of custom video content on different platforms. The feedback may then be analysed and used for training and fine-tuning the Al based models and also future generation and rendering of custom video content without manual intervention to enhance the user experience.
[0042] The method further encompasses automatic conversion of the mode of a generated custom video content from a 16:9 video into a 9: 16 video using computer vision technology without requiring specialized editors.
[0043] As encompassed by the present invention, on reading the source video and receiving a request to generate the custom video content such as a Summary, a Chapter, a Full length package and key moments from the source video, the method processes the input video to generate the custom video content depending on the form selected by the user. Along with the generation of the custom video content in the requested form, the present invention generates the transcript, the metadata, one or more headlines for one or more platforms, one or more options for thumbnails for each of the custom video content. The metadata includes headlines, keywords, descriptions, hashtags, emojis, and summarized content for the input video. The user may then select and edit any custom video content by editing the transcript of said video to generate the desired custom video content to be published. Further, the user may edit the start and stop time and playing time length while editing the transcript of the custom video content. The custom video content, if approved by the user, will be published at the platforms selected by the user while uploading the input video.
[0044] Referring to FIGs. 4a-4g, illustrate exemplary graphical user interfaces for various functions carried out by the system in accordance with the present disclosure. Figures 4a and 4b depict receiving an input long form video by uploading or by entering a URL of the long form input video from a user. Said long form video is processed for generating custom video content. The user at the time of providing the input video also provides selection of the form of the output custom video content required by the user such as one or more of full-length package, key moments, summary and chapters fromthe input video. The user / publisher may then review the uploaded input video. In accordance with the required output type, the system then processes the input video to generate the custom video content such as full-length package, key moments, summary and chapters.
[0045] There shown in figure 4c are various exemplary custom video content generated by the present system including a summary of the input video and chapters from the input video. The user who inputs the video is presented with said custom video content in the desired forms along with the associated transcript such that the user may either publish a custom video content as it is or edit the custom video content by editing the transcript of said video. The user is enabled to select and edit any generated video simply by editing a transcript associated with the selected custom video content.
[0046] Figure 4d illustrates the metadata generated by the system for each of the generated custom video content including one or more headlines meant for one or more platforms and one or more images to be used as thumbnail for the respective custom video content. The system allows the user to edit said one or more headlines for one or more platforms. The system further allows the user to edit / select a thumbnail or alternatively upload an image to be used as thumbnail for the custom video content. The system further allows selecting a template for said custom video content to be published.
[0047] Figure 4e illustrates the custom video content incorporating the options selected in the previous step displayed in a 9:16 format along with the corresponding transcript displayed side-to-side such that the effect of the editing in the transcript is displayed dynamically before publishing the video.
[0048] Figures 4f and 4g depict correction and editing of the transcript associated with the selected custom video content by the user to modify a portion of the custom video content. The changes produced by the editing / correction can be previewed in the video displayed side by side in real-time. The text in the transcript is searchable and modified to generate specific content in the video. The final custom video content may then be reviewed by the user for publication on various platforms and social media sites withouthaving to separately login and / or accessing the dashboard thereof. A video published on a platform may differ in the headline, thumbnail, duration, content, transcript etc. from a corresponding video published on another platform as desired by the user.
[0049] The above-described system and method may be implemented in different manners to automatically generate custom video content / clips / segments in accordance with user’s desired preferences.
[0050] The inventors have discovered that the present system is able to achieve significant increase in speed and volume and cost reduction as compared to existing systems, e.g. the system is able to produce 150% more content in approximately 20% of the time that is currently taken in doing the same tasks by existing system requiring manual assistance.
[0051] Referring to FIG. 5 now, a flowchart of a method 500 illustrating a method for generating custom video content is disclosed, in accordance with one embodiment of the present invention. At step 502, the method comprises receiving a source video, wherein the source video contains a plurality of frames. In other words, the source video is divided into a plurality of frames. At step 504, the method comprises transcribing an audio of the source video to text to generate transcribed text, wherein the transcribed text corresponds to the timestamp of the plurality of frames in the source video.
[0052] At step 506, the method comprises detecting regions of interest like objects, labels, text on frames corresponding to the timestamp of the plurality of frames in the source video using computer vision, wherein regions of interest is detected to generate visual meta data. As explained above, the object detection can be performed by object detection algorithms, such as SAM, Florence, YOLO, SSD and R-CNN. At step 508, the method comprises receiving a request to generate the custom video content. The request may be received from a user wanting to generate reels, shorts, short duration video, key moments, highlights, etc.
[0053] At step 510, the method comprises processing the source video, wherein the processing comprises feeding the transcribed text to a large language model, feeding the visual meta data to the large language model, identifying the type of content in the source video based on an output from the large language model and extracting relevant frames based on received request, type of content and the output from the large language model. As explained above, the large language model (LLM) is responsible for identifying meaningful relationship between the words present in the transcribed text. Based on the analysis of the transcribed text, the LLM can identify the type of content, i.e., whether the text relates to news / bulletin / speech / debate / interview, etc. Also, the LLM is responsible for analyzing the visual meta data, i.e., meta data relating to the visual appearance of the source video. The meta data may include, but not limited to, objects, labels, texts present in the source video.
[0054] At step 512, the method comprises generating custom video content based on the extracted relevant frames. At step 514, the method comprises generating the custom video content based on the extracted relevant frames. Thus, the custom video content will include the frames relevant to generate the custom video content and are extracted after processing of the source video. At step 516, the method comprises transforming the generated custom video content based on the detected region of interest in the source video.
[0055] Referring to FIG. 6 now, an exemplary computing device 600 in which one or more embodiments of the present invention may operate, according to an embodiment. In the system schematic of figure 6, bus 610 is in physical communication with reading module 602, interface 604, memory 606, and processor 608. The memory 606 may contain computer readable instructions which when executed by the processor 608 causes the device 600 to perform the method 500 as provided above in FIG. 5. Bus 610 includes a path that permits components within computing device 600 to communicate with each other. Examples of reading module 602 include peripherals and / or other mechanism that may enable a user to input information to computing device 600, including a keyboard, computer mice, buttons, touch screens, voice recognition, and biometric mechanisms.
[0056] Examples of interface 604 include mechanisms that enable computing device 600 to communicate with other computing devices and / or systems through network connections. Examples of memory 606 include random access memory (RAM), readonly memory (ROM), flash memory, and the like. The memory 606 store information and instructions for execution by processor 608. The processor 408 includes, but not limited to, a microprocessor, an application specific integrated circuit (ASIC), or a field programmable object array (FPOA) and the like. The processor 608 interprets and executes instructions retrieved from memory 606.
[0057] In one embodiment, the computing device 600 may be responsible for implementing the above-mentioned steps. For example, input parameters of the user may be received using the Input / Output device 602. The machine learning models may be stored in the memory 606 and may be implemented by the processor 608.
[0058] In present invention, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium and executed by a hardware -based processor. The implementations described herein are not limited to any specific combinations of hardware circuitry and software.
[0059] In the drawings and specification, there have been disclosed exemplary embodiments of the invention. Although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation of the scope of the invention.
[0060] Although the present invention has been described in considerable detail with reference to certain preferred embodiments and examples thereof, other embodiments and equivalents are possible. Even though numerous characteristics and advantages of the present invention have been set forth in the foregoing description, together with functional and procedural details, the disclosure is illustrative only, and changes may be made in detail, especially in terms of the structuring and implementation within theprinciples of the invention to the full extent indicated by the broad general meaning of the terms. Thus various modifications are possible of the presently disclosed system and method without deviating from the intended scope and spirit of the present invention.
Claims
CLAIMSWE CLAIM1. A method for generating custom video content, the method comprising: receiving a source video, wherein the source video contains a plurality of frames; transcribing an audio of the source video to text to generate transcribed text, wherein the transcribed text corresponds to the timestamp of the plurality of frames in the source video; detecting regions of interest like objects, labels, text on frames corresponding to the timestamp of the plurality of frames in the source video using computer vision, wherein regions of interest is detected to generate visual meta data; receiving a request to generate the custom video content; processing the source video, wherein the processing comprises: feeding the transcribed text to a large language model, feeding the visual meta data to the large language model; identifying the type of content in the source video based on an output from the large language model; extracting relevant frames corresponding based on received request, type of content and the output from the large language model, generating the custom video content based on the extracted relevant frames; transforming the generated custom video content based on the detected region of interest in the source video.
2. The method as claimed in claim 1, wherein the transforming the custom video content comprises converting the custom video from original aspect ratio to multiple aspect ratio including but not limited to 9:16,4:5,1:1.
3. The method as claimed in claim 1, further comprising:generating metadata and captions for the custom video content based on the transcribed text and regions of interest on frames.
4. The method as claimed in claim 1, further comprising: generating thumbnails for the custom video content based on processing of the source video.
5. The method as claimed in claim 1, wherein the source video is either a stored video file or livestream video.
6. An apparatus for generating custom video content, the apparatus comprising: a reading module configured to receiving a source video, wherein the source video contains a plurality of frames; a processor configured to: transcribe an audio of the source video to text to generate transcribed text, wherein the transcribed text corresponds to the timestamp of the plurality of frames in the source video; detect regions of interest like objects, labels, text on frames corresponding to the timestamp of the plurality of frames in the source video using computer vision, wherein regions of interest is detected to generate visual meta data; receive a request to generate the custom video content; process the source video, wherein the processing comprises: feed the transcribed text to a large language model, feed the visual meta data to a large language model; identify the type of content in the source video based on an output from the large language model; extract relevant frames corresponding based on received request, type of content and the output from the large language model, generate the custom video content based on the extracted relevant frames; transform the generated custom video content based on the detected region of interest in the source video.
7. The apparatus as claimed in claim 6, wherein the transforming the custom video content comprises converting the custom video from original aspect ratio to multiple aspect ratio including but not limited to 9:16,4:5,1:1.
8. The apparatus as claimed in claim 6, further comprising: generating metadata and captions for the custom video content based on the transcribed text and regions of interest on frames.
9. The apparatus as claimed in claim 6, further comprising: generating thumbnails for the custom video content based on processing of the source video.
10. The apparatus as claimed in claim 6, wherein the source video is either a stored video file or livestream video.
Citation Information
Patent Citations
Content system with user-input based video content generation feature
US11769531B1
Customizable framework to extract moments of interest
US20230140369A1