Method and system for providing automated personalized and dynamic content tags to a user
The method and system generate personalized and dynamic content tags using a large language model to enhance user engagement by adapting to user preferences and content context, addressing inefficiencies in existing static tagging methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- JIOSTAR INDIA PTE LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-07-30
AI Technical Summary
Existing content tagging methods are static, resource-intensive, and lack adaptability, failing to cater to individual user preferences and evolving content contexts, leading to inefficient content discovery and engagement.
A method and system using a large language model to generate personalized and dynamic content tags synchronized with media content, adapting to user attributes and context, and displayed during playback to enhance engagement.
Automated content tags improve user engagement by aligning with user interests and content context, optimizing content consumption and discovery.
Smart Images

Figure IN2026050093_30072026_PF_FP_ABST
Abstract
Description
“METHOD AND SYSTEM FOR PROVIDING AUTOMATED PERSONALIZED AND DYNAMIC CONTENT TAGS TO A USER”TECHNICAL FIELD
[0001] The present invention relates to the field of generation of automated content tags for watching media content. More specifically, the invention of the present disclosure relates to a method for providing automated personalized and dynamic content tags to a user for watching a media content.BACKGROUND
[0002] In the current digital media landscape, content catalogues are filled with a vast collection of movies, TV shows, and other forms of media content on various content streaming platforms. As these platforms continue to grow in both size and diversity, presenting right content to the right audience has become a significant challenge. A critical aspect of this challenge lies in the manual and static process of tagging content with relevant content tags, unique selling points (USPs) or reasons-to-watch (RTW) attributes that appeal to users.
[0003] Existing methods of providing content tags involve generation of static, and predefined labels that are often limited in scope and relevance. These tags are based on general information, such as genre, cast, or high-level thematic elements, and are assigned manually by operators or content curators. The existing methods present a number of limitations and inefficiencies as mentioned below:• Lack of diversity for different users: Static RTWs often fail to cater to different interests and preferences of individual users. As a result, the content discovery experience may become “one-size-fits-all”, disregarding the unique preferences and viewing behaviours of specific user cohorts.• Manual and professional operations: The process of providing content tags based on individual user preference is resource-intensive and requires skilled professionals to assess and label each title manually. This process may be slow, error-prone, and inefficient, leading to delays in content availability and missed opportunities for engagement. In addition, the manual tagging of content is often a repetitive task, requiring continual updates and adjustments to reflect new trends, content releases, or changes in user behaviour. This repetitive process may lead to inefficiencies.• Inefficiency in cost and time: Due to the manual intervention, the content tagging process may be time-consuming and costly. As a streaming platform’s catalogue grows, the complexity and expense of maintaining an up-to-date and relevant set of content tags or USPs also increase, resulting in a suboptimal use of resources.• Lack of adaptability: Over time, the context in which content is consumed evolves. For instance, a content title may initially have a USP related to its theatrical release, but as time progresses, new releases or trends in the market may make this original USP less relevant. Additionally, as a piece of content transitions from a theatrical release to an over-the-top (OTT) platform, its appeal may shift, and the associated RTWs may no longer reflect the most compelling reasons for a user to watch the content.• Non-Reusable content tags or USPs: The traditional method of creating USPs often leads to a lack of reusability. Once a USP is defined for a particular content title, it may become outdated or irrelevant overtime as audience preferences change, new content is introduced, or the content itself transitions through different phases (e.g., from cinemas to streaming platforms). This results in USPs that cannot be effectively reused or adapted to future content or changing user needs.
[0004] To overcome the above-mentioned issues / limitations, there exists a need for automating generation of content tags for a user based on the user’s preferences. The present invention addresses these issues by disclosing an innovative method for providing automated personalized and dynamic content tags to a user for watching a media content.OBJECTIVES
[0005] This section is provided to introduce certain objectives and aspects of the present invention in a simplified format, that are then further elaborated upon in the subsequent paragraphs provided in the section of detailed description of the present disclosure.
[0006] In order to overcome at least a few of the problems of the known solutions as provided in the previous section, an objective of the present invention is to substantially reduce the limitations and / or drawbacks of the prior arts as described herein above.
[0007] Another object of the present invention is to provide automated personalized and dynamic content tags to a user for watching media content.
[0008] Yet another object of the present invention is to synchronize the content tags with specific scenes or durations (such as dialogues, or emotional peaks) within the media content to provide the user with reasons to watch the media content.
[0009] Yet another object of the present disclosure is to display context-aware content tags to the user, if a user pauses, or stops watching at a certain scene of the media content to encourage continued engagement.
[0010] Yet another object of the present disclosure is to make use of dynamic content tags as indexes tied to a time stamp or a scene in the media content to provide a navigation option for the user to browse through the media content.SUMMARY
[0011] In an aspect, the present disclosure relates to a method for providing automated personalized and dynamic content tags to a user for watching a media content. The method comprises extracting, by a processing unit, information related to the media content from one or more external sources. The method comprises processing, by the processing unit, the extracted information using a large language model (LLM) to generate one or more dynamic content tags. The one or more dynamic content tags are personalized based on one or more user attributes. Further, the method comprises synchronizing, by a synchronization unit, the one or more dynamic content tags with the media content. Synchronizing the one or more dynamic content tags is done to contextually align the one or more dynamic content tags with one or more scenes of the media content. Also, the method comprises displaying, by a display unit, the synchronized one or more dynamic content tags with the media content to the user during media content discovery and media content playback, in an event the user pauses at a scene of the media content.
[0012] In an embodiment, the method relates to synchronizing the one or more dynamic content tags with the media content comprises analyzing subtitles, audio descriptions and scene segmentation data, and aligning the one or more dynamic content tags with the one or more scenes of the media content based on the analysis.
[0013] In another embodiment, the generated one or more dynamic content tags are stored in a vector database.
[0014] In another embodiment, the one or more user attributes comprise at least a user preference and a user interest.
[0015] In yet another embodiment, the one or more dynamic content tags are linked to a progress bar and serve as one or more indexes for the user to navigate within the media content.
[0016] In yet another embodiment, the method relates to tracking, by the processing unit, the one or more user attributes to update the one or more dynamic content tags in real-time for improving user engagement. The method further relates to updating and adjusting, by the processing unit, the one or more dynamic content tags based on at least the one or more user attributes, media content progress, and the one or more user attributes using a Retrieval -Augmented Generation (RAG) model.
[0017] In yet another embodiment, the method comprises training the RAG model to prioritize a dynamic content tag from the one or more dynamic content tags based on the one or more user attributes.
[0018] In yet another embodiment, the one or more dynamic content tags are tailored to specific user cohorts, wherein the user cohorts are dynamically adjusted based on at least: user demographics, viewing history, and the one or more user attributes.
[0019] In yet another embodiment, the media content progress comprises at least one of: pausing, resuming, forwarding or stopping the media content.
[0020] In yet another embodiment, the media content comprises at least a video content, an audio content, and a text-based content.
[0021] In yet another embodiment, the one or more external sources comprise at least third-party databases, online streaming platforms, social media and communication devices of the user.
[0022] In another aspect, the present disclosure provides a system for providing automated personalized and dynamic content tags to a user for watching a media content. The system comprises a processing unit configured to extract information related to the media content from one or more external sources and processes the extracted information using a large language model (LLM) to generate one or more dynamic content tags. The one or more dynamic contenttags are personalized based on one or more user attributes. Further, the system comprises a synchronization unit configured to synchronize the one or more dynamic content tags with the media content. Synchronization of the one or more dynamic content tags is done to contextually align the one or more dynamic content tags with one or more scenes of the media content. The system further comprises a display unit configured to display the synchronized one or more dynamic content tags with the media content to the user during media content discovery and media content playback, in an event the user pauses at a scene of the media content.BRIEF DESCRIPTION OF DRAWINGS
[0023] FIG. 1 illustrates a block diagram of system for providing automated personalized and dynamic content tags to a user for watching a media content, in accordance with an embodiment of the present disclosure.
[0024] FIG. 2 illustrates an exemplary embodiment of a computing device in which the system may be employed, in accordance with an embodiment of the present disclosure.
[0025] FIG. 3 illustrates a flowchart depicting a method for providing automated personalized and dynamic content tags to the user for watching the media content, in accordance with an embodiment of the present disclosure.
[0026] FIG. 4 illustrates an exemplary workflow diagram for the system, in accordance with an embodiment of the present disclosure.
[0027] FIG. 5 illustrates a use case of the system, in accordance with an embodiment of the present disclosure.DETAILED DESCRIPTION
[0028] In the following description, for the purposes of explanation, various specific details are set forth in order to provide a thorough understanding of embodiments of the present disclosure. It will be apparent, however, that embodiments of the present disclosure may be practiced without these specific details. Several features described hereafter may each be used independently of one another or with any combination of other features. An individual feature may not address any of the problems discussed above or might address only some of the problems discussed above.
[0029] The ensuing description provides exemplary embodiments only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing an exemplary embodiment. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the disclosure as set forth.
[0030] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood by one of ordinary skilled in the art that the embodiments may be practiced without these specific details. For example, circuits, systems, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail.
[0031] Also, it is noted that individual embodiments may be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in any figure.
[0032] The word “exemplary” and / or “demonstrative” is used herein to mean serving as an example, instance, or illustration. For the avoidance of doubt, the subject matter disclosed herein is not limited by such examples. In addition, any aspect or design described herein as “exemplary” and / or “demonstrative” is not necessarily to be construed as preferred or advantageous over other aspects or designs, nor is it meant to preclude equivalent exemplary structures and techniques known to those of ordinary skill in the art. Furthermore, to the extent that the terms “includes,” “has,” “contains,” and other similar words are used in either the detailed description or the claims, such terms are intended to be inclusive — in a manner similar to the term “comprising” as an open transition word — without precluding any additional or other elements.
[0033] As used herein, a “processing unit” or “processor” or “operating processor” includes one or more processors, wherein processor refers to any logic circuitry for processing instructions. A processor may be a general -purpose processor, a special purpose processor, a conventional processor, a digital signal processor, a plurality of microprocessors, one or moremicroprocessors in association with a (Digital Signal Processing) DSP core, a controller, a microcontroller, Application Specific Integrated Circuits, Field Programmable Gate Array circuits, any other type of integrated circuits, etc. The processor may perform signal coding data processing, input / output processing, and / or any other functionality that enables the working of the system according to the present disclosure. More specifically, the processor or processing unit is a hardware processor.
[0034] As used herein, “a user equipment”, “a user device”, “a smart-user-device”, “a smart-device”, “an electronic device”, “a mobile device”, “a handheld device”, “a wireless communication device”, “a mobile communication device”, “a communication device”, “ a client device” may be any electrical, electronic and / or computing device or equipment, capable of implementing the features of the present disclosure. The user equipment / device may include, but is not limited to, a mobile phone, smart phone, laptop, a general -purpose computer, desktop, personal digital assistant, tablet computer, wearable device, television, living room media devices, streaming consoles (firestick, Xbox), headwear units (oculus, apple vision pro), AR / VR / MR devices, shopping kiosk, or any other computing device which is capable of implementing the features of the present disclosure. Also, the user device may contain at least one input means configured to receive an input from at least one of a transceiver unit, a processing unit, a storage unit, a detection unit and any other such unit(s) which are required to implement the features of the present disclosure.
[0035] As used herein, “storage unit” or “memory unit” refers to a machine or computer-readable medium including any mechanism for storing information in a form readable by a computer or similar machine. For example, a computer-readable medium includes read-only memory (“ROM”), random access memory (“RAM”), magnetic disk storage media, optical storage media, flash memory devices or other types of machine-accessible storage media. The storage unit stores at least the data that may be required by one or more units of the system to perform their respective functions.
[0036] All modules, units, components used herein, unless explicitly excluded herein, may be software modules or hardware processors, the processors being a general -purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASIC), Field Programmable Gate Array circuits (FPGA), any other type of integrated circuits, etc.
[0037] As discussed in the background section, the current known solutions have several shortcomings. The present disclosure aims to overcome the above-mentioned and other existing problems in this field of technology by providing a method and system for providing automated personalized and dynamic content tags to a user for watching a media content.
[0038] As used herein, media content is any type of content that is produced, distributed, and consumed through various media channels. This includes, but is not limited to, video, audio, text, images, and other multimedia formats. The media content may be consumed on platforms such as streaming services, television, websites, social media, and mobile apps. The media content encompasses a wide range of formats such as movies, TV shows, music, podcasts, articles, blogs, e-books, and digital advertisements.
[0039] As used herein, dynamic content tags may refer to content summary, content descriptors or unique selling propositions (USPs). In general, content tags are descriptive labels or keywords associated with the media content to provide additional context, categorization, or relevant information about the media content. In the present disclosure, the content tags may refer to unique selling propositions (USPs) that highlight key aspects of the media content, such as "action-packed," "most-watched," "award -winning," "critically acclaimed", and the like.
[0040] Embodiments of the present disclosure relate to a system and method for providing automated personalized and dynamic content tags to a user for watching a media content. The system automates generation and synchronization of personalized, dynamic content tags for the media content using machine learning and large language models (LLMs). The system is configured to analyze user behavior, content metadata, and external data sources to create context-aware tags that are displayed during media discovery and playback, enhancing user engagement and encouraging continued viewing. The system adapts the content tags in real-time based on user interactions and content progress to maintain relevance and optimize content consumption.
[0041] The present disclosure may now be explained in detail with the help of a few preferred embodiments that are described herein below.
[0042] FIG. 1 illustrates a block diagram of a system
[0100] for providing automated personalized and dynamic content tags to a user for watching a media content, in accordance with an embodiment of the present disclosure. In an exemplary embodiment, the media contentmay be content streaming on online streaming platforms (such as over-the-top platforms), websites, social media platforms and the like.
[0043] The system
[0100] includes a processing unit
[0102] , a synchronization unit
[0104] , a large language model (LLM)
[0106] , a media packager
[0108] , and a display unit
[0110] , The processing unit
[0102] is configured to extract information related to the media content from one or more external sources. The extracted information is public information. The media content comprises at least a video content, an audio content, and a text-based content. The extracted information is about the media content (e.g., a movie, TV show, or video) such as genre, plot summary, key scenes, social media sentiment, and the like. The one or more external sources may comprise at least open data sources, commercial tools, third-party databases, online streaming platforms, social media platforms and communication devices of the user. In an example, the one or more external sources may be Internet Movie Database (IMDb), Wikipedia, or rotten tomatoes, that include detailed information about the media content such as synopsis, cast, user ratings, and the like.
[0044] Further, the processing unit
[0102] is configured to process the extracted information using the large language model (LLM)
[0106] to generate one or more dynamic content tags. The one or more dynamic content tags are personalized based on one or more user attributes. The one or more user attributes comprise at least one of: a user preference and a user interest. For example, if a user has a pattern (recognized from past viewing history of the user) of watching short movies and mostly watches comedy media content. So, based on user preference of “short movies” and user interest related to “comedy media content”, the generated dynamic content tags may be “Laugh Your Way through 60 minutes movie”, “1 hour movie that will make you LOL”, etc. The generated one or more dynamic content tags are stored in a vector database (not shown in FIG. 1). Generally, a vector database is a dedicated database designed to store and manage data in the form of vectors, which are mathematical representations of any information. When applied to the generated one or more dynamic content tags, the vector database is used to store the one or more dynamic tags as embeddings such as numerical representations. The term “generated one or more dynamic media content tags” is interchangeably used with “one or more dynamic content tags” throughout the detailed description.
[0045] Further, the processing unit
[0102] may be associated with the synchronization unit
[0104] , In an exemplary embodiment, the processing unit
[0102] may send the generated one or more dynamic content tags to the synchronization unit
[0104] for synchronization with the mediacontent. The synchronization unit
[0104] is configured to synchronize the one or more dynamic content tags with the media content. In addition, the one or more dynamic content tags are synchronized to contextually align the one or more dynamic content tags with one or more scenes of the media content. To synchronize the one or more dynamic content tags, the synchronization unit
[0104] is configured to analyze subtitles, audio descriptions and scene segmentation data of the media content. In an embodiment, the subtitles, audio descriptions and scene segmentation data may be extracted from the LLM model
[0106] ,
[0046] In general, scene segmentation data refers to structured information derived from analyzing the media content (e.g., videos, movies, or TV shows) to divide it into meaningful segments or scenes. Each scene is a distinct portion of the media content, grouped based on changes in visual, auditory, or contextual elements. For example, a movie titled “The Time Traveler’s Journal” may be segmented into Scene 1, Scene 2, Scene 3 and the like and the one or more content tags may be synchronized with the respective scenes such as:Scene 1: The protagonist discovers a time-travel deviceContent Tags for Scene 1: “Time Travel Invention”, “Quantum Physics”, “Key Plot Device”.Scene 2: The protagonist travels to medieval England and meets a princessContent Tags for Scene 2: “Medieval Romance”, “Historical Setting”, “Forbidden Love”.Scene 3 : A dramatic battle to protect the time machine.Content Tags for Scene 3: “Action-Packed Climax”, “Strategic Warfare”, “High Stakes”.
[0047] Further, based on the analysis, the synchronization unit
[0104] is configured to align the one or more dynamic content tags with the one or more scenes of the media content. The one or more scenes may be some of the interesting scenes of the media content such as “a big reveal scene”, “an emotional scene”, “a horror scene”, and the like.
[0048] The one or more dynamic content tags are tailored to specific user cohorts. The user cohorts are dynamically adjusted based on at least: user demographics, viewing history, and the one or more user attributes. In general, user cohorts refer to a group of users with shared interests or preferences. In an example, let’s consider a streaming platform showcasing a movie titled “The Great Expedition”, a historical adventure film featuring exploration, battles, and personal relationships. Using the LLM model
[0106] , three user cohorts (User cohort 1, User cohort 2, andUser cohort 3) are identified, Each of the three user cohorts have viewing history more focused towards historical stories. The processing unit
[0102] generates three types of the one or more dynamic content tags, one type for each user cohort.
[0049] For example, User cohort 1 are adults aged 30 to 50, and their viewing history includes historical documentaries, biopics and war movies. Also, it is predicted, using the LLM model
[0106] , that the user interests of the User cohort 1 include exploration, historical accuracy, famous battles and the like. So, the processing unit
[0102] , in this case, may generate the one or more dynamic content tags such as “18th Century Warfare”, “Famous Jungle Battles”, “Historical Tactics Used”, and the like. These type of content tags emphasize historical relevance and accuracy, catering to the user’s interest in history and are specific to the User cohort 1.
[0050] In addition, User cohort 2 are teenagers aged 13 to 19 and their viewing history includes romantic stories, historical dramas, and friendship stories. Based on their viewing history, user interests are predicted and further the processing unit
[0102] generates the one or more dynamic content tags (specific to user interests of the User cohort 2) such as “Sacrifice for Friendship”, “Unlikely Allies”, “Everything is fair in love and war”, “Emotional Goodbye”, and the like. These content tags for the User cohort 2 focuses more on emotional dynamics and relationships during the battle, to engage a younger audience i.e., the User cohort 2.
[0051] Further, User cohort 3 are people of age group 20 to 30 and their viewing history includes action thrillers, superhero movies, fight sequences, and the like. Based on the viewing history, the LLM model
[0106] predicts the user interests, “high energy scenes, action sequences and suspense”. Based on the user interests, the processing unit
[0102] generates the one or more dynamic content tags such as “Epic Fight Scene”, “High-Stakes Combat”, “Explosive Jungle Warfare”, highlighting adrenaline pumping action sequences and visual excitement.
[0052] The processing unit
[0102] is further configured to track the one or more user attributes to update the one or more dynamic content tags in real-time for improving user engagement. The processing unit
[0102] updates the user cohorts and the one or more dynamic content tags in real-time based on user behaviour. For example, a teenager who initially views romantic dramas but starts exploring historical documentaries may see his or her cohort adjusted. The content tags generated for “The Great Expedition” shift from emotional themes to include more historical references. An action seeker who skips non-action scenes may trigger the processing unit
[0102] to emphasize action-related content tags for similar movies in the future.
[0053] In an example, the user cohorts are adjusted based on user demographics as users from a specific region may prefer media content in their native language or content with cultural relevance. For example, if a user is a teenager in the U.S., their dynamic content tags may highlight fun, action-packed, or trending elements of the media.
[0054] In an embodiment, the synchronization unit
[0104] may be associated with the media packager
[0108] and the display unit
[0110] , In another embodiment, the media packager
[0108] may be associated with each component (processing unit, synchronization unit, LLM model, display unit) of the system
[0100] to enable generation of the one or more dynamic content tags and further display the one more or more dynamic content tags to the user. In an example, the media packager
[0108] is used for organizing and preparing the media content for distribution across various platforms. The media packager
[0108] may involve gathering, editing, and formatting of media assets (such as videos, audio files, images, graphics, and other digital content) to ensure they are compatible with different distribution channels, such as television, streaming services, websites, social media, or mobile applications. In an embodiment, the media packager
[0108] may take an input to create a graphical presentation of the one or more dynamic content tags and push to the display unit
[0110] , to be overlay ed on a content title, watch page or some other location on the display unit
[0110] ,
[0055] The display unit
[0110] is configured to display the synchronized one or more dynamic content tags, with facilitation of the media packager
[0108] , to the user during media content discovery and media content playback, in an event the user pauses at a scene of the media content.
[0056] In an example, the display unit
[0110] displays the one or more dynamic content tags during media content discovery, when a user is exploring the steaming platform. The one or more dynamic content tags are displayed at a communication device of the user alongside thumbnails or previews of the media content. For instance: Content tags such as “Award-Winning”, “Based on a True Story”, or “Romantic Thriller” may be displayed that help the user to understand context of the media content, at a glance. In another example, the display unit
[0110] displays the one or more dynamic content tags during media content playback (when the user pauses while watching a media content). For instance: If a user pauses the media content while watching, the one or more dynamic content tags for a specific scene the user has paused at may be displayed on the communication device of the user. Also, the one or more dynamic content tags for an upcoming scene may be displayed so that the user may directly view the upcoming scene andstay engaged. The one or more dynamic content tags in this scenario may include “mystery unlocking in next 15 minutes”, “upcoming fun ahead”, and the like.
[0057] In addition, the processing unit
[0102] is configured to update and adjust the one or more dynamic content tags based on at least the one or more user attributes and media content progress using a Retrieval-Augmented Generation (RAG) model. The media content progress comprises at least one of pausing, resuming, forwarding or stopping the media content. In an example, let us assume, that a user A is watching “a comedy movie” and suddenly the user A has paused the movie as the user A is not finding the movie interesting, and even after a certain time duration (say, 30 minutes), the user A has not resumed the said movie. The processing unit
[0102] , in this case, may track that the user A has paused the movie for certain time duration and may update content tag for an upcoming scene in the movie so that the user A is convinced to resume the movie after seeing the updated content tag. For example, the processing unit
[0102] may update the content tag to “Big reveal coming up - don’t miss it!”. The processing unit
[0102] may adjust the content tag. For example, early on in a movie, the content tag may be: “Get to know the characters - emotional depth ahead”. Later in the movie, when the user has paused after watching for a few minutes, the processing unit
[0102] may adjust the content tag: “The stakes just got higher - don’t miss the twist coming up”.
[0058] In an example, if a user is a 25-year-old male from India, the processing unit
[0102] may generate few initial dynamic content tags that may highlight cultural elements, action, or popular Bollywood tropes, based on demographic information of the user. Also, the processing unit
[0102] may generate the one or more dynamic content tags based on viewing history of the user. Let’s assume that the user mostly watches comedy films but skips musical sequences. In this scenario, the user cohorts are adjusted to emphasize comedic aspects like “A hilarious twist!” or “Laugh-out-loud moments ahead!” while avoiding musical-related tags. Further, the processing unit
[0102] may generate the one or more dynamic content tags based on the media progress. For example, let’s assume that a user pauses during intense scenes but resumes quickly. In this scenario. The processing unit
[0102] may adjust the existing one or more dynamic content tags to user-preference-based tags such as “This thrilling scene isn’t over yet!” or “Don’t miss the ultimate payoff!”, if they pause during playback.
[0059] The processing unit
[0102] is further configured to train the RAG model to prioritize a dynamic content tag from the one or more dynamic content tags based on the one or more user attributes. In an embodiment, the RAG model may be a part of the large language model (LLM)
[0106] , In another embodiment, the RAG model maybeatype ofLLMmodel
[0106] , In yet another embodiment, the RAG model is same as the LLM model
[0106] , In an embodiment, the RAG model may learn over time which dynamic content tags of the one or more dynamic content tags are most likely to engage the user based on the one or more user attributes. In one implementation, if the user has shown a preference for action-packed media content in the past (based on watching history, or engaging with action scenes for a longer duration), the RAG model may prioritize showing action-related dynamic tag for the user out of the one or more dynamic content tags.
[0060] For example, if the user has paused watching a movie and is near a specific scene (e.g., an action scene), the RAG model may prioritize showing tags related to that scene. In another example, if the user has a history of skipping over scenes that are too slow or emotional, the RAG model may generate tags that highlight exciting or dramatic moments to re-engage the user.
[0061] The one or more dynamic content tags are linked to a progress bar and serve as one or more indexes for the user to navigate within the media content. In an embodiment, once the user clicks on a scene specific dynamic content tag (herein after, content tag) of the one or more dynamic content tags, the content tag may redirect the user to that particular scene of the media content to make sure that the user remains engaged with the media content. For example, imagine a user watching an action-thriller movie, but at a certain point (say 20 minutes in), the user decides to pause. The movie has an exciting “big reveal scene” coming up at 30 minutes, and a "climatic fight scene" at a 60-minute mark. When the user pauses at 20 minutes, a content tag, “Big Reveal - The plot twist changes everything!” may appear, prompting the user to continue watching as something big is coming up. As soon as the user clicks on this content tag, the content tag may redirect the user directly to the 30thminute scene, bypassing less interesting parts of the movie. By linking the one or more dynamic tags to the progress bar and making them clickable, the user is given more control over their viewing experience.
[0062] FIG. 2 illustrates an exemplary embodiment of a computing device
[0200] in which the system 100 may be employed, in accordance with an embodiment of the present disclosure. In an implementation, the computing device
[0200] may also implement a method (explained in FIG. 3) for providing automated personalized and dynamic content tags to a user. In another implementation, the computing device
[0200] itself implements the method for providing automated personalized and dynamic content tags using one or more units configured within thecomputing device
[0200] , wherein said one or more units are capable of implementing the features as disclosed in the present disclosure.
[0063] The computing device
[0200] may include a bus
[0202] or other communication mechanism for communicating information, and a hardware processor
[0204] coupled with bus
[0202] for processing information. The hardware processor
[0204] may be, for example, a general-purpose microprocessor. The computing device
[0200] may also include a main memory
[0206] , such as a random-access memory (RAM), or other dynamic storage device, coupled to the bus
[0202] for storing information and instructions to be executed by the processor
[0204] , The main memory
[0206] also may be used for storing temporary variables or other intermediate information during execution of the instructions to be executed by the processor
[0204] , Such instructions, when stored in non-transitory storage media accessible to the processor
[0204] , render the computing device
[0200] into a special-purpose machine that is customized to perform the operations specified in the instructions. The computing device
[0200] further includes a read only memory (ROM)
[0208] or other static storage device coupled to the bus
[0202] for storing static information and instructions for the processor
[0204] ,
[0064] A storage device
[0210] , such as a magnetic disk, optical disk, or solid-state drive is provided and coupled to the bus
[0202] for storing information and instructions. The computing device
[0200] may be coupled via the bus
[0202] to a display
[0212] , such as a cathode ray tube (CRT), Liquid crystal Display (LCD), Light Emitting Diode (LED) display, Organic LED (OLED) display, etc. for displaying information to a computer user. An input device
[0214] , including alphanumeric and other keys, touch screen input means, etc. may be coupled to the bus
[0202] for communicating information and command selections to the processor
[0204] , Another type of user input device may be a cursor controller
[0216] , such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor
[0204] , and for controlling cursor movement on the display
[0212] , The input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allows the device to specify positions in a plane.
[0065] The computing device
[0200] may implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic which in combination with the computing device
[0200] causes or programs the computing device
[0200] to be a special-purpose machine. According to one implementation, the techniques herein are performed by the computing device
[0200] in response to the processor
[0204] executing oneor more sequences of one or more instructions contained in the main memory
[0206] , Such instructions may be read into the main memory
[0206] from another storage medium, such as the storage device
[0210] , Execution of the sequences of instructions contained in the main memory
[0206] causes the processor
[0204] to perform the process steps described herein. In alternative implementations of the present disclosure, hard-wired circuitry may be used in place of or in combination with software instructions.
[0066] The computing device
[0200] also may include a communication interface
[0218] coupled to the bus
[0202] , The communication interface
[0218] provides a two-way data communication coupling to a network link
[0220] that is connected to a local network
[0222] , For example, the communication interface
[0218] may be an integrated services digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, the communication interface
[0218] may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, the communication interface
[0218] sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0067] The computing device
[0200] can send messages and receive data, including program code, through the network(s), the network link
[0220] and the communication interface
[0218] , In the Internet example, a server
[0230] might transmit a requested code for an application program through the Internet
[0228] , the ISP
[0226] , the local network
[0222] and the communication interface
[0218] , The received code may be executed by the processor
[0204] as it is received, and / or stored in the storage device
[0210] , or other non- volatile storage for later execution.
[0068] FIG. 3 illustrates a flowchart depicting a method
[0300] for providing automated personalized and dynamic content tags to the user for watching the media content, in accordance with an embodiment of the present disclosure. The method
[0300] initiates at step
[0302] ,
[0069] Following step
[0302] , at step
[0304] , the method
[0300] includes extracting, by a processing unit
[0102] , information related to a media content from one or more external sources. The media content comprises at least a video content, an audio content, and a text-based content. The extracted information is about the media content (e.g., a movie, TV show, or video) such asgenre, plot summary, key scenes, social media sentiment, and the like. The one or more external sources comprise at least third-party databases, online streaming platforms, social media platforms and communication devices of the user. In an example, the one or more external sources may be IMDb, Wikipedia, or rotten tomatoes, that include detailed information about the media content such as synopsis, cast, user ratings, and the like.
[0070] At step
[0306] , the method
[0300] includes processing, by the processing unit
[0102] , the extracted information using a large language model (LLM) 106 to generate one or more dynamic content tags. The one or more dynamic content tags are personalized based on one or more user attributes. The one or more user attributes comprise at least one of: a user interaction and a user interest. The generated one or more dynamic content tags are stored in a vector database. Generally, a vector database is a dedicated database designed to store and manage data in the form of vectors, which are mathematical representations of any information. When applied to the generated one or more dynamic content tags, the vector database is used to store the one or more dynamic tags as embeddings such as numerical representations. The term “generated one or more dynamic media content tags” is interchangeably used with “one or more dynamic content tags” throughout the detailed description.
[0071] The one or more dynamic content tags are tailored to specific user cohorts. The user cohorts are dynamically adjusted based on at least: user demographics, viewing history, and the one or more user attributes. In an embodiment, the one or more user attributes comprises information related to user interaction and user interest. The information may be non-personal information or personal information or may be specific to a content or platform and is collected with the user’s consent. In general, user cohorts refer to a group of users with shared interests or preferences. In an example, let’s consider a streaming platform showcasing a movie titled “The Great Expedition”, a historical adventure film featuring exploration, battles, and personal relationships. Using the LLM model
[0106] , three user cohorts (User cohort 1, User cohort 2, and User cohort 3) are identified. Each of the three identified user cohorts have viewing history which is more focused towards historical stories. The processing unit
[0102] generates three types of the one or more dynamic content tags, with one type generated for each identified user cohort.
[0072] For example, User cohort 1 are adults aged 30 to 50, and their viewing history includes historical documentaries, biopics and war movies. Also, it is predicted, using the LLM model
[0106] , that the user interests of the User cohort 1 include exploration, historical accuracy, famous battles and the like. So, the processing unit
[0102] , in this case, may generate the one ormore dynamic content tags such as “18th Century Warfare”, “Famous Jungle Battles”, “Historical Tactics Used”, and the like. These type of content tags emphasize historical relevance and accuracy, catering to the user’s interest in history and are specific to the User cohort 1.
[0073] In addition, User cohort 2 are teenagers aged 13 to 19 and their viewing history includes romantic stories, historical dramas, and friendship stories. Based on their viewing history, user interests are predicted and further the processing unit
[0102] generates the one or more dynamic content tags (specific to user interests of the User cohort 2) such as “Sacrifice for Friendship”, “Unlikely Allies”, “Everything is fair in love and war”, “Emotional Goodbye”, and the like. These content tags for the User cohort 2 focuses more on emotional dynamics and relationships during the battle, to engage a younger audience i.e., the User cohort 2.
[0074] Further, User cohort 3 are people of age group 20 to 30 and their viewing history includes action thrillers, superhero movies, fight sequences, and the like. Based on the viewing history, the LLM model
[0106] predicts the user interests, “high energy scenes, action sequences and suspense”. Based on the user interests, the processing unit
[0102] generates the one or more dynamic content tags such as “Epic Fight Scene”, “High-Stakes Combat”, “Explosive Jungle Warfare”, highlighting adrenaline pumping action sequences and visual excitement.
[0075] The method
[0300] further includes tracking, by the processing unit
[0102] the one or more user attributes to update the one or more dynamic content tags in real-time for improving user engagement. The processing unit
[0102] updates the user cohorts and the one or more dynamic content tags in real-time based on user behaviour. For example, a teenager who initially views romantic dramas but starts exploring historical documentaries may see their cohort adjusted. The content tags generated for “The Great Expedition” shift from emotional themes to include more historical references. An action seeker who skips non-action scenes may trigger the processing unit
[0102] to emphasize action-related content tags for similar movies in the future.
[0076] At step
[0308] , the method
[0300] includes synchronizing, by a synchronization unit
[0104] , the one or more dynamic content tags with the media content. In addition, synchronizing the one or more dynamic content tags is done to contextually align the one or more dynamic content tags with one or more scenes of the media content. In addition, the one or more dynamic content tags are synchronized to contextually align the one or more dynamic content tags with one or more scenes of the media content.
[0077] To synchronize the one or more dynamic content tags, the synchronization unit
[0104] is configured to analyze subtitles, audio descriptions and scene segmentation data of the media content. In an embodiment, the subtitles, audio descriptions and scene segmentation data may be extracted from the LLM model
[0106] , In general, scene segmentation data refers to structured information derived from analyzing the media content (e.g., videos, movies, or TV shows) to divide it into meaningful segments or scenes. Each scene is a distinct portion of the media content, grouped based on changes in visual, auditory, or contextual elements. For example, a movie titled “The Time Traveler’s Journal” may be segmented into Scene 1, Scene 2, Scene 3 and the like and the one or more content tags may be synchronized with the respective scenes such as:Scene 1: The protagonist discovers a time-travel deviceContent Tags for Scene 1: “Time Travel Invention”, “Quantum Physics”, “Key Plot Device”.Scene 2: The protagonist travels to medieval England and meets a princessContent Tags for Scene 2: “Medieval Romance”, “Historical Setting”, “Forbidden Love”.Scene 3 : A dramatic battle to protect the time machine.Content Tags for Scene 3: “Action-Packed Climax”, “Strategic Warfare”, “High Stakes”.
[0078] Further, based on the analysis, the synchronization unit
[0104] is configured to align the one or more dynamic content tags with the one or more scenes of the media content.
[0079] The one or more scenes may be some of the interesting scenes of the media content such as “a big reveal scene”, “an emotional scene”, “a horror scene”, and the like. The one or more dynamic content tags are tailored to specific user cohorts. The user cohorts are dynamically adjusted based on at least: user demographics, viewing history, and the one or more user attributes.
[0080] At step
[0310] , the method
[0300] includes displaying, at a display unit
[0110] , the synchronized one or more dynamic content tags with the media content to the user during media content discovery and media content playback, in an event the user pauses at a scene of the media content. In an example, the display unit
[0110] displays the one or more dynamic content tags during media content discovery, when a user is exploring the steaming platform. The one or more dynamic content tags are displayed at a communication device of the user alongside thumbnailsor previews of the media content. For instance: Content tags such as “Award-Winning”, “Based on a True Story”, or “Romantic Thriller” may be displayed that help the user to understand context of the media content, at a glance. In another example, the display unit
[0110] displays the one or more dynamic content tags during media content playback (when the user pauses while watching a media content). For instance: If a user pauses the media content while watching, the one or more dynamic content tags for a specific scene the user has paused at may be displayed on the communication device of the user. Also, the one or more dynamic content tags for an upcoming scene may be displayed so that the user may directly view the upcoming scene and stay engaged. The one or more dynamic content tags in this scenario may include “mystery unlocking in next 15 minutes”, “upcoming fun ahead”, and the like.
[0081] The method
[0300] further includes tracking, by the processing unit
[0102] , the one or more user attributes to update the one or more dynamic content tags in real-time for improving user engagement. In addition, the processing unit
[0102] is configured to update and adjust the one or more dynamic content tags based on at least the one or more user attributes, and media content progress using a Retrieval- Augmented Generation (RAG) model. The media content progress comprises at least one of pausing, resuming, forwarding or stopping the media content. In an example, let us assume, that a user A is watching “a comedy movie” and suddenly the user A has paused the movie as the user A is not finding the movie interesting and even after a certain time duration (say, 30 minutes), the user A has not resumed the said movie. The processing unit
[0102] , in this case, may track that the user A has paused the movie for certain time duration and may update content tag for an upcoming scene in the movie so that the user A is convinced to resume the movie after seeing the updated content tag. For example, the processing unit
[0102] may update the content tag to “Big reveal coming up - don’t miss it!”. The processing unit
[0102] may adjust the content tag. For example, early on in a movie, the content tag may be: “Get to know the characters - emotional depth ahead”. Later in the movie, when the user has paused after watching for a few minutes, the processing unit
[0102] may adjust the content tag: “The stakes just got higher - don’t miss the twist coming up”.
[0082] The method
[0300] terminates at step
[0312] ,
[0083] FIG. 4 illustrates an exemplary workflow diagram
[0400] for the system
[0100] , in accordance with an embodiment of the present disclosure. The workflow diagram
[0400] represents a high-level workflow for providing one or more dynamic content tags to a user forwatching an upcoming media content. The upcoming media content is a new media content that is ingested into the system
[0100] ,
[0084] The upcoming media content such as a new movie or TV series is analyzed and information related to the media content is extracted, at step
[0402] , The information may be metadata such as genre, scenes, synopsis, and the like. At step
[0404] , the extracted information is analysed using the LLM model
[0106] as also explained in FIG. 1. The extracted information is analysed to generate one or more dynamic content tags. At step
[0406] , the one or more dynamic content tags are generated based on the one or more user attributes as explained in FIG. 1.
[0085] Further, at step
[0408] , the generated one or more dynamic content tags are reviewed before finalizing. Reviewing of the generated one or more dynamic content tags ensure that the one or more dynamic content tags align with the media content’s context and meet the user preferences. The review may be done manually or may be done by an intelligent model automatically. In an example, review of the generated one or more dynamic content tags may be in a manual, semi-automated, or fully automated manner. After the review of the generated one or more dynamic content tags, step
[0410] is followed. At step
[0410] , the generated one or more dynamic content tags are ingested to the synchronization unit
[0104] of the system
[0100] ,
[0086] The synchronization unit
[0104] is configured to synchronize the one or more dynamic content tags with the media content. In addition, the one or more dynamic content tags are synchronized to contextually align the one or more dynamic content tags with one or more scenes of the media content. To synchronize the one or more dynamic content tags, the synchronization unit
[0104] is configured to analyze subtitles, audio descriptions and scene segmentation data. In an embodiment, the subtitles, audio descriptions and scene segmentation data may be extracted from the one or more external sources. In another embodiment, the subtitles, audio descriptions and scene segmentation data may be extracted from the LLM model
[0106] , Further, based on the analysis, the synchronization unit
[0104] is configured to align the one or more dynamic content tags with the one or more scenes of the media content. The one or more scenes may be some of the interesting scenes of the media content such as “a big reveal scene”, “an emotional scene”, “a horror scene”, and the like. The one or more dynamic content tags are tailored to specific user cohorts. The user cohorts are dynamically adjusted based on at least: user demographics, viewing history, and the one or more user attributes.
[0087] If errors are found in reviewing of the one or more dynamic content tags, the same is flagged for iterative improvement, at step
[0412] , Instances where the LLM
[0106] generates incorrect, irrelevant, or misleading content tags (referred to as bad cases), errors may occur and the content tags are collected for further analysis, at step
[0412] , At step
[0412] , bad cases collection is done. The errors may occur from issues like misinterpreted metadata, incomplete context, or ambiguous user preferences. In addition, the errors may occur when the one or more dynamic content tags are irrelevant to the context of the media content. Using bad cases as training data, the LLM
[0106] is iteratively improved and retrained to enhance its accuracy and relevance, at step
[0414] , In an example, if a user consistently skip certain tags or scenes, algorithm of the LLM
[0106] may adjust the content tags by prioritizing content tags that highlight more engaging aspects of the media content.
[0088] FIG. 5 illustrates a use case
[0500] of the system
[0100] , in accordance with an embodiment of the present disclosure. The use case
[0500] depicts two different scenarios, a first scenario
[0502] and a second scenario
[0504] , The first scenario
[0502] is for User A and the second scenario
[0504] is for User B. Both User A and User B are viewing a home page of an online streaming platform installed on their respective communication devices.
[0089] The communication devices include but may not be limited to a mobile device, a smart device, a computer, a palmtop, a laptop, a smartwatch, and the like. The home page in the first scenario
[0502] and the second scenario
[0504] shows a media content recommendation on top of the home page. In other words, for both User A and User B, the recommended media content displayed on the home page is same, i.e., a web series “XYZ”. However, for User A, content tag for the movie “XYZ” is shown as “Guinness world record - Most thrilling series” as shown in the first scenario
[0502] , and for User B, content tag for the same web seriesis shown as “Academy Award Winner - Top Suspense Series”. Even though, the web series is same for User A and User B, however, the content tags for both the users are different and is based on the one or more user attributes, such as the user preference and the user interest as explained above in FIG. 1.
[0090] As is evident from the paragraphs above, the present disclosure provides a technically advanced solution for providing automated personalized and dynamic content tags for a user for watching a media content. Thus, in view of the above disclosure, system and method are cost-effective, easy to scale, and assist platform to consistently provide a high-quality personalized service to the users without compromise.
[0091] While considerable emphasis has been placed herein on the disclosed implementations, it will be appreciated that many implementations can be made and that many changes can be made to the implementations without departing from the principles of the present disclosure. These and other changes in the implementations of the present disclosure will be apparent to those skilled in the art, whereby it is to be understood that the foregoing descriptive matter to be implemented is illustrative and non-limiting.
Claims
We claim:
1. A method [300] for providing automated personalized and dynamic content tags to a user for watching a media content, the method [300] comprising:extracting, by a processing unit [102], information related to the media content from one or more external sources;processing, by the processing unit [102], the extracted information using a large language model (LLM) to generate one or more dynamic content tags, wherein the one or more dynamic content tags are personalized based on one or more user attributes;synchronizing, by a synchronization unit [104], the one or more dynamic content tags with the media content, wherein synchronizing the one or more dynamic content tags is done to contextually align the one or more dynamic content tags with one or more scenes of the media content; anddisplaying, by a display unit [110], the synchronized one or more dynamic content tags with the media content to the user during media content discovery and media content playback, in an event the user pauses at a scene of the media content.
2. The method [300] as claimed in claim 1, wherein synchronizing the one or more dynamic content tags with the media content comprises:analyzing subtitles, audio descriptions and scene segmentation data; and aligning the one or more dynamic content tags with the one or more scenes of the media content based on the analysis.
3. The method [300] as claimed in claim 1, wherein the generated one or more dynamic content tags are stored in a vector database.
4. The method [300] as claimed in claim 1, wherein the one or more user attributes comprises at least a user preference and a user interest.
5. The method [300] as claimed in claim 1, wherein the one or more dynamic content tags are linked to a progress bar and serve as one or more indexes for the user to navigate within the media content.
6. The method [300] as claimed in claim 1, further comprising:- tracking, by the processing unit [102], the one or more user attributes to update the one or more dynamic content tags in real-time for improving user engagement; and- updating and adjusting, by the processing unit [102], the one or more dynamic content tags based on at least the one or more user attributes and media content progress using a Retrieval- Augmented Generation (RAG) model.
7. The method [300] as claimed in claim 5, further comprises:training the RAG model to prioritize a dynamic content tag from the one or more dynamic content tags based on the one or more user attributes.
8. The method [300] as claimed in claim 4, wherein the one or more dynamic content tags are tailored to specific user cohorts, wherein the user cohorts are dynamically adjusted based on at least: user demographics, viewing history, and the one or more user attributes.
9. The method [300] as claimed in claim 6, wherein the media content progress comprises at least one of: pausing, resuming, forwarding or stopping the media content.
10. The method [300] as claimed in claim 1, wherein the media content comprises at least a video content, an audio content, and a text-based content.
11. The method [300] as claimed in claim 1, wherein the one or more external sources comprises at least third-party databases, online streaming platforms, social media and communication devices of the user.
12. A system [100] for providing automated personalized and dynamic content tags to a user for watching a media content, the system [100] comprising:a processing unit [102] configured to:o extract information related to the media content from one or more external sources; ando process the extracted information using a large language model (LLM) to generate one or more dynamic content tags, wherein the one or more dynamic content tags are personalized based on one or more user attributes;a synchronization unit [104] configured to:o synchronize the one or more dynamic content tags with the media content, wherein synchronization of the one or more dynamic content tags is done to contextually align the one or more dynamic content tags with one or more scenes of the media content; anda display unit [110] configured to:o display the synchronized one or more dynamic content tags with the media content to the user during media content discovery and media content playback, in an event the user pauses at a scene of the media content.
13. The system [100] as claimed in claim 12, the synchronization unit [104] is further configured to:analyze subtitles, audio descriptions and scene segmentation data;align the one or more dynamic content tags with the one or more scenes of the media content based on the analysis.
14. The system [100] as claimed in claim 12, the generated one or more dynamic content tags are stored in a vector database.
15. The system [100] as claimed in claim 12, wherein the one or more user attributes comprises at least a user preference and a user interest.
16. The system [100] as claimed in claim 12, wherein the one or more dynamic content tags are linked to a progress bar and serve as one or more indexes for the user to navigate within the media content.
17. The system [100] as claimed in claim 12, wherein the processing unit [102] is further configured to:- track the one or more user attributes to update the one or more dynamic content tags in real-time for improving user engagement; and- update and adjust the one or more dynamic content tags based on at least the one or more user attributes and media content progress using a Retrieval-Augmented Generation (RAG) model.
18. The system [100] as claimed in claim 17, wherein the processing unit [102] is further configured to: train the RAG model to prioritize a dynamic content tag from the one or more dynamic content tags based on the one or more user attributes.
19. The system [100] as claimed in claim 14, wherein the one or more dynamic content tags are tailored to specific user cohorts, wherein the user cohorts are dynamically adjusted based on at least: user demographics, viewing history, and the one or more user attributes.
20. The system [100] as claimed in claim 17, wherein the media content progress comprises at least one of: pausing, resuming, forwarding or stopping the media content.
21. The system [100] as claimed in claim 12, wherein the media content comprises at least a video content, an audio content, and a text-based content.
22. The system [100] as claimed in claim 12, wherein the one or more external sources comprises at least third-party databases, online streaming platforms, social media and communication devices of the user.