Systems and methods for facilitating context-based media management
The integration of real-time caption capture and machine learning models in the system addresses the challenge of ensuring brand safety in live-streaming environments by providing timely and accurate ad placement, overcoming the limitations of conventional systems in dynamic live content analysis.
Patent Information
- Application Number
- PCT/US2025/014589
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-11
- Filing Date
- 2025-02-05
- Publication Date
- 2026-04-16
AI Technical Summary
Existing systems fail to provide near real-time content attributes for live-streaming environments, leading to challenges in ensuring brand safety for advertisements, as they cannot effectively analyze and tag live-streaming content on the fly due to the dynamic and unpredictable nature of live broadcasts.
A system that integrates real-time caption capture with machine learning models tuned to specific content types, generating Content Enrichment Platform (CEP) data to ensure timely ad-serving decisions, and efficiently manages metadata in a content management system to maintain brand safety in live-streaming environments.
The system ensures real-time content analysis without slowing down ad-serving processes, providing accurate and timely ad placement, thus maintaining brand safety even in unpredictable live broadcasting scenarios.
Smart Images

Figure US2025014589_16042026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: 00055-0153-00304SYSTEMS AND METHODS FOR FACILITATING CONTEXT-BASED MEDIA MANAGEMENTCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This patent application claims the benefit of priority to U.S. Provisional Application No. 63 / 706,314, filed on October 11 , 2024, the entirety of which is incorporated herein by reference.TECHNICAL FIELD
[0002] Various embodiments of the present disclosure relate generally to the field of online advertising systems, and, more particularly, to systems and methods for ensuring brand safety in the context of live-streaming video content by providing real-time content enrichment data to advertising servers.BACKGROUND
[0003] Advertisers invest significant resources in ensuring their advertisements are displayed alongside content that is aligned with their brand’s values and goals. For instance, advertisers in the news industry have strict brand suitability standards that must be upheld, even in live-streaming environments. Existing systems are unable to provide near real-time attributes about live-streaming content to ad servers for enhanced decision-making. This poses challenges for advertisers and publishers, as there is no way to ensure that advertisements are not shown alongside inappropriate content in real time. The present disclosure is accordingly directed to systems and methods that may provide brand-safe advertising in live-streaming environments.
[0004] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in thisAttorney Docket No.: 00055-0153-00304 application and are not admitted to be prior art, or suggestions of the prior art, by inclusion in this section.SUMMARY OF THE DISCLOSURE
[0005] According to certain aspects of the disclosure, systems and methods are disclosed for providing brand-safe advertising.
[0006] In one aspect, a computer-implemented is provided. The computer- implemented method may include operations including: receiving, at a computer system, caption data associated with a most recent segment of a live media stream; storing, using a processor associated with the computer system, the caption data to a data cache containing previously stored caption data associated with prior segments of the live media stream; providing, using the processor, the caption data to a trained machine learning model responsive to determining that the caption data is relevant to a program broadcast in the live media stream; generating, using the processor and based on output received from the trained machine learning model, a metadata tag to associate with the caption data; and utilizing the metadata tag to identify one or more segments of media content to include in the live media stream during a break in the program.
[0007] In another aspect, a system is provided. The system may include: a memory including instructions; and at least one processor configured to execute the instructions stored in the memory to perform operations comprising: receiving caption data associated with a most recent segment of a live media stream; storing the caption data to a data cache containing previously stored caption data associated with the live media stream; providing the caption data to a trained machine learning model responsive to determining that the caption data is relevant to a program broadcast in the live media stream; generating, based on outputAttorney Docket No.: 00055-0153-00304 received from the trained machine learning model, a metadata tag to associate with the caption data; and utilizing the metadata tag to identify one or more segments of media content to include in the live media stream during a break in the program.
[0008] In yet another aspect, a non-transitory computer-readable medium storing computer-executable instructions is provided. The computer-executable instructions, when executed by a server in network communication with at least one database, cause the server to perform operations including: receiving, at a computer system, caption data associated with a most recent segment of a live media stream; storing, using a processor associated with the computer system, the caption data to a data cache containing previously stored caption data associated with the live media stream; providing, using the processor, the caption data to a trained machine learning model responsive to determining that the caption data is relevant to a program broadcast in the live media stream during the most recent segment; generating, using the processor and based on output received from the trained machine learning model, a metadata tag to associate with the caption data; and utilizing the metadata tag to identify one or more segments of media content to include in the live media stream during a break in the program.
[0009] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various exemplary embodiments and together with the description, serve to explain the principles of the disclosed embodiments.Attorney Docket No.: 00055-0153-00304
[0011] FIG. 1 depicts components of a system of a brand-safe advertising solution, according to one or more embodiments of the present disclosure.
[0012] FIG. 2 depicts an exemplary captions capture and management workflow, according to one or more embodiments of the present disclosure.
[0013] FIG. 3 depicts a diagram corresponding to an exemplary use-case of brand-safe ad insertion, according to one or more embodiments of the present disclosure.
[0014] FIG. 4 presents a process flow for capturing and leveraging caption data to promote brand-safe advertisements, according to one or more embodiments of the present disclosure.
[0015] FIG. 5 presents a diagram according to another exemplary use-case of brand-safe ad insertion, according to one or more embodiments of the present disclosure.
[0016] FIG. 6 presents a diagram according to another exemplary use-case of brand-safe ad insertion, according to one or more embodiments of the present disclosure.
[0017] FIG. 7 presents a diagram according to another exemplary use-case of brand-safe ad insertion, according to one or more embodiments of the present disclosure.
[0018] FIG. 8 presents a diagram according to another exemplary use-case of brand-safe ad insertion, according to one or more embodiments of the present disclosure.
[0019] FIG. 9 depicts an exemplary computing server for enabling offline advertisement consumption and data reconciliation, according to one or more embodiments of the present disclosure.Attorney Docket No.: 00055-0153-00304DETAILED DESCRIPTION OF EMBODIMENTS
[0020] The terminology used below may be interpreted in its broadest reasonable manner, even though it is being used in conjunction with a detailed description of certain specific examples of the present disclosure. Indeed, certain terms may even be emphasized below; however, any terminology intended to be interpreted in any restricted manner will be overtly and specifically defined as such in this Detailed Description section. Both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the features, as claimed.
[0021] In this disclosure, the term “based on” means “based at least in part on.” The singular forms “a,” “an,” and “the” include plural referents unless the context dictates otherwise. The term “exemplary” is used in the sense of “example” rather than “ideal.” The terms “comprises,” “comprising,” “includes,” “including,” or other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, or product that comprises a list of elements does not necessarily include only those elements, but may include other elements not expressly listed or inherent to such a process, method, article, or apparatus. Relative terms, such as, “substantially” and “generally,” are used to indicate a possible variation of ±10% of a stated or understood value.
[0022] The term “user”, “subscriber,” and the like generally encompasses consumers who are subscribed to a streaming service (e.g., streaming platform) associated with the system described herein. The term “streaming service” (e.g., streaming platform) may refer to subscription-based video-on-demand (SVoD) services such as television shows, films, documentaries, and the like. The term “user” may be used interchangeably with “user profile,” “profile,” and the likeAttorney Docket No.: 00055-0153-00304 throughout this application. The phrase “registered with” may be used interchangeably with “subscribed to” and the like throughout this application. The phrase “multimedia content” may be used interchangeably with “multimedia content item,” “article of multimedia content,” “segment of multimedia content,” and the like throughout this application. The terms “advertisement,” “ad,” and / or “commercial” may be used interchangeably to refer to media content occurring between sections of a live media stream.
[0023] In the world of online advertising, brand safety is important for advertisers, especially in sensitive environments, such as news broadcasting. Brand safety refers to ensuring that advertisements are not displayed alongside content that may harm the reputation of the brand. For example, advertisers may want to avoid scenarios where an ad for an airline is shown next to a news segment about a plane crash. In pre-recorded video content, systems have been developed to analyze and tag content with metadata that indicates its “aboutness,” ensuring ads align with the content’s context. However, ensuring brand safety in live-streaming environments remains a challenge due to the real-time nature of the content.
[0024] Currently, there are no systems in place to provide near real-time content data to ad servers during live-streaming broadcasts. Conventional systems used for pre-recorded content rely on various platforms to generate metadata from captions and provide that data to the ad server. In these setups, the metadata includes subject matter, sentiment analysis, and brand safety indicators, which the ad server can use to choose appropriate ads. However, these conventional systems are designed for static content, where there is ample time to analyze and generate data before ads are served. In a live-streaming scenario, such as a breaking newsAttorney Docket No.: 00055-0153-00304 broadcast, the content is continually generated in real time, making it difficult to analyze and tag on-the-fly.
[0025] Attempts to extend the conventional systems to live streams face several issues. First, there is the challenge of managing the context window of the stream. More particularly, live content is dynamic, and the system needs to process new information constantly, meaning the metadata generated must be updated frequently. Second, there is a timing issue in that the system must provide sufficient time for the ad server to make decisions, but the rapid pace of live broadcasts leaves little room for delays. Furthermore, there is an additional challenge of maintaining ad counters and ensuring smooth ad insertion during midroll breaks, as any errors in timing can result in missed or improperly placed ads.
[0026] To address the above-noted problems, the present disclosure describes systems and methods that provide real-time content enrichment data to ad servers specifically for live-streaming content. Effectively, the concepts described herein may integrate live-streaming caption capture with machine learning models that are specifically tuned to the type of content being broadcast (e.g., news channel content, sports broadcast content, etc.). These models may generate Content Enrichment Platform (CEP) data (e.g., “aboutness” tags, sentiment indicators, brand safety metadata, etc.) as the content is being streamed, ensuring that the system can make timely ad-serving decisions. Additionally to the foregoing, the system may be designed to store the CEP data in a content management system (CMS) that may be frequently updated, ensuring that up-to-date metadata is always available. By determining an optimal frequency for updating the metadata, the system may strike a balance between ensuring real-time accuracy and avoiding overwhelming the CMS with too many updates.Attorney Docket No.: 00055-0153-00304
[0027] The approach summarized above and further elaborated upon herein overcomes the issues faced by conventional attempts. It allows real-time content analysis without slowing down the ad-serving process, ensuring that the ads are appropriately timed to match the content being displayed, and provides a scalable solution that can handle the dynamic nature of live streams. By providing advertisers with the necessary real-time data to make informed decisions, the system ensures that brand safety may be maintained even in unpredictable live broadcasting environments.
[0028] The systems and methods described herein represent a variety of technical improvements to computer technology. For instance, the systems’ ability to analyze live-stream content in near real-time and generate CEP data to inform ad decisioning processes represents an advancement in the technical field. In the conventional standard, systems designed for brand safety operate well for prerecorded or static content, where metadata may be generated without stringent time constraints. However, live-streaming environments pose a unique challenge due to the constant flow of data and the need for immediate responses. The concepts described herein improve the technical capability of existing systems by introducing real-time caption capture, machine learning-based analysis, and immediate generation of brand safety indicators. This advancement enhances the system’s ability to keep up with the dynamic and unpredictable nature of live content, addressing the latency and timing issues that would otherwise hinder real-time ad insertion.
[0029] Another technical improvement relates to the system’s ability to manage and store content metadata (e.g., CEP data) efficiently while maintaining up-to-date information for ad servers. In the context of live-streaming, it is importantAttorney Docket No.: 00055-0153-00304 that metadata is updated frequently to reflect the changing content, but excessive updates may overwhelm content management systems (CMS) and lead to performance issues. The concepts described herein introduce a technical optimization by determining an ideal frequency for storing and updating CEP data, which balances the need for real-time accuracy with the system’s operational load. This improvement addresses a common challenge in computer systems - how to manage large volumes of rapidly changing data without causing bottlenecks or system inefficiencies. The system described herein achieves a higher degree of scalability by ensuring that the CMS can handle live updates without degradation in performance.
[0030] Additionally to the foregoing, the system’s use of bespoke machine learning models that are fine-tuned for specific content represents a further improvement in the field of content analysis and brand safety. Unlike generic models, which may not provide the granularity needed to ensure brand safety in live news broadcasts, these custom models may be specifically designed to process the context and sentiment of news content in real-time. This technical advancement improves upon existing machine learning implementations by offering a more precise, domain-specific solution for real-time content analysis.
[0031] Additionally to the foregoing, it is important to note that the concepts presented in this patent application cannot practically be performed in the human mind due to the complexity, scale, and real-time processing requirements inherent to the system. More particularly, a human mind cannot constantly monitor, capture, and process live captions to identify specific subjects, sentiments, and contextual information, especially in a rapidly changing live-stream environment. Additionally, because the systems and methods described herein leverage purposely-trainedAttorney Docket No.: 00055-0153-00304 machine learning models that perform complex functions such as pattern recognition, data classification, and natural language processing (NPL) tasks, each of which involve sophisticated algorithms and massive dataset that are updated and trained over time, the human mind cannot replicate the functions of these algorithms, which are capable of parsing large amounts of textual data and returning highly specific results in a very short period of time (e.g., milliseconds).
[0032] The subject matter of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific exemplary embodiments. An embodiment or implementation described herein as “exemplary” is not to be construed as preferred or advantageous, for example, over other embodiments or implementations; rather, it is intended to reflect or indicate that the embodiment(s) is / are “example” embodiment(s). Subject matter may be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any exemplary embodiments set forth herein; exemplary embodiments are provided merely to be illustrative. Likewise, a reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, or systems. Accordingly, embodiments may, for example, take the form of hardware, software, firmware, or any combination thereof. The following detailed description is, therefore, not intended to be taken in a limiting sense.
[0033] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment” or “in some embodiments” as used herein does not necessarily refer to the same embodiment and the phrase “in anotherAttorney Docket No.: 00055-0153-00304 embodiment” as used herein does not necessarily refer to a different embodiment. It is intended, for example, that claimed subject matter include combinations of exemplary embodiments in whole or in part.
[0034] The terminology used below may be interpreted in its broadest reasonable manner, even though it is being used in conjunction with a detailed description of certain specific examples of the present disclosure. Indeed, certain terms may even be emphasized below; however, any terminology intended to be interpreted in any restricted manner will be overtly and specifically defined as such in this Detailed Description section.
[0035] The system(s) described herein may include one or more processors that are configured to execute instructions to train and / or implement a machine learning model. As used herein, a “machine-learning model” generally encompasses instructions, data, and / or a model configured to receive input, and apply one or more of a weight, bias, classification, or analysis on the input to generate an output. The output may include, for example, an analysis based on the input, a prediction, suggestion, or recommendation associated with the input, a dynamic action performed by a system, or any other suitable type of output. A machine-learning model is generally trained using training data, e.g., experiential data and / or samples of input data, which are fed into the model in order to establish, tune, or modify one or more aspects of the model, e.g., the weights, biases, criteria for forming classifications or clusters, or the like. Aspects of a machine-learning model may operate on an input linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration.
[0036] The execution of a machine-learning model may include deployment of one or more machine-learning techniques, such as k-nearest neighbors, linearAttorney Docket No.: 00055-0153-00304 regression, logistical regression, random forest, gradient boosted machine (GBM), support-vector machine, deep learning, a deep neural network, and / or any other suitable machine-learning technique that solves problems in the field of Natural Language Processing (NLP). Supervised, semi-supervised, and / or unsupervised training may be employed. For example, supervised learning may include providing training data and labels corresponding to the training data, e.g., as ground truth. Unsupervised approaches may include clustering, classification, or the like. K-means clustering or K-Nearest Neighbors may also be used, which may be supervised or unsupervised. Combinations of K-Nearest Neighbors and an unsupervised cluster technique may also be used. Any suitable type of training may be used, e.g., stochastic, gradient boosted, random seeded, recursive, epoch or batch-based, etc.
[0037] Prior to introduction to a machine learning infrastructure, data may be processed and normalized. As used herein, the term “normalize” may refer to the transformation of a value or a set of values to a common frame of reference for comparison purposes. In this regard, one or more normalization algorithms or techniques (e.g., min-max normalization, z-score normalization, decimal scaling, logarithmic transformation, root transformation, etc.) may be leveraged to bring all data attributes in the context data onto the same scale. Such a process may correspondingly improve the performance of the machine learning model by reducing the impact of any outliers and by improving the accuracy of a trained machine learning model associated therewith.
[0038] In some embodiments, a machine-learning model based on neural networks includes a set of variables, e.g., nodes, neurons, filters, etc., that are tuned, e.g., weighted or biased, to different values via the application of training data. In other embodiments, a machine learning model may be based on architectures suchAttorney Docket No.: 00055-0153-00304 as support-vector machines, decision trees, random forests or Gradient Boosting Machines (GBMs). Alternate embodiments include using techniques such as transfer learning, wherein one or more pre-trained machine learning models on large common or domain specific dataset may be leveraged for analyzing the training data.
[0039] In supervised learning, e.g., where a ground truth is known for the training data provided, training may proceed by feeding a sample of training data into a model with variables set at initialized values, e.g., at random, based on Gaussian noise, a pre-trained model, or the like. The output may be compared with the ground truth to determine an error, which may then be back-propagated through the model to adjust the values of the variable.
[0040] Training may be conducted in any suitable manner, e.g., in batches, and may include any suitable training methodology, e.g., stochastic or non-stochastic gradient descent, gradient boosting, random forest, etc. In some embodiments, a portion of the training data may be withheld during training and / or used to validate the trained machine-learning model, e.g., compare the output of the trained model with the ground truth for that portion of the training data to evaluate an accuracy of the trained model. The training of the machine-learning model may be configured to cause the machine-learning model to learn semantic associations between the raw data and the context with which it is associated with (e.g., aspects of the industrial or professional field that the raw data is associated with, etc.), such that the trained machine-learning model is configured to provide output features that are contextually relevant for a user’s purpose.
[0041] In various embodiments, the variables of a machine-learning model may be interrelated in any suitable arrangement in order to generate the output. ForAttorney Docket No.: 00055-0153-00304 example, in some embodiments, the machine-learning model may include signal processing architecture that is configured to identify, isolate, and / or extract features, patterns, and / or structure in a text. For example, the machine-learning model may include one or more convolutional neural network (“CNN”) configured to identify characteristics of caption data obtained from a live media stream, and may include further architecture, e.g., a connected layer, neural network, etc., configured to determine a relationship between the words and phrases to summarize the most important aspects of the captioned content.
[0042] FIG. 1 depicts components of an exemplary system 100 of a brandsafe advertising solution for live-streamed media content, according to one or more aspects of the present disclosure. In an aspect, the components encompassed by this architecture may include each of: client application 105, ad-insertion component 110, ad router 115, hub server 120, context proxy 125, media proxy 130, ML CEP- captions-service (CCC) service 135, CEP service 140, and ad server 145. Although each of the foregoing components is included in FIG. 1 , in other aspects, a subset of the foregoing components may be utilized.
[0043] Client application 105 may correspond to a device or platform where end-users consume live-streamed content, which may include web browsers, connected TV (CTV) devices, mobile applications, and the like. These devices / platforms represent various user interfaces and devices that allow viewers to access live streams, such as desktop and mobile browsers, smart TVs, dedicated streaming devices, and mobile applications running on iOS or Android platforms. In an aspect, client application 105 may be configured to support different types of media, including video streams, captions, and interactive ad components, while maintaining a smooth and responsive user interface. This may involve handlingAttorney Docket No.: 00055-0153-00304 various technical requirements such as buffering, ad insertion points, and user interaction tracking (e.g., skipping ads or interacting with content). In an aspect, when a user accesses a live stream on one of these platforms, client application 105 may be responsible for receiving and displaying the stream in real-time. For instance, the live stream may be embedded within client application 105, and during its consumption, system 100 may ensure that advertisements are seamlessly stitched into the stream during designated ad breaks, without disrupting the user experience.
[0044] Ad-insertion component 110 may correspond to a service that is responsible for managing the dynamic ad insertion (DAI) within the live-streaming content. More particularly, when a user begins watching a live stream on client application 105, ad-insertion component 110 may act as the intermediary service between the live stream and the ad delivery system. It may be configured to monitor the stream in real-time, anticipating upcoming ad breaks and preparing to inject ads into the stream. This timing is important because live streams, unlike pre-recorded content, do not have predetermined start and end points for ad breaks, thereby requiring ad-insertion component 110 to work dynamically to ensure the correct timing for ad requests and insertions. Additionally to the foregoing, ad-insertion component 110 may be configured to stitch the selected ads into the live stream video. This may involve integrating the ads into the video feed such that the transition between live content and the ad feels natural to the user, minimizing buffering or playback interruptions. Ad-insertion component’s ability to integrate with CEP data and other metadata, as further described herein, ensures that the ads it requests are contextually appropriate and aligned with the real-time content of the live stream, thereby enabling brand-safe advertising.Attorney Docket No.: 00055-0153-00304
[0045] Ad router 115 may be a component of system 100 that serves as the intermediary responsible for managing and routing ad requests from the streaming service to the appropriate ad server 145. More particularly, when a live stream reaches an ad break, ad-insertion component 110 may trigger an ad request, which is then passed to ad router 115. Upon receiving the ad request, ad router 115 may communicate with hub server 120 to gather metadata. More particularly, hub server 120 is responsible for interfacing with the underlying content management systems to gather required information, including real-time data about the live stream’s content. In this regard, hub server 120 may fetch this data by, for example, querying the live stream’s video resource data, which may include all relevant content metadata, including CEP parameters. These parameters may be utilized to understand the context of the live stream at any given moment. For example, the CEP data may include tags related to the content’s subject matter (e.g., including keywords, topics, etc.), brand safety information, and / or sentiment analysis, some or all of which may be utilized in determining the suitability of ads to be displayed.
[0046] Hub server 120 may obtain the video resource data by sending a request to context proxy 125. In an aspect, this request may include the live stream’s Media ID, which is a unique identifier that links to the most up-to-date information about the stream. In an aspect, context proxy 125 may function as a layer that efficiently manages the flow of real-time metadata from the content management system to the ad-serving components. It ensures that the data is accessible in a format that can be quickly consumed by hub server 120, and subsequently ad router 115.
[0047] In an aspect, media proxy 130 may be a service responsible for capturing and streaming live captions from the live-streamed content in real-time. ItAttorney Docket No.: 00055-0153-00304 may be configured to continuously operate to gather the textual data associated with a live broadcast, such as captions or subtitle, and then push this data to subscribed services, e.g., the ML CCC service 135, for further processing. Effectively, media proxy 130 provides the foundational raw data (i.e., the live captions) that are analyzed to generate the CEP data, which ultimately determines the suitability of ads based on the context of the live content. In an aspect, media proxy 130 may operate by capturing live-stream captions at predetermined intervals (e.g., every 6 seconds, etc.). This continuous data flow ensures that real-time updates about the live stream’s content are always available for processing by the machine learning models. The API then pushes this live caption data to the ML CCC service 135, which may be responsible for analyzing the text and generating metadata tags that describe the content in terms of “aboutness,” sentiment, and / or brand safety.
[0048] In an aspect, ML CCC service 135 may leverage one or more specifically tailored machine learning models to generate enriched metadata, also known as CEP data. The CEP data includes key elements that inform and guide the selection of ads that are appropriate for the live-streamed content. In an aspect, the machine learning models that are utilized may be fine-tuned to understand the nuances of the live stream’s subject matter, identifying key topics, relevant terms, and even the sentiment conveyed by the content. For instance, the machine learning models may be configured to detect whether the content is of a sensitive nature and attach corresponding brand safety tags that prevent inappropriate advertisements from being displayed along such content. ML CCC service 135 may additionally be configured to output other forms of enriched metadata, such as sentiment tags, which may indicate whether the content is positive, neutral, or negative in tone. After the CEP data is generated by ML CCC service 135, it may be provided to CEPAttorney Docket No.: 00055-0153-00304 service 140, which may be configured to ensure that this data is readily accessible to ad servers. More particularly, CEP service 140 may ensure that this data is stored efficiently and is accessible whenever an ad request is made. The CEP service 140 may continuously update the metadata, saving it at regular intervals (e.g., every minute or as needed) to ensure that the most current information about the live stream is always available.
[0049] Additional details regarding components media proxy 130, ML captions capture service 135, and CEP service 140 (depicted in FIG. 2 as media proxy 210, ML captions capture service 205, and CEP 225, respectively), and the interactions between the same, are further elaborated below in connection with diagram 200 in FIG. 2.
[0050] In an aspect, the video resource data may be returned to ad router 115, which may enrich the ad request with the retrieved CEP data and transmit the enriched ad request to ad server 145. Ad server 145 may leverage the information contained within the CEP data against available advertisements from its inventory. In this regard, ad server 145 takes into account not only the technical specifications of the ad request but also the contextual data derived from the live captions. Once ad server 145 processes the ad request, it may be configured to select the most appropriate ads from its inventory based on a combination of real-time content context, user targeting parameters, and advertiser preferences. It may then send the selected ads back to ad-insertion component 110, which is responsible for stitching the ads into the live stream at the appropriate mid-roll ad breaks. This process may be configured to happen seamlessly and quickly to ensure that viewers experience a smooth transition between live content and ads. In an aspect, ad server 145 mayAttorney Docket No.: 00055-0153-00304 also be configured to handle the reporting and tracking of ad performance, providing feedback on how the ads are served, viewed, and engaged with by the audience.
[0051] Referring now to FIG. 2, workflow 200 represents a captions capture and CEP workflow for managing live-stream captions, enriching them using machine learning tools, and storing the relevant metadata, according to one or more aspects of the present disclosure. In an aspect, the components involved in workflow 200 include each of: ML captions capture service 205, media proxy 210, captions cache 215, enriched flag cache 220, CEP endpoint 225, ML processing component 230, content management system (CMS) 235.
[0052] In an aspect, at step 1 of workflow 200, ML captions capture service 205 may subscribe to media proxy 210, which continuously monitors live-stream events and transmits, at step 2, notifications (e.g., push notifications) whenever new captions or program change events are available. These captions provide text data that describe the content being broadcast and are leveraged later in the workflow where they will be processed and analyzed to generate CEP data. By subscribing to this service, the system ensures that it remains up-to-date with the latest captions being generated from the live stream. Additionally to the foregoing, in an aspect, the notifications may allow the system to track changes in the program schedule, such as when one segment of a show ends and a new one begins, or when advertisements are about to air. This information may be utilized to filter out irrelevant or disruptive caption data, particularly those that occur during ad breaks that may otherwise interfere with the accurate analysis of the live program’s content.
[0053] At step 3, the system may retrieve previously stored captions data from captions cache 215. In an aspect, captions cache 215 may manifest as an object storage service (e.g., such as AMAZON simple storage service (S3)) that actsAttorney Docket No.: 00055-0153-00304 as a continuously updated storage mechanism that holds a rolling history of the captions for a live stream. This rolling captions cache 215 may be designed to store not just the captions themselves, but also additional metadata such as timestamps, upcoming advertisement schedules, and program change information. This may allow the system to maintain context about the live stream, ensuring that new caption data is properly processed in relation to what has already been captured. In an aspect, each time ML captions capture service 205 receives a new batch of captions from media proxy 210, it fetches the corresponding cache from caption cache 215. This cache retrieval process may be utilized to ensure that the system has continuity and access to all the relevant captions from the current live stream. For instance, if the live stream is covering a continuous event, the cache will hold a timeline of the captions, making it possible for the system to later analyze and interpret the sequence of events.
[0054] At step 4, once ML captions capture service 205 has fetched the current rolling captions cache from caption cache 215, it may integrate the new batch of captions it just received from media proxy 210. This integration may involve appending the newly received captions data to the existing cache, ensuring that all relevant information is preserved and accessible for later processing. More particularly, in an aspect, during this update the system may be configured to process the incoming captions to ensure that they meet the criteria for inclusion in the cache. This process, for instance, may involve determining whether the captions are derived from advertisements presented during an advertisement window and / or from another program. Using the advertisement event information stored in captions cache 215, the system may cross-reference the timing of the new captions with the scheduled ad breaks. If the captions were generated during an advertisement break,Attorney Docket No.: 00055-0153-00304 they may be excluded from captions cache 215 to prevent ad-related content from contaminating the live-stream’s program captions, which are utilized for understanding the context of the program itself.
[0055] Additionally to the foregoing, in an aspect, the system may use program change data, also stored in captions cache 215, to decide whether it should remove old captions. More particularly, if a new program has started (e.g., as indicated by program change metadata), captions cache 215 may be cleared of any old caption data, thereby ensuring that captions from a prior program are not mistakenly retained. This process is important because each program segment represents new content, and retaining outdated captions may skew the system’s ability to provide accurate metadata and CEP data. In an aspect, once the new captions are validated, the updated data structure may, at step 5, be saved back to caption cache 215. This ensures that the rolling cache is always up to date and contains only the most relevant captions, ready to be processing for further enrichment.
[0056] At step 6, ML captions capture service 205 may determine whether the current batch of captions should be sent for further enrichment. To facilitate this determination, the ML captions capture service 205 may leverage an “enrich” flag that is stored in another data cache, e.g., enriched flag cache 220, and acts as a trigger that determines whether the system should process the caption data further or pause until the next batch of captions is received. In an aspect, the enrich flag may be fetched each time new caption data is captured, thereby allowing the system to dynamically adjust its behavior based on real-time conditions.
[0057] At step 7, ML captions capture service 205 may evaluate the flag’s current value - either TRUE or FLASE - and take the appropriate action based onAttorney Docket No.: 00055-0153-00304 the flag’s state. For instance, if the flag is TRUE, it means that the system has determined that the current captions are part of meaningful program content that should be analyzed further, and ML captions capture service 205 proceeds with preparing the captions for enrichment. Conversely, if the flag is FALSE, it signals that the current content does not require further processing (e.g., as a result of the content occurring during an ad break, a program change, or another moment where the captions are not programmatically significant). In this case, ML captions capture service 205 may not send the current captions for further enrichment, ensuring that irrelevant or redundant data is filtered out, thereby preventing unnecessary processing. This may help to optimize the workflow by saving system resources and avoiding cluttered or inaccurate metadata. In an aspect, after ML captions capture service 205 evaluates the enrich flag, it may update it in enriched flag cache 220 based on the current conditions. By updating enriched flag cache 220, the system ensures that its most up-to-date decision regarding enrichment is persisted in a reliable and centralized location. This enables other components of the workflow, or even subsequent processes that depend on this data, to access the current state of the flag and make informed decisions based on it. In aspect, after the enrich flag is checked and updated, it may, at step 8, be written back to enrich flag cache 220 to preserve the latest status of the flag.
[0058] In an aspect, when the enrich flag is TRUE, ML captions capture service 205 may, at step 9, take the live caption data and transmit it to CEP 225. At step 10, CEP 225 may format the received captions data into a structured format that adheres to ML processing component’s 230 input requirements to ensure that it effectively process the caption data. More particularly, the raw caption data received from the live stream includes text, timestamps, and metadata, but in its original form,Attorney Docket No.: 00055-0153-00304 so it may not be readily usable by ML processing component 230. To resolve this, CEP 225 may restructure the caption data into a format that is compatible with ML processing component’s 230 analysis algorithms. This process may involve organizing the captions into a well-defined data structure (e.g., such as XML, JSON, etc.) to ensure that ML processing component 230 can interpret the content correctly. This data structure may include one or more elements such as caption text (e.g., the actual words spoken during the live stream), timestamps (e.g., markers that indicate when each caption was spoken, ensuring that the analysis is aligned with the correct moment in the live stream), and / or program metadata (e.g., contextual information such as the current program segment or event, etc.). The foregoing process is important because it ensures that ML processing component 230 receives the data in a structured and consistent manner, allowing its algorithms to accurately extract relevant terms.
[0059] Once the data is formatted, it may be transmitted, at step 11 , to ML processing component 230 for processing. In an aspect, the role of ML processing component 230 is to identify key topics, entities, and / or other contextual elements from the caption data, which it may return as a set of related terms. These terms, which describe the “aboutness” of the content, are relied on for enriching the live- stream metadata. To facilitate this, ML processing component 230 may leverage natural language processing (NLP) algorithms to parse the caption text and identify relationships between the words and phrases. It may also assess the broader context in which the captions were generated, understanding both the explicit content (e.g., such as names of people, please, and things) and the implicit themes (e.g., such as political discourse, entertainment, or global events) from the caption data.Attorney Docket No.: 00055-0153-00304
[0060] In an aspect, the results returned from ML processing component 230, also at step 11 , may include a set of terms or tags that summarize the most important aspects of the captioned content. For instance, if the live stream involves a news report about climate change, ML processing component 230 may return terms such as “global warming,” “carbon emissions,” or “environmental policy.” These terms may be highly valuable because they provide an understanding of the topics being discussed in the live stream at any given moment. In an aspect, after receiving the analyzed data from ML processing component 230, the system may further manipulate the data into a more refined form, referred to herein as a CEP tag. This may involve taking the raw terms output by ML processing component 230 and structuring them in a way that fits the needs of downstream processes, such as advertising systems and content management platforms. For example, the terms may be categorized into different levels of importance or grouped according to specific themes like brand safety, sentiment, or user targeting preferences.
[0061] At step 12, once the interaction with ML processing component 230 is complete and CEP 225 has generated the enriched metadata (e.g., the CEP tags), the system may evaluate the current state of the live stream to determine whether future batches of captions require further enrichment. CEP 225 may determine whether to update enrich flag in enrich flag cache 220 based on one or more factors. For instance, if the content being streamed remains relevant for further analysis, such as a continuous new event or a discussion that needs to be categorized, the enrich flag in enriched flag cache stays TRUE, ensuring that upcoming captions are sent to CEP for enrichment. As another example, during an ad break, the system may update the enrich flag to FALSE, indicating that captions occurring during advertisements do not need to be enriched, as they are not part of the core programAttorney Docket No.: 00055-0153-00304 content. In yet another aspect, if the required metadata has already been generated for a certain segment of the live stream and no further enrichment is needed, the system may set the enrich flag to FALSE to pause further processing until the next program or relevant segment begins. Effectively, the foregoing updating process discussed in association with step 12 ensures that the system is continually making decisions based on real-time conditions, adjusting the enrichment workflow dynamically as the content evolves. By updating the enrich flag, the system avoids redundant processing of captions during periods where enrichment is unnecessary (e.g., such as during repetitive content or during non-program segments, etc.) and conserves resources for more critical content.
[0062] CEP 225 may then, at step 13, transmit the CEP tags to CMS 235 where they may be stored alongside the content as metadata. In an aspect, once the tags are stored, they may be accessed by other systems involved in the ad delivery and content management processes. For instance, when an ad server (e.g., such as ad server 145 in FIG. 1 ) needs to make decisions about which ads to show during a live stream, it may retrieve the CEP tags from CMS 235 to ensure that the selected ads are contextually relevant and brand-safe. Moreover, by storing the CEP tags in CMS 235, the metadata becomes persistent and may be accessed or updated as needed, even after the live stream has ended. In an aspect, this may be valuable in the performance of various downstream actions, e.g., in the generating of reports, conducting of post-event analysis, or supporting on-demand versions of the content where enriched metadata may continue to enhance the user experience.
[0063] Referring now to FIG. 3, diagram 300 illustrates a practical use case of brand-safe ad insertion during a live stream. Specifically, diagram 300 outlines how the system may leverage a rolling window to ensure that ads are appropriatelyAttorney Docket No.: 00055-0153-00304 selected and avoid being placed alongside potentially sensitive or controversial topics discussed in a live media stream. The implementation of the concepts represented by diagram 300 in FIG. 3 may be facilitated using some or all of the components and processes described above in relation to FIGS. 1 and 2.
[0064] In the scenario depicted in FIG. 3, a major news network may broadcast a live show 305. Throughout the broadcast, the system may be configured to monitor the captions of live show 305 in a rolling window 310, which may enable the system to evaluate the recent span of live show 305 (e.g., the most recent 10- minute span), ensuring that advertisement decisions made during break 315 in live show 305 are based on the most current and relevant content. Initially, live show 305 may include a first segment 320, e.g., that is associated with providing weather updates. As first segment 320 progresses, the system may determine that no sensitive topics are being discussed and may classify the content as being safe for general advertisements. However, first segment 320 may be interrupted by breaking news that is embodied in second segment 325, which contains reporting about a plane crash. The system, having captured and analyzed the live captions, may recognize that the topic has changed to a sensitive subject. It may flag the content as not suitable for certain ads, such as those from airlines or travel companies. The system may therefore ensure that no airline ads or travel-related promotions are shown during break 315, which immediately proceeds coverage of the plane crash in second segment 325, thereby protecting brand safety. Instead, appropriate ads, such as those for non-sensitive products (e.g., tech gadgets or household goods), may be displayed.
[0065] Referring now to FIG. 4, an exemplary flow 400 is described for identifying media content to include in a live media stream during a break in a mediaAttorney Docket No.: 00055-0153-00304 program. Aspects of the exemplary flow 400 may be performed in accordance with some or all components described in FIGS. 1 and 2.
[0066] At step 405, a computer system may capture caption data from a live media stream. This caption data may be associated with the most recent segment of the stream and may be collected continuously as dictated by a predetermined collection window, e.g., every 6 seconds. The live stream’s captions may provide textual descriptions or transcriptions of the spoken content in the video. These captions may be utilized for analyzing the content of the stream in near real-time, thereby enabling the system to make informed decisions about brand safety and ad placement.
[0067] At step 410, once the caption data is received, the computer system may store it in a data cache that contains previously captured caption data from the live media stream. In an aspect, this cache may be continuously updated with new caption data, creating a rolling window of both current and past captions that may be referenced for analysis. This rolling cache enables the system to have context about the entire live stream, not just the most recent segment. The cache acts as temporary storage before the data is processed further, ensuring that all recent and relevant caption data is readily available for analysis.
[0068] At step 415, after storing the data in the cache, the system may assess whether the caption data is relevant to the program being broadcast in the live stream. If determined at step 415 to be irrelevant, then the caption data may, at step 420, be discarded. Conversely, if the caption data is determined at step 415 to be relevant, the caption data may, at step 425, be provided to a trained machine learning model. In an aspect, the machine learning model is specifically trained to understand and analyze content from the live media stream (e.g., where the liveAttorney Docket No.: 00055-0153-00304 media stream is a live broadcast channel associated with news, sports, etc.). This model may determine the “aboutness” of the content, identifying key themes, topics, and even sentiment. In an aspect, prior to introduction to the trained machine learning model, the caption data may be formatted to facilitate better processing by the model, as previously described above. In some aspects, step 415 may be performed before or concurrently with step 410. More particularly, before the caption data is stored to data cache at step 410, the relevance of the caption data to a particular broadcast may be determined and the caption data may only be stored to the data cache if it is relevant.
[0069] At step 430, once the machine learning model processes the caption data, it may generate output that is used by the computer system to create a metadata tag. This metadata tag is designed to reflect the key attributes of the content as identified by the machine learning model. For example, the tag may indicate the subject matter (e.g., a plane crash) and / or the corresponding sentiment (e.g., negative). This metadata tag may provide a way to classify the content of the live stream segment and is important for ensuring that only appropriate advertisements or media content are associated with events occurring in the broadcast program.
[0070] At step 435, the metadata tag may thereafter be leveraged to select media content for inclusion in the live stream during an upcoming break. More particularly, the system may utilize the metadata tag to ensure that a segment of media content (e.g., such as advertisements, promotional content, etc.) may be appropriate for the context of the program. For instance, if the metadata tag indicates that the live segment is discussing a tragic event, the system will ensure that certain types of advertisements that share some context with the event (e.g., that areAttorney Docket No.: 00055-0153-00304 associated with products or objects involved in the event, that are associated with people involved in the event, that are associated with a location where the event occurred, etc.) are not displayed during the break. Effectively, the metadata tag guides the ad decisioning process, ensuring brand safety and contextual relevance for the media content shown in the live stream.
[0071] Diagrams 500 - 800 in FIGS. 5 - 8 represent a plurality of different implementations of the concepts described herein that are aimed at preserving brand safety. Each of these diagrams is discussed in greater detail below.
[0072] Referring now to FIG. 5, diagram 500 illustrates a non-limiting example of how brand-safe advertisements may be inserted into a live media stream. The implementation of the concepts represented by diagram 500 in FIG. 5 may be facilitated using some or all of the components and processes described above in relation to FIGS. 1 and 2.
[0073] In an aspect, diagram 500 demonstrates the operational flow of ad insertion, where a brief interval of time 525 (e.g., sometimes referred to as a “bumper”) is introduced into a break 510 of a live media stream 505 prior to the actual presentation of advertisements 515, 520. This provides a smooth transition between the live media stream 505 and the advertisements 515, 520 presented during the break 510. In an aspect, the interval of time 525 may be virtually any time interval (e.g., 5 seconds, 6 seconds, etc.). In some aspects, the length of the interval of time 525 may be based on the sensitivity of the content presented during the live media stream 505. For instance, if live media stream 505 was a news broadcast that presented a breaking story that was especially sensitive (e.g., likely to trigger an emotional reaction from many viewers), the interval of time 525 may be set to be longer than if less sensitive content was broadcast. In an aspect, Society of CableAttorney Docket No.: 00055-0153-00304Telecommunications Engineers (“SCTE”) signals 530 may be used to mark key points in the broadcast, such as the start of the advertisement break 510. In accordance with the processes represented in diagram 500, the SCTE signal 530 may trigger the backend components (e.g., ad-insertion component 110) to request ads from the ad server and to stitch these ads into the live media stream 505 during the break 510 subsequent to the presentation of the bumper interval of time 525.
[0074] Referring now to FIG. 6, diagram 600 illustrates another non-limiting example of how brand-safe advertisements may be inserted into a live media stream. The implementation of the concepts represented by diagram 600 in FIG. 6 may be facilitated using some or all of the components and processes described above in relation to FIGS. 1 and 2.
[0075] In an aspect, diagram 600 demonstrates the operational flow of ad insertion, where a human operator, e.g., a producer, plays a role in triggering the ad event shortly before a commercial break. As shown in diagram 600, a live media stream 605 may be interrupted by break 610, which may contain one or more advertisements, e.g., advertisements 615, 620, and 625. A predetermined period of time 635 prior to break 610 (e.g., 5 seconds, 6 seconds, n seconds, etc.), human operator 630 may initiate an ad event trigger. This lead time may provide the system with enough time to handle the necessary processes for ad selection and ensure that the ads chosen are appropriate based on the context of the live media stream 605. In this regard, the system may assess the recent content presented in the live media stream 605 and ensure that ads 615, 620, and 625 shown are appropriate for the context of the recent portions of the live media stream 605. In an aspect, each of advertisements 615, 620, and 625 may be made brand safe, e.g., each may be associated with content that is not insensitive or inappropriate given the recentAttorney Docket No.: 00055-0153-00304 context of the live media stream 605. In another aspect, only the first ad of a series (e.g., advertisement 615) of ads 615, 620, and 625 may be ensured to be brand safe. This may enable more freedom in ad selection during the duration of break 610, while still ensuring that the advertisement with the closest proximity to the most recently presented content in the live media stream 605 is not inappropriate or insensitive.
[0076] In yet another aspect, the system may be capable of anticipating the content that is likely to be featured after break 610 concludes and the live media stream 605 resumes. For instance, if prior to break 610 the live media stream 605 was featuring a breaking news story about a natural disaster, the system may be configured to anticipate that the same story will be featured after break 610 concludes. Accordingly, the system may ensure that the advertisements occurring directly proximate to the live media stream 605 (e.g., advertisements 615 and 625 occurring at the initiation and conclusion of break 610, respectively) are not contextually inappropriate or insensitive, while still having some ad-selection freedom with respect to the non-proximate advertisements.
[0077] Referring now to FIG. 7, diagram 700 illustrates another non-limiting example of how brand-safe advertisements may be inserted into a live media stream. The implementation of the concepts represented by diagram 700 in FIG. 7 may be facilitated using some or all of the components and processes described above in relation to FIGS. 1 and 2.
[0078] In an aspect, diagram 700 represents the operational flow of ad insertion by leveraging advertisement pods (“pods”). More particularly, as shown in diagram 700, a live media stream 705 may be interrupted by break 710, which may contain one or more advertisements, e.g., advertisements 715, 720, and 725. TheseAttorney Docket No.: 00055-0153-00304 advertisements may be divided and grouped into distinct pods (e.g., ad pod 1 730 and ad pod 2 735), where some pods may contain advertisements that meet brand safety standards and others may not. This division of ads into pods may offer greater flexibility in managing the types and placement of ads shown during the break 710. For instance, ad pod 1 730 may be configured to contain one or more “house” ads, which are advertisements and / or other types of media (e.g., a diagram of upcoming weather in a region, images of anchors during a news broadcast, etc.) that typically promote the network’s own content or services. In situations where the content being discussed in the live stream is too sensitive for external advertisers, the system may automatically be configured to play one or more ads 715 from ad pod 1 730 at the start of break 710 to ensure that something relevant is still shown. Thereafter, the system may present one or more ads from ad pod 2 735 that it has determined are contextually appropriate, before resuming the live media stream 705 at the conclusion of the break 710.
[0079] Referring now to FIG. 8, diagram 800 illustrates another non-limiting example of how brand-safe advertisements may be inserted into a live media stream. The implementation of the concepts represented by diagram 700 in FIG. 7 may be facilitated using some or all of the components and processes described above in relation to FIGS. 1 and 2.
[0080] In an aspect, diagram 800 represents the operational flow of ad insertion by leveraging the use of a delayed live point. More particularly, live media stream 805 may contain a delay 830 between the time the live content is recorded and when it is broadcast to viewers to provide the system with a window to process real-time data, analyze the live content, and select the appropriate, brand-safe advertisements. In an aspect, delay 830 may be virtually any predetermined lengthAttorney Docket No.: 00055-0153-00304(e.g., 6 seconds, 30 seconds, 1 minute, etc.) but may be configured to be minimal enough to ensure the viewer’s experience feels live, while still being significant enough to give the backend components time to process the content and select the most contextually appropriate ads, e.g., advertisements 815, 820, and 825, to introduce during a break 810 in the live media stream 805. Effectively, delay 830 allows the system to ensure that ad decisions are based on the most accurate and up-to-date context of the live media stream 805.
[0081] In general, any process discussed in this disclosure that is understood to be computer-implementable, such as the processes illustrated in FIG. 9, may be performed by one or more processors of a computer system, such as computer system 100 described above. A process or process step performed by one or more processors may also be referred to as an operation. The one or more processors may be configured to perform such processes by having access to instructions (e.g., software or computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The instructions may be stored in a memory of the computer server. A processor may be a central processing unit (CPU), a graphics processing unit (GPU), or any suitable types of processing unit.
[0082] A computer system, such computer system 100, may include one or more computing devices. If the one or more processors of the computer systems are implemented as a plurality of processors, the plurality of processors may be included in a single computing device or distributed among a plurality of computing devices. If any computer systems 100 comprises a plurality of computing devices, the memory of the computer systems 100 may include the respective memory of each computing device of the plurality of computing devices.Attorney Docket No.: 00055-0153-00304
[0083] FIG. 9 is a simplified functional block diagram of a computer system 900 that may be configured as a computing device for executing any of the processes FIGS. 1 - 8, according to exemplary embodiments of the present disclosure. FIG. 9 is a simplified functional block diagram of a computer that may be configured as the computer system 100 according to exemplary embodiments of the present disclosure. In various embodiments, any of the systems herein may be an assembly of hardware including, for example, a data communication interface 920 for packet data communication. The platform also may include a central processing unit (“CPU”) or processor 902, in the form of one or more processors, for executing program instructions. The platform may include an internal communication bus 908, and a storage unit 906 (such as ROM, HDD, SDD, etc.) that may store data on a computer readable medium 922, although the system 900 may receive programming and data via network communications via electronic network 925 (e.g., voice, video, audio, images, or any other data over the electronic network 925). The system 900 may also have a memory 904 (such as RAM) storing instructions 924 for executing techniques presented herein, although the instructions 924 may be stored temporarily or permanently within other modules of system 900 (e.g., processor 902 and / or computer readable medium 922). The system 900 also may include input and output devices 912 and / or a display 910 to connect with input and output devices such as keyboards, mice, touchscreens, monitors, displays, etc. The various system functions may be implemented in a distributed fashion on a number of similar platforms, to distribute the processing load. Alternatively, the systems may be implemented by appropriate programming of one computer hardware platform.
[0084] Program aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of executable code and / orAttorney Docket No.: 00055-0153-00304 associated data that is carried on or embodied in a type of machine-readable medium. “Storage” type media include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer of the mobile communication network into the computer platform of a server and / or from a server to the mobile device. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links, or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0085] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the disclosure being indicated by the following claims.
[0086] In general, any process discussed in this disclosure that is understood to be performable by a computer may be performed by one or moreAttorney Docket No.: 00055-0153-00304 processors. Such processes include, but are not limited to: the process shown in FIGS. 2 and 4, and the associated language of the specification. The one or more processors may be configured to perform such processes by having access to instructions (computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The one or more processors may be part of a computer system (e.g., one of the computer systems discussed above) that further includes a memory storing the instructions. The instructions also may be stored on a non-transitory computer-readable medium. The non-transitory computer-readable medium may be separate from any processor. Examples of non-transitory computer-readable media include solid-state memories, optical media, and magnetic media.
[0087] It should be appreciated that in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment. Thus, the claims following the Detailed Description are hereby expressly incorporated into this Detailed Description, with each claim standing on its own as a separate embodiment of this invention.
[0088] Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and formAttorney Docket No.: 00055-0153-00304 different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0089] Thus, while certain embodiments have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as falling within the scope of the invention. For example, functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present invention.
[0090] The above disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other implementations, which fall within the true spirit and scope of the present disclosure. Thus, to the maximum extent allowed by law, the scope of the present disclosure is to be determined by the broadest permissible interpretation of the following claims and their equivalents, and shall not be restricted or limited by the foregoing detailed description. While various implementations of the disclosure have been described, it will be apparent to those of ordinary skill in the art that many more implementations are possible within the scope of the disclosure. Accordingly, the disclosure is not to be restricted except in light of the attached claims and their equivalents.
Claims
Attorney Docket No.: 00055-0153-00304What is claimed is:
1. A computer-implemented method, the computer-implemented method comprising operations including: receiving, at a computer system, caption data associated with a most recent segment of a live media stream; storing, using a processor associated with the computer system, the caption data to a data cache containing previously stored caption data associated with prior segments of the live media stream; providing, using the processor, the caption data to a trained machine learning model responsive to determining that the caption data is relevant to a program broadcast in the live media stream; generating, using the processor and based on output received from the trained machine learning model, a metadata tag to associate with the caption data; and utilizing the metadata tag to identify one or more segments of media content to include in the live media stream during a break in the program.
2. The computer-implemented method of claim 1 , wherein the most recent segment corresponds to a recent predetermined length of the live media stream.
3. The computer-implemented method of claim 1 , wherein the storing comprises:Attorney Docket No.: 00055-0153-00304 determining, using the processor, whether the caption data was generated during an advertisement window or another program broadcast during the live media stream; and storing the caption data to the data cache responsive to determining that the caption data was not generated during the advertisement window or the another program.
4. The computer-implemented method of claim 1 , further comprising removing at least a subset of the stored caption data responsive to determining that the subset of the stored caption data is no longer associated with a current program broadcast during the live media stream.
5. The computer-implemented method of claim 1 , wherein the providing comprises: transmitting the caption data to a content enrichment platform prior to provision to the trained machine learning model; formatting, utilizing the content enrichment platform, the caption data to a structured format that is compatible for processing by the trained machine learning model; and providing the formatted caption data to the trained machine learning model.
6. The computer-implemented method of claim 1 , wherein the trained machine learning model is trained using training data optimized for the live media stream.Attorney Docket No.: 00055-0153-003047. The computer-implemented method of claim 1 , wherein the utilizing the metadata tag comprises: identifying a context associated with the most recent segment of the live media stream using the metadata tag; identifying a sensitivity category of the most recent segment of the live media stream based on the identified context; and determining the one or more segments of media content to include in the live media stream during the break based on the identified sensitivity.
8. The computer-implemented method of claim 7, wherein the sensitivity category corresponds to a low sensitivity category or a high sensitivity category.
9. The computer-implemented method of claim 7, wherein the determining the one or more segments of media content comprises: identifying subject matter of each of the one or more segments of media content; comparing the identified subject matter to the context associated with the most recent segment of the live media stream; and precluding any of the one or more segments of media content from being included in the live media stream during the break if: i) the sensitivity category corresponds to a high sensitivity category; and ii) the identified subject matter is associated with the context.
10. The computer-implemented method of claim 9, further comprising including in the live media stream during the break a subset of the one or moreAttorney Docket No.: 00055-0153-00304 segments of media content, wherein the subject matter of each media segment in the subset is not associated with the context associated with the most recent segment of the live media stream.11 . A system comprising: a memory including instructions; and at least one processor configured to execute the instructions stored in the memory to perform operations comprising: receiving caption data associated with a most recent segment of a live media stream; storing the caption data to a data cache containing previously stored caption data associated with the live media stream; providing the caption data to a trained machine learning model responsive to determining that the caption data is relevant to a program broadcast in the live media stream; generating, based on output received from the trained machine learning model, a metadata tag to associate with the caption data; and utilizing the metadata tag to identify one or more segments of media content to include in the live media stream during a break in the program.
12. The system of claim 11 , wherein the most recent segment corresponds to a recent predetermined length of the live media stream.
13. The system of claim 11 , wherein the operations for storing further comprise:Attorney Docket No.: 00055-0153-00304 determining, using the processor, whether the caption data was generated during an advertisement window or another program broadcast during the live media stream; and storing the caption data to the data cache responsive to determining that the caption data was not generated during the advertisement window or the another program.
14. The system of claim 11 , wherein the operations further comprise: removing at least a subset of the stored caption data responsive to determining that the subset of the stored caption data is no longer associated with a current program broadcast during the live media stream.
15. The system of claim 11 , wherein the operations for providing comprise: transmitting the caption data to a content enrichment platform prior to provision to the trained machine learning model; formatting, utilizing the content enrichment platform, the caption data to a structured format that is compatible for processing by the trained machine learning model; and providing the formatted caption data to the trained machine learning model.
16. The system of claim 11 , wherein the trained machine learning model is trained using training data optimized for the live media stream.
17. The system of claim 11 , wherein the operations for utilizing the metadata tag further comprise:Attorney Docket No.: 00055-0153-00304 identifying a context associated with the most recent segment of the live media stream using the metadata tag; identifying a sensitivity category of the most recent segment of the live media stream based on the identified context; and determining the one or more segments of media content to include in the live media stream during the break based on the identified sensitivity.
18. The system of claim 17, wherein the operations for determining the one or more segments of media content further comprise: identifying subject matter of each of the one or more segments of media content; comparing the identified subject matter to the context associated with the most recent segment of the live media stream; and precluding any of the one or more segments of media content from being included in the live media stream during the break if: i) the sensitivity category corresponds to a high sensitivity category; and ii) the identified subject matter is associated with the context.
19. The system of claim 18, wherein the operations further comprise: including in the live media stream during the break a subset of the one or more segments of media content, wherein the subject matter of each media segment in the subset is not associated with the context associated with the most recent segment of the live media stream.Attorney Docket No.: 00055-0153-0030420. A non-transitory computer-readable medium storing computer-executable instructions which, when executed by a server in network communication with at least one database, cause the server to perform operations comprising: receiving, at a computer system, caption data associated with a most recent segment of a live media stream; storing, using a processor associated with the computer system, the caption data to a data cache containing previously stored caption data associated with the live media stream; providing, using the processor, the caption data to a trained machine learning model responsive to determining that the caption data is relevant to a program broadcast in the live media stream during the most recent segment; generating, using the processor and based on output received from the trained machine learning model, a metadata tag to associate with the caption data; and utilizing the metadata tag to identify one or more segments of media content to include in the live media stream during a break in the program.
Citation Information
Patent Citations
System and method for searching contents in accordance with advertisements
KR102411095B1
Closed-caption processing using machine learning for media advertisement detection
US10958982B1
Targeted advertising by context of media content
US20110179445A1
Implementing moments detected from video and audio data analysis
US20220335719A1
Managing metadata enrichment of digital asset portfolios
US20240185302A1