Adaptive dynamic virtual product placement stitching for playback
By integrating virtual product placements directly into video streams with real-time rendering and encoding, the challenges of resource-intensive and disruptive ad insertion are addressed, achieving seamless and personalized advertising across platforms with reduced costs and improved playback quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- RYFF INC
- Filing Date
- 2026-01-14
- Publication Date
- 2026-07-23
AI Technical Summary
Conventional methods for integrating advertisements into video content are resource-intensive, time-consuming, and often disrupt playback quality, failing to achieve seamless integration and real-time adaptation to live campaigns or viewer preferences, leading to increased storage costs and operational complexity.
Integrate virtual product placements directly into video streams as part of the original scene, dynamically rendering and encoding only specific frames with real-time synchronization to match platform specifications, eliminating the need for full media file re-rendering and re-encoding.
Enables seamless, non-interruptive advertising with reduced processing overhead and storage costs, allowing real-time personalization and frame-accurate synchronization across multiple streaming platforms.
Smart Images

Figure US2026011266_23072026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: 301370400013 W000(20)ADAPTIVE DYNAMIC VIRTUAL PRODUCT PLACEMENT STITCHING FOR PLAYBACKCROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims priority to and the benefit of U. S. Provisional Application 63 / 746,568, filed on January 17, 2025, the entire contents of which are incorporated herein as a part of the specification.TECHNICAL FIELD
[0002] The present disclosure relates to modifying video streams to display advertisements, and more particularly to integrating virtual product placements (VPPs) into video streams, enabling seamless advertising experiences through real-time rendering and adaptive frame- accurate encoding.BACKGROUND
[0003] Conventional methods for inserting advertisements into video content, such as replacing entire media files or server-side ad insertion, are resource-intensive, time-consuming, and often disrupt playback quality. These approaches either require re-encoding and redistributing massive video files for even minor changes, or fail to integrate seamlessly with the original video content, leading to noticeable mismatches in visual and audio quality, for example, blurry images or lossy audio.
[0004] Additionally, maintaining multiple versions of mezzanine files for different ad campaigns significantly increases storage costs and operational complexity. Real-time adaptation to live campaigns or viewer preferences is unfeasible under these traditional workflows. Some of the conventional methods are described below.Attorney Docket No.: 301370400013 W000(20)
[0005] Traditional Encoding Pipelines
[0006] Some companies replaced entirely the original files with newly generated files using traditional encoding pipelines. This approach is used by major streaming platforms such as Amazon Prime Video and Netflix, and relies on workflows that process entire mezzanine files. These mezzanine files, which contain all media tracks such as video, audio, and subtitles, are encoded into multiple resolutions and bitrates to create adaptive streaming ladders. Altering even a single segment within the content typically requires regenerating the mezzanine file and re-encoding all variations, which is resource-intensive and time-consuming. This approach becomes inefficient, as even minor changes demand full workflow execution across all platforms.
[0007] Server-Side Ad Insertion (SSAI)
[0008] Server-Side Ad Insertion (SSAI) includes technologies such as Amazon Web Services (AWS) MediaTailor. SSAI modifies media manifests to insert ads into streams as pre-rolls, mid-rolls, or post-rolls. While this approach enables dynamic ad insertion, it is fundamentally limited in several ways. SSAI focuses on adding external ad content rather than modifying the primary content itself. As a result, it does not guarantee frame-accurate synchronization or visual continuity with the original content. Playback often suffers from noticeable transitions or mismatched attributes, such as variations in colorspace, bitrate, or audio levels.
[0009] Hybrid SSAI with Mezzanine Replacement
[0010] Some companies attempt a hybrid approach by creating new mezzanine files with advertising content and encoding them into adaptive ladders. This hybrid approach maintains different versions of the same content and generates a media manifest at playback. While this approach allows for some level of content replacement, it comes with significant challenges. Maintaining multiple versions of mezzanine files for different variations drastically increases storage costs and operational complexity. Additionally, this approach is inherently static, asAttomey Docket No.: 301370400013 W000(20) content updates are tied to batch processing workflows, making real-time adaptation unfeasible.
[0011] Another example is Amagi THUNDERSTORM which facilitates SSAI, enabling personalized ads to be stitched into video streams. Amagi THUNDERSTORM integrates with third-party ad networks and supports pre-, mid- and post-roll ads which are not frame accurate and are simply inserted as separate video content into the playing video by visibly splitting the video content. Its focus is on inserting pre-encoded ads rather than dynamically rendering and synchronizing altered video segments at playback.
[0012] Client-Side Ad Insertion
[0013] Client-side ad insertion, commonly used for mobile and web platforms, relies on the player to download and display ads as separate video files. This approach often leads to inconsistent experiences, with mismatches in visual and audio quality between the main content and the ads. Device-specific limitations further exacerbate playback issues, and personalization remains constrained to pre-selected ad libraries.SUMMARY
[0014] In contrast to the approaches described above, the systems and methods described in the present disclosure integrate VPPs directly into video streams as part of the original scene, offering a seamless, non-interruptive advertising format. Unlike traditional ad breaks, where separate video ads disrupt the main content, the present systems and methods ensure that VPPs (e.g., brand elements such as billboards, logos, designs products, etc.) appear naturally within the video as if filmed that way. Instead of altering the entire video asset, the present systems and methods enable only the modification of specific frames with VPPs. The modified frames are presented at playback while matching the platform specification. This approach eliminates the need to redeliver entire media assets, which may often exceed 0.5 TB, or to reencode entire media assets, which may take 1-3 days. By dynamically rotating brands for the same placements, the present systems and methods enable personalized ad delivery withoutAttorney Docket No.: 301370400013 W000(20) requiring updates to the original media files or incurring the cost of unnecessary full-file encoding processes.
[0015] The present systems and methods are distinguished from other technologies employed by self-managed companies, such as Amazon, that utilize adaptive encoding and timeframe alignment for altering video content. While these companies may reduce computational overhead and turnaround time by altering only specific segments of a video rather than reencoding the entire media source, their approach often focuses on producing alternate prerendered versions of the content. In contrast, the present systems and methods dynamically integrate virtual product placements into video streams in real-time. This feature eliminates the need for pre-rendered alternate versions and allows real-time modifications based on live marketing campaigns and viewer-specific attributes. The following are key advantages of the present systems and methods.
[0016] Selective Rendering and Encoding
[0017] The present systems and methods dynamically render and encode in real-time only the specific video segments requiring brand integrations, matching the platform's original video attributes (e.g., resolution, bitrate, codec profile, etc.), which eliminates the need for full media file re-rendering and re-encoding (for example, by altering a mezzanine file) and significantly reduces processing overhead.
[0018] Frame-Accurate Synchronization
[0019] The present systems and methods achieve frame-level synchronization to ensure seamless playback of video streams. Frame-level synchronization is critical for achieving seamless playback, especially when integrating altered fragment frames into existing video streams. This synchronization is critical for maintaining visual and temporal continuity, preventing glitches, playback disruptions, or mismatches between video and audio.Attorney Docket No.: 301370400013 W000(20)
[0020] A critical aspect of this process is key frame alignment. Key frames, also know n as I-frames in video compression, are standalone frames that contain all the necessary data to display a complete image without relying on other frames. Key frames serve as reference points for video players, enabling random access, video seeking, and proper playback of dependent frames. 1-frames are part of a Group of Pictures (GOP) structure, which also includes predictive frames (P-frames) and bidirectional frames (B-frames) that depend on adjacent frames to reconstruct the video. To ensure seamless integration, the GOP structure of the VPP fragments must match that of the original video. For instance, if the original video employs a GOP structure tike IBBPBBT, the VPP fragments must adopt the same pattern, maintaining decoding continuity for the video player.
[0021] Synchronization also involves matching encoding attributes, such as resolution, bitrate, colorspace, and codec profile to guarantee visual consistency and avoid playback quality degradation. The present systems and methods dynamically extract these encoding attributes from the original segments of the media manifest at playback and applies them to the VPP fragment frames during encoding.
[0022] Another critical aspect is handling video and audio offsets, which are timing shifts between video and audio streams. If left unaddressed, these offsets can result in desynchronized playback where audio leads or lags behind the video, disrupting the viewer experience. To resolve this issue, non-video tracks (e.g., audio, subtitle and metadata tracks) are extracted and remixed without re-encoding, preserving their original quality while accounting for any timing offsets, ensuring synchronized playback.
[0023] In addition to attribute matching, the system performs precise frame-level alignment to ensure that the fragment frames integrate seamlessly with the original video stream. Frame-level alignment involves extracting stream-level information down to individual frames, identifying exactly which frames require VPP integration. The precise frame numberAttorney Docket No.: 301370400013 W000(20) information is dynamically applied in video editing tools to render only the necessary fragment frames, avoiding the need for full mezzanine rendering. By working at the frame level, this approach guarantees that the rendered fragment frames align perfectly with the original video stream, ensuring seamless playback and maintaining temporal continuity.
[0024] To guarantee playback quality, the present systems and methods incorporate an automated quality control process. This process verifies that the encoding attributes of the VPP fragment frames match those of the original video stream before playback, and also checks for exact matches in non-altered tracks, addressing potential discrepancies such as mismatched colorspace or incorrect bitrate before the viewer notices..
[0025] Real-Time Personalization Capabilities
[0026] The present systems and methods enable real-time personalization capabilities at playback start. This real-time adaptation allows seamless integration of VPPs tailored to live marketing campaigns and viewer-specific preferences. The present systems and methods support personalized targeting by utilizing factors such as geographic location, viewer demographics, browsing history, and behavioral data to dynamically select and integrate relevant brands or products. For example, location-based brands can be prioritized, displaying local advertisements that resonate with viewers in a specific region.
[0027] Universal Compatibility
[0028] The present systems and methods are designed for platform adaptability, operating independently of platform-specific encoding processes while ensuring seamless integration across multiple independent streaming services. The present systems and methods emulate platform- or channel-specific encoding requirements in real time, enabling the same source video to be dynamically adapted for various streaming formats (e.g., HLS. DASH, and CTV) without requiring additional encoding or integration steps. This approach eliminates the need for separate encoding workflows tailored to individual streaming services..Attomey Docket No.: 301370400013 W000(20)
[0029] Streamlined Workflow
[0030] Unlike traditional encoding pipelines, the present systems and methods avoid creating or re-encoding entire mezzanine files, dynamically encoding only the required segments during playback, which reduces storage and computational costs. For example, the present systems and methods reduce encoding time by about 90% compared to traditional full mezzanine reencoding workflows by encoding only the VPP fragment frames, as well as 95% of the storage requirements by avoiding storing alternative mezzanine files. By rendering only the required frames, 95% of the computational resources are saved.
[0031] According to one aspect, the present disclosure is directed to a method for integrating fragment frames into video streams comprising: (a) receiving a media manifest associated with a video stream; (b) identifying, based on the media manifest, one or more segments of the video stream where fragment frames are to be integrated; (c) extracting encoding attributes of the identified one or more segments from the media manifest; (d) rendering the fragment frames corresponding to the identified one or more segments; (e) encoding the rendered fragment frames to match the extracted encoding attributes of the identified one or more segments to yield encoded fragment frames; (f) updating the media manifest to reference the encoded fragment frames in place of the identified segments of the video stream; and (g) transmitting the updated media manifest and the encoded fragment frames to a viewer playback device for playback, wherein the playback device integrates the encoded fragment frames into the video stream.
[0032] In some cases, the encoding attributes comprise at least one of resolution, codec profile, bitrate, colorspace, or group of pictures (GOP) structure.
[0033] In some cases, the fragment frames display one or more virtual product placements.
[0034] In some cases, the virtual product placements comprise at least one of billboards, branded products, logos, text overlays, or original elements of the video stream that areAttorney Docket No.: 301370400013 W000(20) augmented.
[0035] In some cases, the method further comprises dynamically prioritizing the virtual product placements to be displayed in the fragment frames based on marketing campaign criteria and viewer-specific attributes.
[0036] In some cases, the identifying in step (b) comprises extracting a timeline of the media manifest and assessing whether an active ad campaign is applicable based on marketing campaign criteria and viewer-specific attributes.
[0037] In some cases, the method further comprises pre-rendering and caching commonly used fragment frames to reduce latency during playback.
[0038] In some cases, the method further comprises (h) verifying through a quality control process that the extracted encoding attributes of the encoded fragment frames match the encoding attributes of the video stream.
[0039] In some cases, step (h) occurs after step (e).
[0040] In some cases, step (a) occurs at a start time of the playback on the viewer playback device.
[0041] In some cases, the method further comprises extracting non-video tracks from the media manifest, where the non-video tracks comprise audio, subtitle, and metadata tracks, to ensure synchronization of the non-video tracks with the fragment frames.
[0042] In some cases, the method further comprises remixing non-video tracks without reencoding them.
[0043] In some cases, the encoded fragment frames are encoded in a platform-compatible format at playback.
[0044] According to another aspect, the present disclosure is directed to a system for integrating fragment frames into video streams comprising: a controller comprising one or more processors and non- transitory memory storing computing instructions which whenAttorney Docket No.: 301370400013 W000(20) executed by the one or more processors is configured to: (a) receive a media manifest associated with a video stream; (b) identify, based on the media manifest, one or more segments of the video stream where fragment frames are to be integrated; (c) extract encoding attributes of the identified one or more segments from the media manifest; (d) render the fragment frames corresponding to the identified one or more segments; (e) encode the rendered fragment frames to match the extracted encoding attributes of the identified one or more segments to yield encoded fragment frames; (f) update the media manifest to reference the encoded fragment frames in place of the identified segments of the video stream; and (g) transmit the updated media manifest and the encoded fragment frames to a viewer playback device for playback, wherein the playback device integrates the encoded fragment frames into the video stream.
[0045] It should be noted that the technical effects obtainable through the present disclosure are not limited to the above-described effects, and other effects that are not mentioned herein will be clearly understood by those skilled in the art from the following descriptions.BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings illustrate exemplary aspects of the present disclosure and, together with the following detailed description, serve to provide further understanding of the technical spirit of the present disclosure. However, the present disclosure is not to be construed as being limited to the drawings.
[0047] FIG. 1 is a conceptual image showing an identified segment of a video stream where fragment frames are to be integrated in accordance with an aspect of the present disclosure.
[0048] FIG. 2 is a conceptual image showing the identified segment of FIG. 1 modified as a fragment frame with a virtual product placement in accordance with an aspect of the present disclosure.
[0049] FIG. 3 is a flow diagram showing a system for integrating fragment frames into video streams in accordance with an aspect of the present disclosure.Attomey Docket No.: 301370400013 W000(20)
[0050] FIG. 4 is a schematic diagram showing the implementation of the system of FIG. 3 on various computing devices in accordance with an aspect of the present disclosure.
[0001] FIG. 5 is a flow diagram showing a method of integrating fragment frames into video streams in accordance with an aspect of the present disclosure.DETAILED DESCRIPTION OF THE DISCLOSURE
[0052] The present disclosure can be variously changed and have various aspects, and the specific aspects disclosed herein in detail are used to facilitate an understanding of the present disclosure to those skilled in the art.
[0053] Therefore, it should be understood that there is no intention to limit the present disclosure to the particular aspects disclosed, and on the contrary', the present disclosure covers all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure.
[0054] In this application, it should be understood that terms such as “include” or “have” are intended to indicate the presence of a feature, number, step, operation, component, part, or a combination thereof described on the specification, and they do not preclude the possibility of the presence or addition of one or more other features or numbers, steps, operations, components, parts or combinations thereof.
[0055] FIG. l is a conceptual image 100 showing a segment of a video stream that is identified for modification with fragment frames with a virtual product placement, in accordance with an aspect of the present disclosure. The video stream can include media, for example, a film, a TV show, a streaming video, a live broadcast, or any other type of media content capable of being played on a video platform. The segment can be a group of frames of any time frame (for example, from 1 second to 30 minutes). The number of frames in the segment can depend on the time frame of the segment and the frame rate at playback, for example, 24 frames-per-Attorney Docket No.: 301370400013 W000(20) second (FPS), 30 FPS, or 60 FPS. The identified segment can correspond to specific sections within the video stream where VPPs are intended to be visible, such as a billboard in the background or a branded item in a character’s hand.
[0056] The segment can be a scene in the video stream or can be part of a scene in the video stream. A scene can represent a continuous block of visual and / or narrative action occurring within a specific time frame and setting. For example, the image 100 can represent a scene that occurs in a kitchen, and can include elements such as an oven, a stove, a refrigerator, a sink, a microwave, cabinets and an island. A scene can capture events or interactions that occur without interruption in time or logical flow, for example, a conversation between two characters in a room or a chase sequence through a city. The scene can take place in a single location or setting, such as a living room, office, or outdoor park. Changes in location can signify the start of a new scene (but not always). A scene can include multiple camera angles, zooms, or cuts that are unified by the same location or time continuity, and transitions between scenes can be marked by fades, cuts, or other editing techniques. A scene can range in length from a few seconds (e.g., a quick cutaway shot) to several minutes (e.g., an extended dialogue or action sequence).
[0057] The identification or selection of the segment can be influenced by the visual context of a scene, ensuring that the fragment frames integrating the VPP align seamlessly with the aesthetics and narrative of the video stream. For example, if the video stream includes a Halloween scene, and the VPP corresponds to an advertisement for a brand of candy, the segment can be identified in a scene showing elements associated with a Halloween event such as costumes, cobwebs, and jack-o-lantems.
[0058] In some cases, the segment can be pre-identified or pre-selected manually. In some cases, the visual context of a scene can be identified using a visual machine learning algorithm that classifies the elements of a scene. This classification can involve analyzing features suchAttorney Docket No.: 301370400013 W000(20) as colors, obj ects, and textures within the frames to detect thematic elements, for example, such as orange hues and pumpkin shapes indicative of Halloween. Furthermore, metadata associated with the video stream, such as scene descriptions or timestamps, can be cross- referenced to refine the identification process. By combining visual analysis and metadata, the system can ensure that the selected segment is relevant to both the narrative and the advertising goals.
[0059] Furthermore, the segment can be dynamically identified during playback based on marketing campaign criteria or viewer-specific attributes, enabling personalized content delivery without requiring manual pre-selection of the segment. Marketing campaign criteria can include factors such as the type of product being advertised, the target audience demographics, the timing or duration of the advertisement, and the specific moments within the video stream where the VPP achieves the highest visibility or engagement (in other words, the most views or impressions). For instance, a campaign for a sports drink can prioritize segments featuring athletic action, while a luxury brand can focus on scenes with upscale settings or affluent characters. Viewer-specific attributes can include data such as geographic location, age, gender, viewing history, or real-time engagement metrics such as whether a viewer has paused or rewound a particular segment. This feature enables the identification of segments that are more likely to resonate with the individual viewer, such as showing a locally relevant brand or a product that aligns with the viewer’s previous preferences.
[0060] In some cases, instead of video stream, a static image (for example, to be shown on a website) can modified with a virtual product placement in accordance with the present disclosure. The static image can include photographs, digital artwork, or graphic designs that serve as promotional materials, social media posts, or website banners. The VPP can be seamlessly integrated into the image by analyzing the image's existing elements, such as lighting, perspective, and color palette, to ensure the VPP appears natural and cohesive.Attorney Docket No.: 301370400013 W000(20)
[0061] FIG. 2 is a conceptual image 200 showing the identified segment of FIG. 1 modified as a fragment frame with a virtual product placement 202, in accordance with an aspect of the present disclosure. The VPP 202 is depicted as a container of pea soup positioned on the kitchen island. This placement aligns naturally with the kitchen setting, where the presence of food is both expected and contextually appropriate. The VPP 202 can be rendered with lighting, shadows, and reflections that match the surrounding environment, ensuring seamless integration with the scene so that the VPP 202 appears as if originally part of the video stream. Additionally, the positioning of the VPP 202 can be selected to maximize visibility without disrupting the narrative flow of the scene. Although the conceptual image 200 shows a static image, the identified segment can be a scene, and the VPP 202 can be viewed from other angles (so that the sides or rear of the VPP 202 are visible) as the scene progresses. For example, as the camera pans across the kitchen, the rendering of the VPP 202 can adjust to reflect changing perspectives, ensuring that proportions, dimensions, and visual effects remain accurate from all angles.
[0062] In some cases, one or more virtual product placements displayed in a fragment frame can comprise at least one of billboards, branded products, logos, text overlays, or original elements of the video stream that are augmented. Billboards can include large, attentiongrabbing advertisements on planar surfaces, and can be placed in outdoor or urban settings within the video stream, such as a cityscape or sports stadium. Branded products can include physical items, such as beverages, electronics, or clothing, integrated into the scene in a manner consistent with the setting and action, such as a branded laptop on an office desk or a branded coffee cup held by a character. Logos can include graphics or symbols placed within the video stream, such as on clothing, equipment, or building signage. Text overlays can include messages such as slogans, hashtags, or promotional information and can appear as part of a digital screen or embedded in a background element. Augmented original elements can involveAttorney Docket No.: 301370400013 W000(20) enhancing or modifying obj ects already present in the video stream to include branded content, such as replacing a generic can label with a specific product's branding or digitally altering a poster on a wall to feature an advertisement. These various forms of VPPs can be carefully placed to maintain the realism of the video stream while maximizing the impact and relevance of the advertisement to the viewer.
[0063] FIG. 3 is a flow diagram showing a system 300 for integrating fragment frames with a VPP into video streams in accordance with an aspect of the present disclosure.
[0064] A viewer 336 can request 338 a video stream from a playback module 340 (e.g., running on server computing device[s]) using a viewer playback device (e g, a computing device, a smartphone, a laptop, etc.) running a playback application (e.g, internet-connected video player software). The playback module 340 can include a playback resource service 342 and a playback schedule sen-ice 346. The playback resource service 342 can be responsible for managing the resources required to play a video stream. The playback resource service 342 can ensure that the server computing device(s) provides the necessary bandwidth, computing resources, and storage to deliver the video stream, fetch or process video / audio assets, ensure that users are authorized to access the requested video stream (for example, by applying DRM [Digital Rights Management] controls), and provide smooth playback by pre-fetching or caching content.
[0065] The playback resource service can generate 344 a media manifest and transmit the manifest to the playback schedule service 346. The manifest file can be a structured representation of how the video stream should be streamed to the user. When the viewer 336 requests 338 video stream, the server computing device(s) dynamically generate 344 the manifest tailored to the viewer playback device, network, and preferences. The manifest lists available video segments, their durations, resolutions, and bitrates, and helps the viewer playback device decide which segment to download based on network conditions. In HLSAttorney Docket No.: 301370400013 W000(20) (HTTP Live Streaming), the manifest file is an M3U8 file that lists .ts segments of the video stream. In DASH (Dynamic Adaptive Streaming over HTTP), the manifest file is an MPD (Media Presentation Description) file.
[0066] The playback schedule service 346 can handle the timing and order of media playback, and can ensure that a scheduled playlist of video streams is presented in the correct order and / or that live streams start on time. The playback schedule service 346 can coordinate with other services to provide seamless transitions between pieces of video streams, and can insert advertisements at the right times (pre-roll, mid-roll, or post-roll) based on a playback timeline. The playback schedule service 346 can implement user playback controls, such as "resume from last position," "chapter-based navigation," or other timeline-based functionalities, can manage interruptions and ensure playback resumes correctly after a pause or disruption, and can suggest and queue up next videos or playlists based on user behavior and history.
[0067] The manifest can be transmitted to a manifest modification module 306. The manifest can be modified 348 by a manifest modification service 350 by identifying segments of the video stream where fragment frames are to be integrated with a VPP. The manifest modification service 350 can assess whether any active marketing campaigns (z.e., ad campaigns) are applicable for the specific video stream that is requested 338 by the viewer. The segments of the video stream can be identified based on marketing criteria (e.g.. provided by a brand and / or marketing agency) and user-specific attributes (e.g., collected from the viewer).
[0068] The manifest modification service 350 can communicate with a supply-side platform (SSP) 310 and / or a demand-side platform (DSP) 308 to request 352 live marketing campaigns and dynamically decide which VPPs should be shown to viewers in real-time by integrating the VPPs into the fragment frames. The manifest modification service 350 can fetch the most relevant targeted advertisements to display, and can retrieve information about marketingAttorney Docket No.: 301370400013 W000(20) campaigns that match user-specific attributes (e.g., geolocation, time of day). For example, if a viewer in New Y ork watches a live sports event, the live campaign request 352 ensures that the viewer sees a VPP for a car dealership that is local to New York.
[0069] An advertiser 302 (e.g., brand user) can enable 304 a marketing campaign through a DSP 308 and / or SSP 310. The DSP 308 can be a digital advertising platform used to purchase advertisement inventory across multiple publishers through auctions, and can help target specific audiences based on demographics, behavior, geolocation, and more. The DSP 308 can implement Real-Time Bidding (RTB) which bids on advertisement inventory from an SSP 310 and can match advertisements to users using data like cookies or device IDs. The DSP 308 can provide tools for managing budgets and performance metrics. For example, a retail company can use a DSP 308 to bid for advertisement space on streaming services, ensuring their Black Friday sale ads are shown to relevant users watching videos.
[0070] The SSP 310 can be a digital advertising platform used by publishers (e.g, websites, apps, streaming platforms) to sell their advertisement inventory in an automated, programmatic way. The SSP 310 can help publishers maximize revenue by making their advertisement inventory available to a wide range of buyers (advertisers, DSPs, ad exchanges, etc.), and manages pricing, availability, and targeting options for advertisement placements. For example, a streaming service (e.g., Netflix) can use the SSP 310 to sell advertisement slots during a live TV broadcast, allowing advertisers to bid for the spots in real-time. An example of an SSP 310 is SPHEERA (trademark of Ryff, Inc., Los Angeles. CA).
[0071] If a marketing campaign is detected by the manifest modification service 350, the timeline of the manifest can be extracted to identify segments requiring modifications with fragment frames integrating the VPPs. These segments can correspond to portions of the video where the VPPs can be most effectively integrated, such as prominent scenes or visually impactful moments. The modified manifest can then be transmitted 360 to the viewerAttomey Docket No.: 301370400013 W000(20) playback device.
[0072] In addition, encoding atributes can be extracted from the original segments of the video stream, including information such as resolution, group of pictures (GOP) structure, bitrate, codec profile, and visual characteristics such as colorspace and pixel format. Extracting resolution can ensure that the fragment frames match the pixel dimensions of the original video stream to avoid visual inconsistencies. Extracting GOP structure can match compression paterns to presen e temporal coherence between the original video stream and the fragment frames. Extracting bitrate can ensure a consistent data rate to smooth streaming and playback without quality degradation or buffering. Extracting codec profile can enable the fragment frames to adhere to the specific video compression standards (e.g., H.264 or HEVC) of the original video stream to ensure compatibility with playback devices. Extracting colorspace and pixel format can ensure that the fragment frames blend seamlessly with the visual style and quality of the original video.
[0073] Additionally, non-video tracks such as audio, subtitle, and metadata tracks can be extracted from the manifest to ensure synchronization with the modified fragment frames. These non-video tracks can include audio tracks, which contains dialogue, music, or sound effects; subtitle tracks, which can include closed captions or translations for accessibility; and metadata tracks, which can include information such as chapter markers or descriptive content. The synchronization process can involve aligning the timing and sequence of these non-video tracks with the fragment frames to ensure a cohesive playback experience. The non-video tracks can be remixed without re-encoding them, preserving their original quality and minimizing computational overhead. Remixing ensures that the audio, subtitles, and metadata remain perfectly synchronized with the modified fragment frames, avoiding any playback delays or mismatches that could disrupt the viewer experience. Remixing eliminates the need for re-encoding, which could otherwise degrade the quality of the non-Attorney Docket No.: 301370400013 W000(20) video tracks or introduce additional latency. The remixed non- video tracks can be transmitted to the viewer playback device with the modified fragment frames and the modified manifest.
[0074] The segments identified by the manifest modification sendee 350 can be transmitted to a fragments catalog module 312. The fragment frames can be rendered 314 with the VPPs which represent the targeted brands 316. The rendering process can incorporate advanced techniques, such as matching lighting, shadows, reflections, and perspective, to ensure that the VPPs appear natural and cohesive within the scene. The rendering process can leverage distributed resources, such as render farms, which can comprise multiple high-performance computing devices working in parallel to process the video frames. In some cases, the rendering process is deployed using a cloud service provider (e.g., AWS, Google Cloud, Microsoft Azure, etc ).
[0075] To further enhance efficiency, intermediate rendering results (e.g., without a VPP inserted) can be cached, such that different VPPs may be easily inserted when necessary. This feature is particularly useful when multiple variations of the same segment are modified with different VPPs. Additionally, commonly used fragment frames with a specific VPP can be pre-rendered and cached to reduce latency during playback. Caching can enable the reuse of pre-processed data and reduces the need to recompute identical fragment frames.
[0076] The original mezzanine file associated with the video stream can be transmitted to a catalog module 326 and can be encoded 320 by an encoding module 328. The original mezzanine file can be a high-quality, intermediate video file used as the master source for generating various output formats, and can typically retain high-resolution video and audio to preserve the best possible quality’. Encoding 320 the mezzanine file can involve compressing the high-quality video and audio content into formats optimized for specific platforms or devices, such as streaming services, broadcast television, or mobile playback. This process can use codecs (e.g., H.264, HEVC) to reduce the file size while maintaining acceptable visualAttorney Docket No.: 301370400013 W000(20) and audio quality. The encoding 320 generates multiple versions of the video content at different resolutions, bitrates, and frame rates, enabling adaptive streaming and compatibility across various networks and devices. In some cases, the mezzanine file is DRM encoded 330 using a DRM service 332. The DRM service 332 can embed encryption to protect the video content from unauthorized access or piracy, and can ensure that only authorized users can decrypt and play the content.
[0077] The media distribution service 356 can be a platform responsible for serving 354 (z. e. , delivering) the video stream to the viewer playback device. The encoded video content can be ingested 334 by a media distribution service 356. Ingesting 334 can involve uploading the encoded video content into the media distribution sendee 356, which then transcodes the video content into various formats and stores the video content for distribution.
[0078] Once the video content is ingested 334, the video stream is then made available for delivery to the viewer playback device. The media distribution service 356 serves 354 the video stream by streaming or downloading the video content to the viewer playback device based on the request 338, and can involve adaptive streaming where the best resolution and bitrate are selected depending on the viewer playback device and network conditions.
[0079] In a parallel process, the rendered fragment frames integrated with the VPPs representing the targeted brands 316 can be encoded 318 by an encoding module 322. The encoding module 322 can encode the fragment frames using the encoding attributes extracted from the manifest to ensure seamless integration with the original video stream. The encoded fragment frames can then be ingested 324 and served 352 by a fragments media distribution service 354. The encoded video content and the encoded fragment frames can be transmitted 358 to the viewer playback device for playback. The manifest can be updated or modified to reference the encoded fragment frames in place of the identified segments of the video stream, enabling the viewer playback device to seamlessly integrate the encoded fragment frames intoAttorney Docket No.: 301370400013 W000(20) the video stream during playback.
[0080] In some cases, a quality control process can verity’ that the extracted encoding attributes of the encoded fragment frames match the encoding attributes of the video stream. This verification can include comparing parameters such as resolution, bitrate, codec profile, colorspace, pixel format, and GOP structure to ensure seamless integration during playback. Any discrepancies identified during this process, such as mismatched colorspace or frame rate, can trigger an automated correction mechanism to re-encode the fragment frames with the correct attributes. Additionally, the quality control process can involve playback simulation to confirm that the encoded fragment frames transition smoothly with the original video stream without introducing visual artifacts, stuttering, or desynchronization issues.
[0081] FIG. 4 is a schematic diagram illustrating one or more viewer playback computing devices 402 (i.e., viewer playback devices) and one or more media server computing devices (z.e., media servers) implementing the system 300 described with respect to FIG. 3, according to an aspect of the present disclosure.
[0082] Each of the viewer playback computing device(s) 402 and server computing device(s) 410 can respectively include one or more processors 404 and 412 (i.e., processing modules) configured to execute program instruct! ons maintained on a memory 406 and 414 (i. e. , memory modules). In this regard, the one or more processors 404 and / or 412 can execute any of the various methods, processes, steps, programs, functions, modules or algorithms described throughout the present disclosure. The memory 406 can store a viewer playback application 408 which plays the video stream modified with the fragment frames. The memory 414 can store the playback module 340, manifest modification module 306, catalog module 326, and fragments catalog module 312. The viewer playback computing device(s) 402 can communicate with the media server computing device(s) 410 via a network connection.
[0083] The computing device(s) 402 and / or 410 can comprise a desktop computer, laptopAttorney Docket No.: 301370400013 W000(20) computer, server computer, mainframe computer system, workstation, image computer, parallel processor, smartphone, tablet, or any other computer system (e.g., networked computer). The processors 404 and / or 412 can include any processing element know n in the art. In this sense, the processors 404 and / or 412 can include any microprocessor-type device configured to execute algorithms and / or instructions, for example, application specific integrated circuit (ASIC), field programmable gate array (FPGA), parallel processor, graphics processing unit (GPU), central processing unit (CPU), a logical circuit, an electronic processor and / or other chipsets. It is further recognized that the term “processor” can be broadly defined to encompass any device having one or more processing elements, which execute program instructions from anon-transitory memory.
[0084] The memory 406 and / or 414 can include any storage medium known in the art suitable for storing program instructions executable by the associated processors 404 and / or 412. For example, the memory 406 and / or 414 can include a non-transitory memory medium. By way of another example, the memory 406 and / or 414 can include, but is not limited to, a read-only memory, a random access memory, a magnetic or optical memory device (e.g., disk), a magnetic tape, a solid state drive, etc. It is further noted that memory 406 and / or 414 can be housed in a common housing with the respective processors 404 and / or 412. In some cases, the memory 406 and / or 414 can be located remotely with respect to the physical location of the respective processors 404 and / or 412 and the respective computing devices 402 and / or 410.
[0085] FIG. 5 is a flow diagram showing a method 500 of integrating fragment frames into video streams in accordance with an aspect of the present disclosure. The method 500 can implement the system 300 described with respect to FIG. 3.
[0086] At step 502, a media manifest can be generated after a viewer requests a video stream for playback, and the manifest associated with the video stream can transmitted (e.g.. from the playback module 340 to the manifest modification module 306).Attorney Docket No.: 301370400013 W000(20)
[0087] At step 504, one or more segments of the video stream where fragment frames are to be integrated can be identified based on the media manifest (for example, by the manifest modification module 306).
[0088] At step 506, encoding attributes of the identified one or more segments can be extracted from the manifest (for example, by the manifest modification service 350).
[0089] At step 508, the fragment frames can be rendered corresponding to the identified one or more segments (for example, by the fragments catalog module 312).
[0090] At step 510, the rendered fragment frames can be encoded to match the extracted encoding attributes of the identified one or more segments to yield encoded fragment frames (for example, by the fragments encoding module 322).
[0091] At step 512, the manifest can be updated to reference the encoded fragment frames in place of the identified segments of the video stream (for example, by the manifest modification service 350).
[0092] At step 514, the updated media manifest and the encoded fragment frames are transmitted to a viewer playback device for playback, where the playback device integrates the encoded fragment frames into the video stream.
[0093] It is noted that the steps 502-514 can be performed in a sequential (serial) manner.However, the present disclosure is not limited thereto, and, in some cases, the steps 502-514 can be performed in a parallel manner.
[0094] In the above, the present disclosure has been described in more detail through the drawings and aspects. However, the configurations described in the drawings or the aspects in the specification are merely aspects of the present disclosure and do not represent all the technical ideas of the present disclosure. Thus, it is to be understood that there can be various equivalents and variations in place of them at the time of filing the present application which are encompassed by the claims.
Claims
Attorney Docket No.: 301370400013 W000(20)CLAIMSWhat is claimed is:
1. A method for integrating fragment frames into a video stream, comprising:a) receiving a media manifest associated with the video stream;b) identifying, based on the media manifest, one or more segments of the video stream where fragment frames are to be integrated;c) extracting encoding attributes of the identified one or more segments from the media manifest;d) rendering the fragment frames corresponding to the identified one or more segments;e) encoding the rendered fragment frames to match the extracted encoding attributes of the identified one or more segments to yield encoded fragment frames;f) updating the media manifest to reference the encoded fragment frames in place of the identified segments of the video stream; andg) transmitting the updated media manifest and the encoded fragment frames to a viewer playback device for playback, wherein the viewer playback device integrates the encoded fragment frames into the video stream.
2. The method of claim 1, wherein the encoding attributes comprise at least one of resolution, codec profile, bitrate, colorspace, or group of pictures (GOP) structure.
3. The method of claim 1 or 2. wherein the fragment frames display one or more virtual product placements.
4. The method of any one of claims 1 to 3, wherein the one or more virtual product placements comprise at least one of billboards, branded products, logos, text overlays, or original elements of the video stream that are augmented.
5. The method of any one of claims 1 to 4, further comprising dynamically prioritizing the one or more virtual product placements to be displayed in the fragment frames based on marketing campaign criteria and viewer-specific attributes.Attorney Docket No.: 301370400013 W000(20)6. The method of any one of claims 1 to 5, wherein the identifying in step (b) comprises extracting a timeline of the media manifest and assessing whether an active ad campaign is applicable based on marketing campaign criteria and viewer-specific attributes.
7. The method of any one of claims 1 to 6, further comprising pre-rendering and caching commonly used fragment frames to reduce latency during playback.
8. The method of any one of claims 1 to 7, further comprising (h) verifying through a quality control process that the extracted encoding attributes of the encoded fragment frames match the encoding attributes of the video stream.
9. The method of claim 8, wherein step (h) occurs after step (e).
10. The method of any one of claims 1 to 9, wherein step (a) occurs at a start time of the playback on the viewer playback device.
11. The method of any one of claims 1 to 10. further comprising extracting non-video tracks from the media manifest, wherein the non-video tracks comprise audio, subtitle, and metadata tracks, to ensure synchronization of the non-video tracks with the fragment frames.
12. The method of any one of claims 1 to 11, further comprising remixing non-video tracks without re-encoding them.
13. The method of any one of claims 1 to 12, wh erein the encoded fragment frames are encoded in a platform-compatible format at playback.
14. A system for integrating fragment frames into a video stream, comprising:a controller comprising one or more processors and non-transitory memory storing computing instructions which when executed by the one or more processors is configured to:a) receive a media manifest associated with the video stream;b) identify, based on the media manifest, one or more segments of the video stream where fragment frames are to be integrated;c) extract encoding attributes of the identified one or more segments from the media manifest;Attorney Docket No.: 301370400013 W000(20)d) render the fragment frames corresponding to the identified one or more segments;e) encode the rendered fragment frames to match the extracted encoding attributes of the identified one or more segments to yield encoded fragment frames;f) update the media manifest to reference the encoded fragment frames in place of the identified segments of the video stream; andg) transmit the updated media manifest and the encoded fragment frames to a viewer playback device for playback, wherein the viewer playback device integrates the encoded fragment frames into the video stream.
15. The system of claim 14, wherein the encoding attributes comprise at least one of resolution, codec profile, bitrate, colorspace, or group of pictures (GOP) structure.
16. The system of claim 14 or 15, wherein the fragment frames display one or more virtual product placements.
17. The system of any one of claims 14 to 16, wherein the one or more virtual product placements comprise at least one of billboards, branded products, logos, text overlays, or original elements of the video stream that are augmented.
18. The system of any one of claims 14 to 17, wherein the controller is further configured to dynamically prioritize the virtual product placements to be displayed in the fragment frames based on marketing campaign criteria and viewer-specific attributes.
19. The system of any one of claims 14 to 18, wherein the identifying in step (b) comprises extracting a timeline of the media manifest and assessing whether an active ad campaign is applicable based on marketing campaign criteria and viewer-specific attributes.
20. The system of any one of claims 14 to 19, wherein the controller is further configured to pre-render and cache commonly used fragment frames to reduce latency during playback.Attorney Docket No.: 301370400013 W000(20)21. The system of any one of claims 14 to 20, wherein the controller is further configured to (h) verify through a qualify control process that the extracted encoding attributes of the encoded fragment frames match the encoding attributes of the video stream.
22. The system of claim 21, wherein step (h) occurs after step (e).
23. The system of any one of claims 14 to 22, wherein step (a) occurs at a start time of the playback on the viewer playback device.
24. The system of any one of claims 14 to 23, wherein the controller is further configured to extract non-video tracks from the media manifest, wherein the non-video tracks comprise audio, subtitle, and metadata tracks, to ensure synchronization of non-video tracks with the fragment frames.
25. The system of any one of claims 14 to 24, wherein the controller is further configured to remix non-video tracks without re-encoding them.
26. The system of any one of claims 14 to 25, wherein the encoded fragment frames are encoded in a platform-compatible format at playback.