A system and method for targeted content

EP4690810A1Pending Publication Date: 2026-02-11BROADCAST VIRTUAL PTY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024777336
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-31
Filing Date
2024-03-31
Publication Date
2026-02-11

AI Technical Summary

Technical Problem

Current methods fail to seamlessly integrate targeted content into live broadcasts, making it difficult to provide customized information or advertising that appears naturally within the camera image frame.

Method used

A system and method that utilize a client device, upstream, and downstream devices to determine client identification information, generate targeted content, and seamlessly integrate it into the broadcast by using unique frame signatures, timestamps, and tracking data to augment the primary image frame with virtual content.

Benefits of technology

Enables viewers to receive customized advertising and relevant information by seamlessly integrating targeted content into live broadcasts, enhancing viewer engagement and relevance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2024050302_03102024_PF_FP_ABST
    Figure AU2024050302_03102024_PF_FP_ABST
Patent Text Reader

Abstract

A system for augmenting a broadcast with virtual content, the system including a client device for viewing the broadcast, at least one camera for receiving a primary image frame of the broadcast, at least one upstream device in communication and associated with the at least one camera, and, at least one downstream device in communication with the at least one upstream device, wherein the system is configured to determine client identification information associated with the client device, determine one or more virtual content based on the client identification information and augment the primary image frame with the virtual content to be displayed on the client device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A System and Method for Targeted Content

[0002] Technical Field

[0003] The present invention relates to a method and system for providing / generating content for a broadcast. In one particular example, the method and system described herein provides targeted content for a live broadcast program.

[0004] Background of the Invention

[0005] The following references to and descriptions of prior proposals or products are not intended to be and are not to be construed as, statements or admissions of common general knowledge in the art. In particular, the following prior art discussion does not relate to what is commonly or well known by the person skilled in the art, but assists in the understanding of the inventive step of the present invention of which the identification of pertinent prior art proposals is but one part.

[0006] Presently, live broadcasts, either via television or digitally through the internet, are still common for people to watch, such as the news or live sporting events. These broadcasts often have the same images displayed to all viewers.

[0007] Targeting content to a viewer in a live broadcast has many advantages in that a user can be shown content particularly relevant or customised to them. This can include advertising but may also include other information for a specific viewer such as events near them or the weather associated to where the viewer is.

[0008] However, there are technical challenges in providing targeted content during a live broadcast, especially content that is seamlessly integrated in an image frame for the viewer. That is, it is difficult to augment an image received from a camera with additional content, where the content does not look like it is surrounding or overlaid on the camera image but a part of the image, especially for a live camera feed. Thus, there is no known method or system by which viewers of a live broadcast can be presented targeted content.

[0009] The present invention seeks to provide a system and method for providing content which may ameliorate the foregoing shortcomings and disadvantages, or which will at least provide a useful alternative.

[0010] Summary of the Invention

[0011] According to one aspect of the invention, there is provided herein systems and methods for providing / inserting / generating content to a broadcast. According to one example, the content is generated on a live broadcast. In a further example, the content is generated on a broadcast replay. In yet a further example, the content is integrated seamlessly with the broadcast. In yet a further example, the content can be virtual advertising specifically customised (or targeted) for the viewer of the broadcast.

[0012] According to one aspect, there is provided a system for providing / inserting / generating content for a live broadcast on a client device, the system including at least one camera for receiving an image / a frame of the live broadcast device; at least one upstream device operatively connected to the at least one camera; and, at least one downstream device operatively connected to the at least one upstream device; wherein the system is configured to determine client identification information associated with the client device, determine one or more adverts associated with the client identification information; and generate targeted advertising on the received image / frame to be displayed on the client device.

[0013] In one example, the upstream device is configured to: receive the frame; generate downstream data stream including any one or a combination of: a unique frame signature associated with the frame a unique timestamp for received frame; and, frame tracking data for received frame; and, send the downstream data to the downstream device.

[0014] In one form, the downstream device is configured to: receive downstream data from the upstream device; receive transmission signal with graphics or without graphics; and, augment transmission signal with graphical content.

[0015] In yet another example, the downstream device is further configured to: regenerate a frame signature for the received frame; qquery timestamp and tracking data associated with frame; and, compare and determine whether virtual insertion is required based on the timestamp and tracking data.

[0016] According to a further aspect, there is provided herein a system for augmenting a broadcast with virtual content, the system including: a client device for viewing the broadcast; at least one camera for receiving a primary image frame of the broadcast; at least one upstream device in communication and associated with the at least one camera; and, at least one downstream device in communication with the at least one upstream device; wherein the system is configured to determine client identification information associated with the client device, determine one or more virtual content based on the client identification information; and augment the primary image frame with the virtual content to be displayed on the client device. In one form, the upstream device is configured to: receive the primary image frame; generate downstream data stream associated with the primary image frame including any one or a combination of: a unique frame signature associated with the primary image frame; a unique upstream timestamp for received primary image frame; frame tracking data for received primary image frame; and, and insertion instructions; store the downstream data stream associated with the primary image frame in a data store; and, send the downstream data stream to the downstream device.

[0017] In yet a further example, the downstream device is configured to: receive downstream data stream from the upstream device; receive a secondary image frame; compare with the down stream data and, augment the primary image with graphical content.

[0018] According to a further form, the downstream device is further configured to: receive the secondary image frame at a time interval after the upstream device receives the clean signal; regenerate a frame signature for the secondary image frame; query timestamp and tracking data associated with secondary image frame; and, query the data store and compare with the primary image frame and determine whether virtual insertion is required based on the timestamp and tracking data and insertion instructions.

[0019] In one example, the secondary image frame is generated by any one of or a combination of: a transmission signal with graphics; and, a transmission signal without graphics.

[0020] According to another example, the primary image frame is generated by a clean signal transmitting the image frame.

[0021] In a further example, the upstream system applies a hashing algorithm to the video frame.

[0022] According to another example, the tracking data includes any one or a combination of intrinsic camera data and extrinsic camera data.

[0023] In a further form, the upstream system further confirms whether operator changes are required to the primary image frame and updates the primary image frame accordingly.

[0024] According to another example, the downstream system further generates a globally unique timestamp by combining the received upstream timestamp, a connection counter and a unique downstream timestamp. In a further form, the downstream system stores the globally unique timestamp in the datastore against the tracking data.

[0025] According to another example, the secondary image frame includes any one or a combination of a clean image frame and a final production image frame.

[0026] In a further form, when the secondary image frame is compared to the primary image frame, if there is a match, it is determined whether virtual insertion or targeted content is required, and the primary image frame is augmented accordingly.

[0027] According to another example, if there is no exact match, either the primary image is not augmented or neighbouring frames are checked to determine whether there is a match and if the primary image is to be augmented.

[0028] In one example, the time interval is 1 second, although any time delay is possible.

[0029] According to yet another example, the system has a plurality of cameras, each of the plurality of cameras having an associated upstream system.

[0030] In another form, the downstream system includes a processing system and a data store associated with each of the plurality of cameras.

[0031] According to a further aspect, there is provided herein a method for augmenting a broadcast with virtual content, the method including: receiving a primary image frame of the broadcast from at least one camera; determining client identification information associated with a client device for viewing the broadcast; and, determining one or more virtual content based on the client identification information; and, augmenting the primary image frame with the virtual content to be displayed on the client device.

[0032] According to a further aspect, there is provided herein a method for augmenting a broadcast with virtual content, where in an upstream device, the method includes: receiving a primary image frame; generating downstream data stream associated with the primary image frame including any one or a combination of: a unique frame signature associated with the primary image frame; a unique upstream timestamp for received primary image frame; frame tracking data for received primary image frame; and, and insertion instructions; storing the downstream data stream associated with the primary image frame in a data store; and, sending the downstream data stream to the downstream device. In yet a further aspect, there is provided herein a method for augmenting a broadcast with virtual content, where in a downstream device, the method includes: receiving downstream data stream from the upstream device; receiving a secondary image frame; comparing with the down stream data and, augmenting the primary image with graphical content.

[0033] According to a further example, the method includes, in the downstream device: receiving the secondary image frame at a time interval after the upstream device receives the clean signal; regenerating a frame signature for the secondary image frame; querying timestamp and tracking data associated with secondary image frame; and, querying the data store and compare with the primary image frame and determine whether virtual insertion is required based on the timestamp and tracking data and insertion instructions.

[0034] It will be appreciated that the system and method of the present invention can include any of the aspects, forms, and examples described herein in any combination.

[0035] Brief Description of the Drawings

[0036] The invention may be better understood from the following non-limiting description of a preferred embodiment, in which:

[0037] Figure 1 shows an example of a flow diagram for an aspect of the present invention;

[0038] Figure 2 shows another example flow diagram of an aspect of the present invention;

[0039] Figure 3 shows a further example flow diagram of another aspect of the present invention;

[0040] Figures 4A and 4B show further example flow diagrams of another aspect of the present invention;

[0041] Figure 5 shows another example flow diagram of another aspect of the present invention;

[0042] Figure 6 shows an example system diagram of an aspect of the present invention; and,

[0043] Figure 7 shows an example system diagram of another aspect of the present invention. Detailed Description

[0044] The present invention can allow for a viewer watching a particular program on any device to receive targeted content directly to their device. As an example, there is provided herein a system for providing / inserting / generating targeted content for a live broadcast on a client device, the system including at least one camera for receiving an image / a frame of the live broadcast device, at least one upstream device operatively connected to the at least one camera, and, at least one downstream device operatively connected to the at least one upstream device. Thus, for example, the system is configured to determine client identification information associated with the client device, determine one or more content information associated with the client identification information, and generate targeted content on the received image / frame to be displayed on the client device.

[0045] Notably, in the examples described herein, the user / viewer is shown customised advertising. However, it will be appreciated that the customised content provided to the user can be any form of content such as weather, news, events, and advertising which the viewer may find relevant and specific to them.

[0046] Typically, the system includes a client device, an upstream device associated with a camera (that is configured to provide a live broadcast), and a downstream device. However, it will be appreciated that the system can also include multiple cameras and multiple upstream devices associated with the camera to provide different views of an event (for example, a sporting event).

[0047] Further examples of the interactions between these devices are provided below.

[0048] Example: Client Device

[0049] Figure 1 shows an example of a client device such as a tablet, phone, laptop, computer, telecommunication device or any device that can receive live TV signals such as a sports game or the like.

[0050] In this example, at step 100, the client device can connect to a content server such as an advert server, at step 110, the system identifies the client device, and can determine client information. The system can further identify the event that the client is viewing. At step 120, the system will then also identify the event frame and at step 130, the system typically then looks up or determines the advert that is to be displayed on the client device based on the event frame, and client information. At step 140, the system can the generate and display the determined advert on the client’s device. Typically, a client device would have two data streams to choose between a standard video stream containing basic untargeted branding or a video stream with additional metadata and key signal to allow per-device rendering.

[0051] In this example, the advert server provides the client device with adverts such as sponsor information, artwork, etc., that is to be inserted. The client device can then receive the data streams from the system and then further receive the adverts that will be displayed by the client device to the user. For example, if the user is watching rugby league, the adverts associated with or particular to the user can then appear on the grass on the field.

[0052] Thus, for example, a client device can be a smart TV, set top box or mobile device used for viewing a sporting event, where the event can be viewed live or as a replay.

[0053] The system’s augmented data streams can thus be hosted by the streaming platform in a format compatible with the platform’s infrastructure, using industry standard formats. Content can thus be provided through an encoding process to produce the two feeds outlined above. Notably, the standard stream typically requires no changes to the existing workflow and can be viewed on the client’s device unaltered.

[0054] According to one example, where an advert is required to be inserted on an event, once the event is selected for viewing, the device establishes a connection to the advert server. Upon connection the following information is determined:

[0055] - An authorisation check is performed to establish if the streaming platform can provide targeted branding for the selected match / streaming platform / user.

[0056] - A user identity or profile is sent to the server to be associated with the connection. This can be used to provide demographically or individually targeted content. The level of user targeting (country, region, age range, household, individual person) typically depends on the capabilities of the streaming platform and privacy considerations.

[0057] The selected match is then typically associated with connection. Once a connection has been established to the advert server the following information can be retrieved on an on-going periodic basis:

[0058] - Encryption keys can be retrieved to decrypt the data feed; and,

[0059] - the match and user profile can be used to serve artwork (images and movies) to the client device for branding insertion. The artwork is tailored based on the user profile, match and any other pertinent conditions. Periodically the client device can also report viewing statistics back to the advert server. This only needs to consist of a summary of parts of the stream viewed by the client device. Notably, the advert server can determine branding exposure directly as it knows which advertising slots are available for any frame of video and what artwork was supplied to the client device.

[0060] For each frame that the client device produces it can either:

[0061] - decode the clean video feed and key signal;

[0062] - decrypt the tracking data;

[0063] - render virtual branding content and composite it accordingly; or

[0064] - render the standard fallback stream.

[0065] The client device can fall back to the standard stream in case of performance or connectivity issues, thus ensuring a minimum level of branding at all times.

[0066] Example: Upstream

[0067] Figure 2 shows an example of the process of the upstream device in the system. In this example, at step 200 the system receives a frame, and at step 205, the system can generate a unique frame signature associated with the received frame. At step 210, a unique time-stamp is generated for the received frame, which is sent to the downstream device. Further, in the upstream device, at step 220, the operator can either update content graphics in the frame or alternatively generate a graphical placeholder for a targeted content. At step 225, time stamp data is further generated, and a snapshot can be taken at step 230 and the respective data is sent to the downstream device at step 225.

[0068] Thus in one example, when a video frame is captured, the following information can be generated:

[0069] - A unique frame signature. As an example, a hash code which is computed using a variant of perceptual hashing that produces a bitstring.

[0070] - A unique timestamp integer, which in one example can typically include a monotonically increasing counter that counts the number of video frame ticks that have occurred since the system was last powered on. A tick being the smallest unit of time of the video signal i.e. a field in interlaced formats and a frame in progressive formats - depends on the frame rate of the video

[0071] For each video frame tracking data is computed either via various techniques (typically mechanical tracking or image-based tracking). According to one example, tracking data can include the camera pose (where the camera is) and it’s optical parameters (focal length, lens distortion, etc). Notably, operator driven updates to the state of the graphics (location, artwork) can be time stamped via the same unique timestamp mentioned above. The graphics snapshot contains data pertaining to all possible combinations of graphics that can be applied on the downstream devices. In one particular example, the graphics snapshot state comprises of:

[0072] - Location, size and appearance of all graphical elements to be inserted across all feeds (a feed is a subset of graphical elements relevant to a particular market e.g. Australia or Europe).

[0073] - Control parameters including lookup tables for keying operations (eg. Cutting people out of the grass) - these are required to perform rendering subsequently.

[0074] - Per technology metadata e.g. lighting estimates for virtual billboards or ball-tracking techniques, etc.

[0075] Typically, an operator watching the match live controls the artwork, and can choose where the artwork is placed and control the overall appearance. When the operator makes the decision, an entire snapshot is taken and timestamped. The contains a complete description of all the possible graphical elements which covers all the variations of branding.

[0076] In this example, any time the operator makes an action, a complete snapshot can be taken, timestamped, and sent to the downstream systems.

[0077] As required a snapshot of the graphics state along with the timestamp of the snapshot are sent to all connected downstream devices. Furthermore, at every tick, the frame signature, timestamp and tracking data can be sent to all connected downstream devices.

[0078] In this example, downstream devices are provided with a graphics snapshot when they first connect (or re-connect) to the upstream device, in order to make the transition between the devices as streamlined as possible between downstream and upstream. Notably (as shown in Figures 6 and 7, for example) the devices are generally physically separate as this can provide redundancy against connectivity issues.

[0079] In one example, in order to enable client specific branding as a targeted advert, a generic placeholder for graphical elements can be inserted in the content. An operator can then be informed that there is a graphic being provided in the process as it’s a targeted advert. Thus, a targeted advert graphics feed can be added which contains generic place holder locations for graphical elements rather than specific details around which artwork is selected. Example: Downstream

[0080] Figures 3 to Figure 5 show an example process flow for a downstream device. In these examples, as shown in Figure 3, the downstream device receives data from the upstream at step 300 at step 305, the data can be stored in a data store or the like, before the process moves to the process of Figure 4A / 4B at 315. Also, the downstream device can receive either content video feed with graphics at step 310 or content video feed without graphics at step 315, before the process continues to the example process of Figures 4A / 4B at 315. Typically, after the process of Figures 4A / 4B, the signal being transmitted (the TX signal) has been augmented with at least a subset of graphical content at step 320.

[0081] Referring further to Figure 4A, at step 400, the system can regenerate the frame signature received for a frame, at step 405 the system can determine whether virtual insertion is required. In order to find the frame where insertion is required, the process goes to step 410 or 415. That is, at step at step 410 the system finds an exact match, or at step 415 if no exact match is found, the system, at step 420 checks the neighbouring frames or visual effects to determine a frame match.

[0082] Once the frame has been determined, the process can move to step 425 where the timestamp and tracking data associated with the frame are queried. At step 430, the associated graphics state and snapshot are determined. At step 435, it is determined whether virtual content is required. If no content is required, the process moves to step 440 and the frame can be replayed with no virtual content added. If virtual content is required, the process proceeds, at step 445 to the process in Figure 4B and / or Figure 5.

[0083] Referring to Figure 4B, in this example, the process continues from Figure 4A at step 450, where the system, at step 455 selects a graphics snapshot, at step 460 determines a production graphics state, and at step 465 selects tracking data. At step 470, the system can then insert virtual content such as branding and production graphics as required. At step 475, the system can generate a final augmented output feed.

[0084] Figure 5 shows an example of the process continuing from Figure 4A at step 500, where at step 510 it is determined whether a target advert is required. At step 520 a generic graphical snapshot is selected and generated. At step 530, a primary key mask is generated and at step 540 the process converts the tracking data such that at step 550, the system can generate secondary key mask including primary key mask and production graphics.

[0085] Notably, the secondary key allows the system to avoid / minimise inserting content on production graphics. Typically, the virtual insertion of content sits underneath the production graphics. Accordingly, the secondary key can identify where the production graphics are on an image. The primary key can identify the physical regions where the content can be inserted. For example, rectangle on the grass where there might be space for inserting content. Typically, the two masks are combined for rendering for the final image.

[0086] It will be appreciated that a downstream device can be connected to one or more upstream devices. There is at least one upstream device for every camera that requires virtual branding - this can be typically, two cameras but can also include any multiple such as up to six cameras. The system can include an upstream device per camera and a downstream device per feed being generated.

[0087] As input the downstream device can receive:

[0088] - data feed per upstream device;

[0089] - video feed of the content being transmitted (TX signal) - final signal out of the broadcast that has all of the graphics; and,

[0090] - video feed of the content being transmitted without any production graphics (clean signal) and additionally another video feed but without any graphics overlaid (for example, no scores, etc).

[0091] A downstream device is configured to augment a TX signal by adding graphical elements for a particular subset of content called a feed. The feed can then typically include branding for a particular market segment (eg. Australian feed, UK feed, European feed). Each feed is generally a subset of the addressable market for the distribution of the content.

[0092] In one example, when data is received from an upstream device a unique timestamp is assigned. This is generated by combining a per upstream reconnection counter with the received timestamp from the upstream device. The upstream device re-connection count is serialised thus allowing a unique timestamp to be generated across upstream and downstream device restarts.

[0093] According to one example, there are three components to the unique timestamp: monatomic increasing counter from upstream, and from downstream, and also a tracked re-connection count. Once the unique timestamp is produced the data along with the timestamp is stored into a database (including the frame signature). Furthermore, data is typically stored either indexed by the unique frame signature or by the combined unique timestamp depending on the usage requirements for the data.

[0094] In this example, two data streams can come from the upstream device: - high frequency data from every frame (time stamp, frame signature and tracking data) - stored in database - where the key is the frame signature); and,

[0095] - operator driven data - stored against the unique timestamp.

[0096] If a target content is required and a video frame presented to the target content system is determined, via the above criteria, to be from a camera requiring a virtual insertion, the system can proceed as follows:

[0097] 1. The relevant generic graphics snapshot is selected (for example, generic target advert snapshot) and minimal snapshot describing the graphic element locations is produced. This generally includes a bounding rectangle and texture coordinates (allowing the image to be shifted around within the particular bounding rectangle) for each graphical element.

[0098] 2. A key mask is produced that indicates where virtual content can be inserted. This is typically a final composited key mask that combines all of the key masks from all of the technologies in use

[0099] 3. Tracking data is converted into a format suitable for use by end client devices that are performing the rending (project / view matrix, lens distortion data, image centre, aspect ratio).

[0100] 4. Production graphical effects (graphics overlays, fades, replay transitions, etc) are summarised and converted to a key mask in point (2) that is combined with the key mask generated in (2) above to produce a final combined key signal.

[0101] 5. If a video frame doesn’t match then it is passed through unchanged and a blank key signal is generated. Placeholder values are generated for the tracking data to maintain a constant data rate.

[0102] Thus the final output for each input frame for targeted advertising from the downstream device is typically:

[0103] 1. Tracking data in a form suitable for a graphics framework

[0104] 2. Bounding boxes describing graphics insertion locations

[0105] 3. The input TX video signal passed back out

[0106] 4. A key mask

[0107] Thus, the output can be summarised as being a minimal version of tracking data, key mask and graphics snapshot. This allows a minimal feed that can then be consumed by a client device such as a mobile phone.

[0108] The exact format of the output signals typically depend on the distribution platform. This could include but is not limited to multiple SDI (Serial Digital Interface, a broadcast video standard) signal of TX and key signals accompanied by a TCP data stream. Alternatively, an encoded video stream of various possible formats with tracking data embedded therein can also be used.

[0109] Journey of Video Frame into the system:

[0110] Input video frames are buffered for a sufficient amount of time that the data sent from the upstream device has been received and stored (allowable time window for any network issues). The amount of time required depends on the specific implementation but is typically in the order of a few tens of tick’s worth of time (where a tick is a video frame).

[0111] For each frame of video presented to the downstream device the unique frame signature is re-computed. This signature is compared against all the cameras currently configured for virtual insertion, in order to establish if reconfiguration is required and if virtual content insertion is required. Notably herein the multiple frames are referred to as primary and secondary image frames. It will be appreciated that if there are more than two cameras, the number of frame comparison can increase.

[0112] In ideal circumstances (no compression, format changes or production effects) the unique frame signature perfectly matches a previously received entry in the database. In this case the match can be directly accepted. An example algorithm that can be used is the hamming distance algorithm.

[0113] Typically, the matching distance between signatures in the database and the presented frame are computed via the hamming distance as the frame signature is typically a binary string.

[0114] In the case of a non-perfect match additional matching is performed, which can include:

[0115] - Examining a small time window of matches to perform temporal alignment and correction (by looking at neighbouring video frames to see if they match). The assumption employed is that neighbouring frames should be close to each other in time. Where time is measured in terms of the tick counter supplied by the upstream device. Special consideration is needed to handle replays where time can be stationary, flowing backwards or skipping in large jumps.

[0116] - Explicitly checking for visual effects like cross fading and wipes by examining the matching distance over time. During a fade the distance metric displays a mostly linear ramp from a high to low value or vice- versa depending on the source and destination of the fade. Look at hamming distance over time - you can then compute the percentage of the fade, and work out which cameras are involved in the cross-fade. Once a corrected match is obtained a final threshold check can be used to determine if the presented frame matches a camera feed requiring virtual insertion. Additional checks on the hamming distance to make the determination

[0117] If a match has been determined then the frame signature can be used to:

[0118] 1. Query the timestamp and tracking data associated with the particular frame (if any);

[0119] 2. Once the timestamp is retrieved then the complete graphics state snapshot that the timestamp falls within can be obtained. As graphics updates happen sporadically, you then have to go back to a last snapshot to grab the relevant state; and,

[0120] 3. Any other technologies can retrieve the necessary data either via the frame signature or timestamp depending on the nature of the data.

[0121] If a video frame presented to the downstream system is determined, via the above criteria, to be from a camera requiring a virtual insertion:

[0122] 1. The graphics snapshot is selected, the relevant subset of graphical elements for the feed being produced is extracted and then activated for rendering;

[0123] 2. The state of the production graphics is determined (fade, replay transition, graphical overlay) by predictions from the database lookup above and by comparing the two input video signals provided (clean and TX - version that has no graphics and those that do); and,

[0124] 3. The graphics state snapshot subset, production graphics state and tracking data is combined to generate a final augmented output feed. Accordingly, the system allows virtual branding to be inserted and the production graphics are overlaid over the top.

[0125] If a video frame does not match, then it is passed through the device unchanged. For downstream content production the final composited feed is provided to the broadcaster for onward distribution and transmission, typically as an SDI signal. Notably, it will be appreciated that matching and storing is primarily for replays.

[0126] Figure 6 shows an example of a system of the present invention. Figure 6 shows at 600 the broadcast production / upstream device generates three separate data streams - a clean signal, the transmission signal (TX), and metadata from the broadcast. The three data streams / signals are sent to a target advert downstream device 610.

[0127] If there is no target advertising to be added, a typically branded transmission signal can be generated and can be put through standard encoding at 620 before being sent to a consumer device or app at 660. Standard encoding can be applied via any known stream encoding software and / or hardware.

[0128] For target advertising, the system can send the transmission signal, key mask, and metadata to augmented encoding for per client targeting at 630 which produces an augmented stream for the client device at 660. Alternatively, or in addition, the system can transmit the transmission signal, key mask and metadata, for segmented targeting at 640. Segmented targeting can include a stream render which can generate demographic targeted branding streams and / or a standard stream. If targeted advertising is required, a segment ID is sent to an advert server at 650 which communicates with the client device 660 to match a unique ID / profile of the client to appropriate or associated adverts, provide movies / images and a determined location and track impressions. The branding artwork is then provided by the advert server 650 such that it can be displayed by the client device at 660. Notably, the decoder can extract separate video parts (such as video, key tracking data) from the stream and decode the compressed video.

[0129] According to a further example, a typical broadcast that would employ the targeting advert system and method that has been described herein can have multiple cameras situated around the field of play. According to one example, the majority of the live coverage is typically performed from between one and three cameras situated in the grandstand at the halfway line. An example of the system described herein showing cameras 710A and 710B is Figure 7.

[0130] As an example, Figure 7 shows that each camera that requires virtual branding can have an upstream system / device connected to it, and each respective upstream device is then connected to the downstream device as described herein. Thus, the system and method described herein can have multiple cameras where feeds from multiple cameras can be augmented with additional virtual content.

[0131] Notably, Figure 7 also shows a mixing desk where an image frame including a clean signal can be received by the mixing desk and the mixing desk can transmit a clean signal (i.e. an image frame with no production graphics) and a transmission signal with production graphics. Typically, this is referred to herein as a primary image frame.

[0132] Accordingly, the system can augment a broadcast with virtual content, and as shown for example in Figure 7, the system can include a client device for viewing the broadcast, at least one camera for receiving a primary image frame of the broadcast, at least one upstream device in communication and associated with the at least one camera; and, at least one downstream device in communication with the at least one upstream device. The system can then determine client identification information associated with the client device, determine one or more virtual content based on the client identification information; and augment the primary image frame with the virtual content to be displayed on the client device.

[0133] The upstream device can be configured to then receive the primary image frame, generate downstream data stream associated with the primary image frame including any one or a combination of a unique frame signature associated with the primary image frame, a unique upstream timestamp for received primary image frame, frame tracking data for received primary image frame, and may also include insertion instructions in relation to where to place the content on the primary image frame.

[0134] The downstream data stream can then be stored in a data store and be sent to the downstream device.

[0135] The downstream device is then configured to receive downstream data stream from the upstream device, receive a secondary image frame, compare with the down stream data and, augment the primary image with graphical content.

[0136] The downstream device can then also receive the secondary image frame at a time interval after the upstream device receives the clean signal, regenerate a frame signature for the secondary image frame, query timestamp and tracking data associated with secondary image frame; and, query the data store and compare with the primary image frame and determine whether virtual insertion is required based on the timestamp and tracking data and insertion instructions.

[0137] Notably, when the signal arrives at the downstream system, there has been time for the upstream to have seen the frame and stored in the datastore.

[0138] Accordingly, multiple versions of the content can be inserted depending on the multiple cameras. For example, in a game of tennis, the system can produce different feeds for different geographical areas of a viewer. That is, content can be added to the feed depending on where the viewer is.

[0139] Example: Two Virtual Camera Rugby Union

[0140] Taking as an example a Rugby Union match where virtual advertising is required on two cameras - cameras 1 and 2 (as shown for example, in Figure 7):

[0141] Camera 1 Live

[0142] As an example, tracking a single frame of video on camera 1 for the timestamp 100, the following steps can be performed by the system described herein: On the upstream system:

[0143] - A hashing algorithm is used to produce a bitstring hash for the video frame

[0144] - The tracking data (camera extrinsics and intrinsics) for timestamp 100 is obtained and bundled with the timestamp and frame hash and sent to any connected downstream systems

[0145] - At timestamp 100 no operator driven changes to the graphic state have been made so no additional data needs to be sent

[0146] On one of the downstream systems configured to communicate with the camera 1 upstream system:

[0147] - The data for timestamp 100 will be received and decoded

[0148] - A globally unique timestamp will be computed by combining the received upstream tick, a connection counter and the unique timestamp of the downstream system

[0149] - The tracking data and global timestamp is stored in a database using the received hash as the index

[0150] Notably, a globally unique timestamp can allow the system to spool back the graphic state as it was when the live feed was showing. That is, if there is a replay, all the graphic instructions are wound back to exactly as they were in the original clip. Typically, the upstream and downstream systems have two independent counters. Thus, at the point the data arrives at the downstream system, the system typically takes the two counters and combines them into a globally unique timestamp for replays.

[0151] The broadcast is currently using camera 1 live, so around 1 second after receiving the data for tick 100 the frame that was previously presented to the upstream system will be presented to the downstream system. Two versions of the frame will be presented:

[0152] 1. Clean - no production graphics overlays present

[0153] 2. TX - the final production frame with additional elements overlaid on the frame from timestamp 100

[0154] The hashing algorithm is applied to the clean signal and used to perform a lookup in the previously mentioned database.

[0155] In this example, the best matching data is selected from the database using a hamming distance comparison between the computed and stored bitstrings.

[0156] In this case there are no significant production effects being applied (for example fades) so there is a direct match to the stored bitstrings.

[0157] Given that a match for a camera requiring virtual insertion has been identified the following actions occur:

[0158] - The appropriate graphical state snapshot for camera 1 and tick 100 is retrieved

[0159] - Any changes in the snapshot compared to the currently loaded graphical state are applied

[0160] As the appropriate graphical state is loaded and the tracking data is available virtual insertion can be performed to produce a final frame for output via SDI.

[0161] Camera 2 Replay with Targeted Advertising

[0162] In this example, a downstream system that is configured to produce targeted advertising downstream is considered.

[0163] The broadcast is currently using a replay of camera 2 from an earlier point in time. When the incident happened live the frame currently under consideration had a timestamp of 500 on the upstream system.

[0164] When camera 2 was viewing the incident live the same flow of events as per the first two paragraphs for the Camera 1 Live case will have occurred. Hence the downstream system currently under consideration has the data for timestamp 500 from camera 2 already stored.

[0165] Two versions of the replay from timestamp 500 on camera 2 will be presented:

[0166] - Clean - no production graphics overlays present

[0167] - TX - the final production frame with additional elements overlaid on the frame from timestamp 100

[0168] The hashing algorithm is applied to the clean signal and used to perform a lookup in the previously mentioned database.

[0169] The best matching data is selected from the database using a hamming distance comparison between the computed and stored bitstrings.

[0170] In this case compression has been applied to the frame due to being stored on the replay device. This results in a non-perfect match and neighbour temporal filtering is applied to ensure that the best and correct match is located.

[0171] Given that a match for a camera requiring virtual insertion has been identified the following actions occur:

[0172] - The appropriate graphical state snapshot for camera 2 and tick 500 is retrieved and the targeted advertising subset is selected

[0173] - Any changes in the snapshot compared to the currently loaded graphical state are applied The combination of target advertising placeholders, tracking data and production graphics key mask (produced by comparing the Clean and TX signals) are encoded as per the targeted advert stream specification and can be sent onwards for distribution to client devices.

[0174] Camera 3

[0175] The broadcast is currently using camera 3 live, a frame is presented to the downstream system. Two versions of the frame will be presented:

[0176] - Clean - no production graphics overlays present

[0177] - TX - the final production frame with additional elements overlaid on the frame

[0178] The hashing algorithm is applied to the clean signal and used to perform a lookup in the previously mentioned database.

[0179] The best matching data is selected from the database using a hamming distance comparison between the computed and stored bitstrings.

[0180] In this case there will be a large minimum hamming distance between the frame under consideration and the data available in the database.

[0181] Given that there is no match for a camera requiring virtual insertion the following actions occur: the graphic state of graphics to be inserted is cleared and as the empty graphical state is loaded the final frame will be identical to the input frame supplied to the downstream system.

[0182] Accordingly, there is provided herein systems and methods whereby the viewer can be provided with customised content whilst viewing a broadcast of an event, or the like.

[0183] It will be understood to persons skilled in the art of the invention that many modifications may be made without departing from the spirit and scope of the invention.

[0184] In the claims which follow and in the preceding description of the invention, except where the context requires otherwise due to express language or necessary implication, the word “comprise” or variations such as “comprises” or “comprising” is used in an inclusive sense, i.e. to specify the presence of the stated features but not to preclude the presence or addition of further features in various embodiments of the invention.

Claims

Claims1. A system for augmenting a broadcast with virtual content, the system including: a client device for viewing the broadcast; at least one camera for receiving a primary image frame of the broadcast; at least one upstream device in communication and associated with the at least one camera; and, at least one downstream device in communication with the at least one upstream device; wherein the system is configured to determine client identification information associated with the client device, determine one or more virtual content based on the client identification information; and augment the primary image frame with the virtual content to be displayed on the client device.

2. The system of claim 1 , where the upstream device is configured to: receive the primary image frame; generate downstream data stream associated with the primary image frame including any one or a combination of: a unique frame signature associated with the primary image frame; a unique upstream timestamp for received primary image frame; frame tracking data for received primary image frame; and, and insertion instructions; store the downstream data stream associated with the primary image frame in a data store; and, send the downstream data stream to the downstream device.

3. The system of claim 2, wherein the downstream device is configured to: receive downstream data stream from the upstream device; receive a secondary image frame; compare with the down stream data and, augment the primary image with graphical content.

4. The system of claim 3, wherein the downstream device is further configured to: receive the secondary image frame at a time interval after the upstream device receives the clean signal;regenerate a frame signature for the secondary image frame ; query timestamp and tracking data associated with secondary image frame; and, query the data store and compare with the primary image frame and determine whether virtual insertion is required based on the timestamp and tracking data and insertion instructions.

5. The system of claim 4, wherein the secondary image frame is generated by any one of or a combination of: a transmission signal with graphics; and, a transmission signal without graphics.

6. The system of any one of claims 1 to 5, wherein the primary image frame is generated by a clean signal transmitting the image frame.

7. The system of claim 2 wherein the upstream system applies a hashing algorithm to the video frame.

8. The system of claim 2, wherein the tracking data includes any one or a combination of intrinsic camera data and extrinsic camera data.

9. The system of claim 2, wherein the upstream system further confirms whether operator changes are required to the primary image frame and updates the primary image frame accordingly.

10. The system of claim 3, wherein the downstream system further generates a globally unique timestamp by combining the received upstream timestamp, a connection counter and a unique downstream timestamp.

11. The system of claim 10, wherein the downstream system stores the globally unique timestamp in the datastore against the tracking data.

12. The system of claim 3 and 4, wherein the secondary image frame includes any one or a combination of a clean image frame and a final production image frame.

13. The system of claim 12, wherein when the secondary image frame is compared to the primary image frame, if there is a match, it is determined whether virtual insertion or targeted content is required, and the primary image frame is augmented accordingly.14.The system of claim 13, wherein if there is no exact match, either the primary image is not augmented, or neighbouring frames are checked todetermine whether there is a match and if the primary image is to be augmented.

15. The system of claim 4, wherein the time interval is 1 second.

16. The system of any one of claims 1 to 15, wherein the system has a plurality of cameras, each of the plurality of cameras having an associated upstream system.

17. The system of claim 16, wherein the downstream system includes a processing system and a data store associated with each of the plurality of cameras.