Dynamic edge watermarking
Patent Information
- Application Number
- GB2024004940
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-05
- Publication Date
- 2026-01-07
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND An encoded video signal is encoded by one or more encoding devices and can be decoded by one or more decoding devices. The encoding is typically carried out by a content provider or creator with decoding typically carried out on a platform local to or accessibly locally to an end user. There may be one or more intermediaries between the content provided or creator and the end user that may transcode, transfer or otherwise process the content as part of receiving and redistributing the content. An enhancement encoded video signal can be decoded (by a decoder) at different levels of quality (e.g., at different spatial dimensions). This is advantageous because it means that a single encoded video signal can be sent to many types of decoding devices (each having different operating capabilities), and each device can decode the encoded video signal in line with the operating capability of the decoder or available bandwidth, et cetera. For example, a first decoding device may only be able to decode and render an encoded video signal using a base codec whereas a second decoding device may be able to decode and render an encoded video signal using a base (or “base layer”) codec and an enhancement (or “enhancement layer”) codec. Creators and content providers of data streams, such as video signals or audio signals, often seek to protect the content produced. Significant effort is therefore applied to content protection. This includes applying digital rights management (DRM) and / or encryption of data streams. As part of the content protection, from time to time, a watermark is applied to, for example, video signals. When used, watermarking can be applied at the source of the signal. This means that any party that seeks to distribute that content will only be able to distribute watermarked content. This similarly applies to DRM and encryption and other forms of content protection. However, none of these forms of content protection assist in identification of the distributor. Watermarking an identifier of the unauthorised party who receives the distributed content is not generally possible because the unauthorised party will typically use encoding software they control, and which is often derived from open source software. In view of this creators and content providers cannot ensure any watermark continues to be applied or is applied at all during the encoding process. It is not possible, at any significant scale, to leverage data stream access points to add party-specific watermarking, such as embedding unique identifiers of an accessing party into each stream when provided to a recipient. This is because it entails significant processing load at the edge, which for the case of modifying and re-encoding is excessively taxing. Further, that re-encoding may compromise quality in a manner unacceptable to the content provider, creator or intended recipient. This is due to the re-encoding after adding watermarking likely not being tuned specifically to the content in question, in comparison to well-tuned encoding produced by the content provider or creator, as well as repeated re-encoding introducing losses and artefacts. There is thus a need for enhanced content protection for encoded video signals. SUMMARY OF INVENTION According to a first aspect, there is provided an apparatus suitable for transferring a video signal to be output as an encoded video signal encoded including a base encoded stream and at least one enhancement stream (the enhancement stream comprising one or more layers of residual data, the residual data being generated based on a comparison of data derived from a decoded version of a base encoded stream and data derived from an input video signal), the enhancement stream being suitable for combining with the base encoded stream to reconstruct the input video signal, and each enhancement stream comprising respective frames, the apparatus being configured to: generate an adjusted enhancement stream by providing, in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream, a deviation, the deviation being (visually) identifiable in an video signal constructed using a base encoded stream and the adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream the deviation thereby being characteristic of (or unique to) the output destination; and output the encoded video signal with the adjusted enhancement stream and the base encoded stream. This provides (binary) data in the frames of the enhancement stream, with the data providing a unique identifier for each output of an encoded video signal, which allows a specific distributor of a decoded video signal to be identified. This is achieved while limiting additional processing load and quality degradation since it is applied in an enhancement stream that already provides quality enhancement and providing a deviation limits the decoding needs for the enhancement stream. By the phrase “output destination”, it is intended to mean unique user and / or device and / or time and / or location to which the encoded video signal is provided. The identifier may be characteristic of the output destination by identifying it, such as by providing a specific ID of a user or user device, or it may be characteristic in a weaker sense, by providing a timestamp the stream was served at or an edge node from which the stream was served. With a unique identifier being embedded in the encoded video signal due to the deviation being application in the adjusted enhancement stream, this thus makes a corresponding identification possible should content including the encoded video signal appear from an unauthorised source. The deviation is optionally visible in the resultant decoded video, produced by inputting the output encoded video into a decoder. In this way, the deviation may be resilient to methods of watermark removal, including removing metadata from captured video, re-encoding captured video, rescaling or introducing noise. Typically, the at least one frame is a plurality of frames. While the deviation is able to be provided in (only) a single frame, this limits the amount of information or data provided by the deviation. By providing the deviation over a plurality of frames, more data can be provided in the enhancement stream, whilst having the additional benefit of distributing the processing required to produce the deviation over time, reducing computational complexity. Typically, the plurality of frames may be a part of a sequence of frames. Further, the sequence of frames may have at least one intervening frame (to which the deviation is not applied) between each frame of the plurality of frames. While the plurality of frames may be a plurality of adjacent or directly consecutive frames in a sequence, having intervening frames spreads the processing load reducing the effect of including the deviation. This also makes the deviation more difficult to detect by a third party because it is distributed rather than provided in a single sequence. Further, this has the benefit of reducing any perceived visual impact of the deviation, as it persists for the shortest possible period of a single frame. There are examples where consecutive frames of a pair of frames of the plurality of frames are (directly) consecutive in the sequence of frames instead of having any intervening frame therebetween. Typically, the deviation may be a variation of a luminance of a whole frame or one or more portions of a frame. An alternative would when the deviation may be a variation in one or more colour planes. By having the deviation as a variation in luminance the impact on the video quality and appearance to the user is limited while still allowing the variation to be detectable. Making adjustments in one or more colour planes would have the appearance to a user of reducing the quality, so is less desirable. The deviation in the luminance may be the addition of a constant shift (i.e. an arithmetic shift, not a logical bit-shift) to the luminance of the frame. The shift may be positive or negative. Typically, a positive shift in the luminance may be interpreted as a binary “1", and a lack of a shift may be interpreted as a binary “0”. In this way arbitrary data may be encoded using shifts in luminance. Alternatively, a positive shift in the luminance may be interpreted as a binary “1”, and a negative shift in the luminance may be interpreted as a binary “0”. In this way, arbitrary data may be encoded using shifts in the luminance, whilst permitting unmodified frames to pass through without significance. The magnitude of such a constant shift in the luminance may be + / - 1% of the average luminance of the frame, for example. Such a shift in luminance may instead be a multiplicative shift instead of an additive shift. This may lead to a smaller visual impact on the resultant video without loss of efficacy. Such a shift may be represented as a 1.1 times multiplier representing a binary “1”, and a 0.9 times multiplier representing a binary “0”. Alternatively, a 1.01 times multiplier representing a binary "I”, and a 0.99 times multiplier representing a binary “0”. It will be understood that the multipliers should be selected to minimise the visual impact of the deviation, whilst maintaining a visible effect that can be captured and decoded to read the encoded data. Alternatively or in addition, the deviation may be used to encode more than simple binary data. A single luminance deviation may encode two bits of information using four allowed shifts: a large positive shift, a small positive shift, a small negative shift and a large negative shift. It will be understood that by using increasing numbers of allowed shifts of luminance in the deviation, increasing quantities of data may be encoded. Typically, the deviation is provided in a value of at least two portions of the at least one frame. While it is possible for the at least a portion to be a full frame, i.e. the at least a portion being the whole frame instead of only a portion of a frame, or only a single portion that has a size smaller than a whole frame (such as a tile or other correspondingly small area, since this would reduce any visual impact), providing the deviation in at least two portions of the at least one frame (i.e. a plurality of portions) of the at least one frame, this allows multiple values in a single frame to be varied when generating the adjustment enhancement stream. This allows the deviation to include a higher density of data per frame to which a deviation is applied. If deviations are applied to multiple frames, this also allows deviations to be applied to less frames to limit the interventions per adjustment enhancement stream that is generated. Typically, the apparatus may include a library of deviation data, and the apparatus being configured to generate an adjusted enhancement includes identifying, based on the intended output destination of an encoded video signal, relevant deviation data in the library and writing the relevant deviation data from the library into the adjusted enhancement stream. This limits computational load by avoiding the deviation data needing to be calculated for each adjustment enhancement stream. Typically, the encoded video signal may have an empty enhancement stream prior to outputting the encoded video signal, the apparatus being configured to output the encoded video signal by writing the adjusted enhancement stream into the empty enhancement stream. This allows the encoded video signal to be stored without an enhancement stream or without at least a part of an enhancement stream since there may be enhancements layers within the enhancement stream or a plurality of enhancement streams. This reduces storage capacity requirements for the encoded video signal and allows an identical base encoded stream to be provided for each output while allowing the (relevant) enhancement stream to be tailored to each output at a minimal computational load. Alternatively, the encoded video signal may be stored with an enhancement stream that is blank, zero or empty. As a first alternative, the apparatus may be configured to output the encoded video signal by overwriting any existing enhancement stream in the encoded video signal. This provides the same benefit as above of allowing the same base encoded stream to be provided for each output while allowing the (relevant) enhancement stream to be tailored to each output. The first alternative would be expected to have a higher computational load. However, this still provides more efficient modification of the video signal to provide a unique identifier that fully decoding the encoded video signal. As a second alternative, the apparatus may be configured to generate the adjustment enhancement stream and output the encoded video signal by decoding at least one layer of an existing enhancement stream of the encoded video signal, applying the deviation to the decoded at least one layer and reapplying the decoded encoding. An existing enhancement stream may be fully decoded (and re-encoded) or only an entropy encoded layer may be decoded (and re-encoded). This still allows the (relevant) enhancement stream to be tailored to each output. The second alternative would again be expected to have a higher computational load, though still significantly lower than decoding the base video layer, modifying it, and re-encoding it. However, this still provides more efficient modification of the video signal to provide a unique identifier that fully decoding the encoded video signal. Typically, the apparatus may be further configured to receive the encoded video signal. This allows the apparatus to only hold encoded video signals that are being output instead of holding many encoded video signals, which would require significant storage capacity. Typicaliy, the apparatus may be further configured to receive a content request from an output destination and retrieve the encoded video signal in response to the content request. This allows an encoded video signal to be provided as it is required instead of preparing it ahead of time, thereby reducing redundancy and storage requirements. Typically, the adjustment enhancement stream may be an LCEVC Level 1 stream. This can also be referred to the LOQ1 stream. While a LCEVC Level 2 or LOQ2 stream could be used, the Level 1 stream avoids the adjustment interfering with the temporal signalling applied in Level 2, potentially requiring a temporal buffer refresh for each subsequent frame to which an adjustment is applied if using Level 2. On the other hand, Level 1 typically does not include temporal signalling, and thus, is more efficient to use. Level 1 layers are also typically not upscaled, simplifying the implementation. Typically, the apparatus is further configured to encrypt or apply digital rights management (DRM) to the output encoded video signal. This may make the streams of the output encoded video signal unavailable to third parties that do not have access to a corresponding decoder. Further, this may prevent access by any third party to the encoded video signal entirely. Such a process may output a resultant decoded video signal as data suitable for display, without access to the undertying streams and structure of the encoded video. This therefore provides standard protection offered by encryption and DRM and also reduces the likelihood of the specifics of deviation from being discoverable without specialist equipment, further protecting the output encoded video signa! from unauthorised access or distribution. Further, this prevents removal of the enhancement stream before decoding and piayback, which prevents the removal of the deviation applied to the video. Typically, the deviation is readable as a plurality of bits. This provides a simple way of recording data in the enhancement stream while allowing the plurality of bits to represent more complex data. Typically the plurality of bits are a representation of a combination of one or more of: a user ID, account ID, time of output of the encoded video signal, other characteristic times, such as the time the video was requested, apparatus ID, IP address, and region of a user. A combination of, for example, user ID, account ID and IP address may be able to represented as a three-tuple by the plurality of bits. Further, the plurality of bits may represent a key to a database, in which is stored one or more of the elements listed in this paragraph. Using a combination of one or more of these elements, allows for the output destination of the encoded video signal to be accurately identified. Generally, the bits may represent metadata, i.e. a unique ID, and may be an indication of other metadata stored for look-up by the content provider. Typically, the apparatus is an edge device, the output being arranged to be output out of a network of which the edge device is a part, and / or from the edge device to a connected user device. This allows the apparatus to be as close as possible to where the output is being provided. Further, this reduces the resources the apparatus may need to hold, since the apparatus will only need to be configured to provide outputs to local users instead of to large quantities of the network’s users. This means any deviations or library only needs to be configured to be capable of serving local users. In other examples, the apparatus may be part of a content delivery network (CDN), or may be a single server or cloud server. Typically, the apparatus may be further configured, during the generation of the adjustment enhancement stream, to apply a frame number offset to the at least one frame to which the deviation is provided, the frame number offset being varied based on output destination. For example, for a first output destination, the deviation may be applied (first) to a (first) frame five seconds into the stream; for a second output destination, the deviation may be applied (first) to a (first) frame ten seconds into the stream, with further output destinations have a different offset. The frame number offset may be a time based offset or may be an offset of a number of frames, i.e. periodic. Alternatively, the deviation may not be applied at constant frame number offsets. Deviations may be applied to frame numbers following a more complex sequence, following no sequence at all, or randomly. The deviation need not be applied consistently between different output encoded video signals. Such deviations can still encode significant data, as the difference with the original, source video can be decoded by comparison. This may have the advantage of being more difficult to detect and / or remove by a third party. Typically, the apparatus may be further configured, during the generation of the adjustment enhancement stream, to repeat the deviation at one or more predetermined intervals over the length of the adjustment enhancement stream. This provides the deviation throughout the duration of the encoded video. There may be more than one predetermined interval or the predetermined may be a variable or non-constant Interval to reduce the likelihood of re-encoding, video capture or re-Interpolation of a frame rate causing the deviation to be removed either accidentally or deliberately. According to a second aspect, there is provided a method of transferring a video signal to be output as an encoded video signal encoded including a base encoded stream and at least one enhancement stream (the enhancement stream comprising one or more layers of residual data, the residual data being generated based on a comparison of data derived from a decoded version of a base encoded stream and data derived from an input video signal), the enhancement stream being suitable for combining with the base encoded stream to reconstruct the input video signal, and each enhancement stream comprising respective frames, the method comprising: generating an adjusted enhancement stream by providing, in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream, a deviation, the deviation being (visually) identifiable in a video signal constructed using a base encoded stream and the adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of (such as unique to) the output destination; and outputting the encoded video signal with the adjusted enhancement stream and the base encoded stream. The method of the second aspect may be configured to provide any one or more of the functions of the apparatus of the first aspect or any feature which the apparatus of the first aspect is configured to perform. According to a third aspect, there is provided a detector suitable for identifying an output destination of an encoded video signal represented in an independent video stream, the detector being arranged to: receive the encoded video signal; receive a frame range over which a deviation in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream is provided, the deviation being (visually) identifiable in a video signal constructed using a base encoded stream and an adjusted enhancement stream, the deviation being detenmined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of (such as unique to) the output destination; receive the representation of the encoded video signal in the independent video stream; compare, in the received frame range, the received representation of the encoded video signal and the received encoded video signa! to identify differences; derive the deviation provided to the encoded video signal represented in the independent video stream based on the identified differences; and identify the output destination of the encoded video signal of the representation of the encoded video signal based on the derived deviation. This allows the output destination, such as the (unique) user and / or device and / or time and / or location, to which the encoded video signal for which the representation Is provided in the independent video signal to be identified. This thus allows tracking of the origin of content should content including the encoded video signal appear from an unauthorised source. According to a fourth aspect, there is provided a method of detecting an identity of an output destination of an encoded video signal represented in an independent video stream, the method comprising: receiving the encoded video signa!; receiving a frame range over which a deviation in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream is provided, the deviation being (visually) identifiable in a video signal constructed using a base encoded stream and an adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of (such as unique to) the output destination; receiving the representation of the encoded video signal in the independent video stream; comparing the received representation of the encoded video signal and the received encoded video signal to identify differences; deriving the deviation provided to the encoded video signal represented in the independent video stream based on the identified differences; and identifying the output destination of the encoded video signal of the representation of the encoded video signal based on the derived deviation. According to a fifth aspect, there is provided a system suitable for transferring a video signal to be output as an encoded video signal and detecting an identity of an output destination of an encoded video signal represented in an independent video stream, the system comprising the apparatus of the first aspect and the detector of the third aspect, the detector being arranged to receive the frame range from the apparatus According to a sixth aspect, there is provided a data stream comprising: a base encoded stream; and an adjusted enhancement stream, the adjusted enhancement stream including a deviation in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream the deviation being (visually) identifiable in a video signal constructed using a base encoded stream and the adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of (such as unique to) the output destination, wherein the base encoded stream and adjusted enhancement stream are combinable to generate a constructed signal. According to a seventh aspect, there is provided a computer program comprising instructions which, when executed, cause an apparatus to perform the method according to the second aspect or fourth aspect or to provide an apparatus according to the first aspect or a detector according to the third aspect or a system according to the fifth aspect. According to an eighth aspect, there is provided a non-transitory computer-readable medium comprising the computer program according to the seventh aspect. BRIEF DESCRIPTION OF DRAWINGS Various examples are described in detail below with reference to the accompanying figures, in which: Figure 1 shows a known, high-level schematic of an LCEVC encoding, decoding and transport process; Figure 2 shows a schematic of an example system; Figure 3 shows a first example process; Figure 4 shows a second example process; and Figures 5A, 5B and 5C show example frame constant values and example deviations. DETAILED DESCRIPTION For context, we will begin by describing an enhancement coding scheme, LCEVC, suitable for use with the concepts of the present disclosure. LCEVC will be described in the context of Figure 1. Throughout the description the terms hierarchical coding and enhancement coding may be used interchangeably and while examples are described in the context of LCEVC, it will be understood that the concepts described may be suitable to any similar hierarchical or enhancement coding scheme. This may include concepts described herein also being suitable for other hierarchical coding schemes such as VC-6 or SHEVC. LCEVC adopts a multi-layer approach where any base codec (e.g. H.264, HEVC, AVI and others), is enhanced via an additional low bitrate stream. LCEVC’s data stream structure is defined by two component streams: a base stream decodable by a hardware decoder; and, an enhancement stream consisting of one or two enhancement layers suitable for software processing implementation with sustainable power consumption. The enhancement provides improved compression efficiency to existing codecs, and reduces encoding and decoding complexity, for on demand and live streaming applications. Figure 1 below illustrates how LCEVC operates on both encoding and decoding pipelines. The base encoding (whether H.264, HEVC or others) is performed on a down-scaled input at a lower resolution, typically a quarter of the desired output resolution. LCEVC enhancement data is calculated at the two resolutions providing two levels of correction and enhancement. The LCEVC encoder generates the enhancement stream from two inputs: the base encoding and the original uncompressed full resolution video, effectively correcting the quality gap between the two. The LCEVC data can be packaged together with the base elementary stream (for example as Supplemental Enhancement Information (SE!) of the Network Abstraction Layer, NAL), as frame metadata in a WebM container or in an additional data Packet Identifier (PID) in a MPEG-2 TS stream. The LCEVC decoder works at an individual video frame level. As input it takes the decoded low-resolution picture from the base video decoder, which is typically provided by a hardware decoder on the device, and the LCEVC enhancement decoded in software to produce a full-resolution picture ready for rendering on the display view. Example implementations of decoding LCEVC are set out in WO 2022 / 023739 and WO 2023 / 118851, which are incorporated by reference. As illustrated in figure 1, in the encoder 100, an input full resolution video, i.e. a source video 102, is processed to generate various encodings. A first encoding (base encoding 110) is produced by feeding a base encoder 106 (e.g., AVC, HEVC, VP9, or any other codec) with a down-sampled version of the input video, which is produced by down-sampling 104 the input video 102. In the example shown, the downsampling is to a quarter resolution but this is optional as will be elaborated on elsewhere. The base encoding 110 may be referred to as a base layer. A second encoding (level 1 encoding 112, an example of an enhancement encoding) is produced to create first level corrections 116 by applying an encoding operation to the residuals obtained by taking the difference between a reconstructed base codec video and the down-sampled version of the input video. The reconstructed base codec video is obtained by decoding the output of the base encoder 106 with a base decoder. In typical implementations, the level 1 encoding 112 is optional. This level 1 encoding 112 may be referred to as a first enhancement layer. A third encoding (level 2 encoding 114, another example of an enhancement encoding) is produced to create first level corrections 120 by processing the residuals obtained by taking the difference between an up-sampled version (i.e. normative upsampling 118) of a corrected version of the reconstructed base coded video and the input video 102. This level 2 encoding 114 may be referred to as a second enhancement layer. The enhancement layer(s) and the base layer are typically combined (i.e. as illustrated by mux 122) and the full resolution video is encoded in layers. This is then typically transmitted using standard packaging and transmission protocols, an example of which will be described below in the context of Figure 2. While often the corrections forming the enhancement layer involve an upsampling to higher resolution, generally any improvement in quality may be provided by the enhancement layer. For example, resolution, visual quality (VQ), bit depth (e.g. 8 to 10b) or colour space (e.g. HDR). At the decoder 140, the encoded video is separated (i.e. as iMustrated by demux 142) into an ancillary data stream 144 and a video stream 146. The decoder receives the layers (a base encoding, an optional level 1 encoding and a level 2 encoding) together with headers containing further decoding information. The base encoding, i.e. in video stream 146, is decoded by a base decoder 148 corresponding to the base decoder used in the encoder. At an enhancement decoder 150, which receives headers and the enhancement layers in the ancillary data stream 144, its output is combined with the decoded residuals obtained by decoding the level 1 encoding (if present). The combined video is up-sampled and further combined with the decoded residuals obtained by applying a decoding operation to the level 2 encoding to output the full resolution video 152. In various examples, the encoded video is passed directly from the encoder 100 to the decoder 140. However, in a number of examples, as shown in Figure 1, the encoded video is passed to an intermediary location 1000. The encoded video is able to be held in the intermediary location until it is wanted by an output location. As shown in Figure 2, this includes an example intermediary location 1000. in Figure 2, this is illustratively identified as a cloud based arrangement. However, this can be provided by alternatives, such as a single computer or server that is able to provide the same general functionality and capable of storing content. Generally illustrated at 200 in Figure 2 is a system. This includes the intermediary location 1000 as identified above, which is an apparatus suitable for transferring a video signal to be output as an encoded video signal encoded. The encoded video signal, in various examples, includes a base encoded stream and at least one enhancement stream, such as an encoded video signal including a base encoded stream and an enhancement stream corresponding to a stream encoded using level 1 encoding referred to above. The system 200 also includes a detector 2000 connected to the intermediary location 1000. In Figure 2, the connection between the detector and the intermediary location is shown by a wired connection. However, this connection may be provided by a wireless connection and / or over a network with wired and / or wireless connections. In Figure 2, various output devices 1520a, 1520b, 1520c are shown connected to the intermediary location 1000. The output devices are able to connect to the intermediary location whatever form it takes. As with the detector 2000, in Figure 2, the connection between the each output device and the intermediary location is shown by a wired connection. However, this connection may be provided by a wireless connection and / or over a network with wired and / or wireless connections. In the example shown in Figure 2, each output device 1520a, 1520b, 1520c is shown connected to an interface 1002a, 1002b. When the apparatus providing the intermediary location 1000 is a network, cloud-based system or platform, such as a content delivery network (CDN), each interface may be provided by an edge server. In other examples, the interface is another suitable device, component or port by which the output devices are able to connect to the apparatus. While two interfaces 1002a, 1002b are shown in the example of Figure 2, the number of interfaces may vary from example to example, as may the number of output devices 1520a, 1520b, 1520c, and the number of output devices connected or connectable to each interface. Each interface 1002a, 1002b is connected to a storage unit 1004. This may be a server, computer, hard disk drive, solid state drive or some other form of memory or storage. The capacity of the storage unit may be fixed or may be variable. While the connection between each interface and the storage unit is represented by a wired connection in the example of Figure 2, the connection between each interface and the storage unit may be provided by a wireless connection and / or over a network with wired and / or wireless connections. Further, the components ofthe apparatus providing the intermediary location 1000 have a distributed functionality in some examples. As such, the location of each device and / or each functionality may not be fixed, may vary and / or may be redistributable. The output devices 1520a, 1520b, 1520c typically include some form of decoder, such as the decoder 140 shown in Figure 1. The output devices of the example shown in Figure 2 thus represent the decode side of Figure 1. Turning to the overall functionality of the system 200, intermediary Iocation1000 and detector 2000, regardless of their specific form, in some examples, the apparatus providing the intermediary location, is capable of outputting an encoded video signal. In various examples this is output to an output device. In a number of examples, in order to allow track or later identification of where or which user, device or location the encoded video signal is output to, a watermark is added to the video signal. This watermark is tailored to wherever the specifics of the output are. Watermarking is known per se, such as in WO 2021 / 064414, WO 2021 / 079147, WO 2021 / 064412 and WO 2023 / 187374, each of which is incorporated by reference. At a general level, the watermarking is achieved by generating an adjusted enhancement stream that is output as an encoded video signal with a base encoded stream. Consistent with the aspects disclosed herein, watermarking in accordance with these aspects is implemented, in some examples, by deliberately modifying a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream to cause that value to deviate from what it would be without the watermarking. This can be referred to as “a deviation”. As referred to above, the deviation introduces a desired variation of one or more stream content parameters, such as luminance, chrominance, audio level, frequency, or one or more other parameters. This can either be applied to a whole frame or whole display or instead to a one or more (smaller) portion of the frame or display. The deviation is preferably specific to the output or group of outputs. As such, in some examples, once the enhancement stream has been created, which may be generated from a library of pre-prepared entropy encoded streams (in typical examples enhancement streams) held in the interface 1002a, 1002b, storage unit 1004 or elsewhere, a mapping into binary numbers is generated from a unique identifier. This is one or more of a user ID, account ID, time of output of the encoded video signal, apparatus / servebedge server ID, IP address, and region of a user. These are then merged into one or more video frames by applying the deviations to represent binary values. The apparatus 1000 may generate the adjusted enhancement stream the interface device 1002a, 1002b, in the storage unit 1004 or elsewhere, or, as noted above, may include pre-prepared entropy encoded streams (in typical examples enhancement streams) in a library. The base encoded stream may be provided from the storage unit 1004 or elsewhere in use. This may be provided with or without an enhancement stream. Where the base encoded stream is provided without an enhancement stream, the adjusted enhancement stream is provided in place of an enhancement stream. When an enhancement stream is present, such as when level 1 encoding 112 is provided by the encoder 100, the enhancement stream can either be overwritten to replace it with the adjusted enhancement stream. Alternatively, the enhancement stream can be at least partially decoded, such as by conducting entropy decoding, the deviations can be applied, and re-entropy encoded. Once provided with the base encoded stream, the adjusted enhancement stream and the base encoded stream are provided in various examples as an output in the form of an encoded video signal in some examples. This is then output to the relevant output device. This may be due to the relevant content having been requested from the output device 1502a, 1502b, 1502c or by a particular user, with the content having been provided from the storage unit 1004 or elsewhere in response to a request for the content from the output device. Considering various specific examples, the deviation provides a form of metadata. The encoded metadata could include one or more of the following: a right-holder ID, an edge cache / server ID, timestamp, event / content ID, recipient ID (address or other ID), or other information or an internal code that to uniquely identity or otherwise characterise the stream. An example entails reparsing just one or two frames every N frames, where N is a constant value, for example, 20, or a variable value calculated in a known fashion, and adding a constant +x and then -x of a stream content parameter to the whole display. In the case of LCEVC encoded video, in some examples, this can be done very easily by adding a constant symbol in LOQ1, so it would cost very little data, and would have no repercussions on subsequent frames (i.e., there is no requirement to clear the temporal buffer, since LOQ1 does not depend on this). Such an example operation would require very minimal computational load, essentially akin to a simple reparsing of just the LCEVC enhancement data, for just those specific frames, if the encoded information is encoded in a small number of frames over a reasonably long time period, such as 10 minutes, the processing load can be spread in such a way to produce irrelevant processing load, and essentially zero quality impact. For example, transmitting 2 bits of single-frame content parameter changes (such as luminance) every 10 seconds would still allow to send four8-bit numbers every 160 seconds, which would allow for a large quantity of data to be encoded. It should be understood that a greater number of bits may be encoded dependent on the specific encoding scheme chosen. With the example just given, to encode an IP address, such as an IPv4 address, every 10 minutes, only a minimal luminance change is required every 40 seconds. In a different example, to encode an identifier every 30 minutes (which still allows for sufficient redundancy in order to identify a third party source), one minimal luminance change every two minutes is required, making it completely effortless from a processing standpoint and imperceptible from a user experience standpoint. If this is done with a more sophisticated approach rather than just a full-screen luminance change, it is possible to transmit more than two bits per “intervention”, thus requiring even fewer individual modifications per stream. From an implementation point of view, a CDN edge server can store pre-prepared entropy encoded streams of the LOQ1 data to add to specific frames in order to embed an identifier, so that the processing operation would just require a simple reparsing of the enhancement data for that frame. This is something that caching servers are well equipped to achieve (which is different from video transcoding, which has a high processing load). Turning to the detector 2000, this is configured to receive the encoded video signal. This is with the intention of identifying an output destination of an encoded video signa! represented in an independent video stream. As well as the encoded video signal, the detector receives, from the apparatus 1000, a frame range in which the deviation is included in the encoded video signa! output by the apparatus; and a representation of the encoded video signal in the independent video stream that is to be tested. Once all this data has been received, the representation of the encoded video signal and the encoded video signal are compared in the received frame range to identify differences. From this a deviation implemented in the encoded video signal for which the representation has been received is derived, i.e. the deviation is identified. The deviation being derived then allows the output destination to be identified, either by comparing the deviation against logs of deviations implemented or because the deviation itself provides or is the relevant identifier. An example mechanism by which the detector may operate is outlined in more detail below in relation to Figure 5. Figure 3 shows an example schematic of a method according to the invention. In step S302, a user device sends a content request to an edge server. The edge server is part of a larger content delivery network, and is responsible for transmitting an encoded video signal (or stream) corresponding to the requested content to the user device. In step S304, the edge server makes a request to a data store for an encoded video signal corresponding to the requested content. In other examples, the video signal may be stored on the edge server itself, or may be provided by other means, for instance, by making a network request on the wider Internet. In step S306, the data store provides to the edge server the requested encoded video signal corresponding to the requested content. In step S308, the edge server requests user-specific deviation data from the data store. In this example, the data store Is acting as a library of deviation data, providing the edge server with an appropriate enhancement layer which encodes an identifier characteristic of the user who made the request. In other examples, the edge server may generate the appropriate enhancement layer itself, or it may in and of itself comprise a library of deviation data and not need to make a request to the data store. In step S310, the data store provides to the edge server the requested userspecific deviation data. In step S312, the edge server generates an adjusted enhancement stream from the user-specific deviation data. In this example, the user-specific deviation data is a portion of an enhancement stream with user-specific deviations in it. The edge server creates from this a full enhancement stream suitable for combining with the original base layer of the original encoded video signal corresponding to the requested content, by appropriate repetition. In step S314, the edge server generates an encoded video signal by combining the adjusted enhancement stream generated in step S312 with the base encoded stream of the original encoded video signal. In step S316, the edge server sends the resultant encoded video signal to the user device. In step S318, the user device decodes the encoded video signal, and plays the resultant video. The video played exhibits the deviations provided by the adjusted enhancement stream, and is therefore watermarked. If the resultant video were to be captured and shared, the deviations could be decoded to retrieve the encoded data characteristic of the user who received it originally. This is expounded further below, in the discussion of Figures 5A to 5C. Figure 4 shows an example schematic of a method according to the invention. In step S400, a user device sends a content request to an edge server. The discussion of step S302 above applies equally. In step S402, the edge server requests a user-specific adjusted video signal from a compute server. This differs from the example provided by Figure 3, where a compute server is not used. Indeed, steps of the method of Figure 4 performed by the compute server could be, in a different example, performed by the edge server itself. However, as some steps of the method of Figure 4 are more computationally demanding, it may be advantageous to offload these to a server specifically suited, eg. one with greater CPU or GPU resources. In step S404, the compute server requests encoded video signal corresponding to the requested content from the data store. This step is analogous to S304 and the discussion above applies equally. In step S406, the data store provides the requested video signal. This step is analogous to S306 and the discussion above applies equally. In step S408, the compute server requests user-specific deviation data from the data store. This step is analogous to S308 and the discussion above applies equally. In step S410, the data store provides the user-specific deviation data to the compute server. This step is analogous to S310 and the discussion above applies equally. In step S412, the compute server decodes a layer of the existing enhancement stream in the requested video signal corresponding to the requested content. In this exampie, the existing enhancement layer is decoded by the compute server, so that it can be combined with the required deviation. In other examples, the encoding of the enhancement layer may be such that it can be combined with the deviation without decoding. In step S414, the compute server applies the user-specific deviation to the decoded enhancement layer. In this example, this corresponds to increases and decreases of the luminance for particular frames, on top of the effect already applied by the existing enhancement layer, to produce an adjusted enhancement stream. In step S416, the compute server re-encodes the adjusted enhancement stream such that it can be combined with the original base encoded stream as a video signal. In step S418, the compute server combines the adjusted enhancement stream with the original base encoded stream to produce an adjusted encoded video signal that, when decoded and played, will exhibit both the effects of the original enhancement stream, / .e. increased fidelity of playback, and the deviations, i.e. encoded information in the luminance of certain frames. In step S420, the compute server provides the encoded video signal to the edge server. In step S422, the edge server sends the encoded video signal to the user device. In step S424, the user device decodes the encoded video signal, and plays the resultant video. The video played exhibits both the enhancements of the original enhancement stream and the deviations provided by the adjusted enhancement stream, and is therefore improved and watermarked. If the resultant video were to be captured and shared, the deviations could be decoded to retrieve the encoded data characteristic of the user who received it originally. This is expounded further below, in the discussion of Figures 5Ato 5C. Figure 5A depicts a graph of an example average luminance of frames of an original video. The luminance, in this example, is given as an arbitrary, normalized dimensionless quantity for ease of explanation. It can be seen from the graph that the luminance varies from frame to frame, following a general trend over many frames. In this example, the luminance shown is an average of the luminance of each pixel in the whole frame of video. However, in other examples, a one or more regions of the frame may be averaged instead, or a single pixel may be measured. Other examples for which this would be applicable could be regions, blocks, tiles or some other portion smaller than the full frame. Additionally, in other examples, a parameter other than the luminance may be extracted, such as chrominance, audio level, frequency, etc. The original video is produced by decoding an original encoded video signal, which is encoded using the LCEVC codec. As the LCEVC codec comprises an enhancement codec, the original encoded video signa! comprises a base encoded stream and an enhancement stream. In Figure 5B, an example average luminance of frames of an adjusted video, produced by decoding an adjusted encoded video signal comprising an adjusted enhancement stream and an original base encoded stream, is shown. It should be understood that the difference between the luminance shown in Figure 5Aand 5B is exaggerated for ease of explanation. The adjusted enhancement stream has the effect of producing a deviation in the luminance of certain frames in the decoded adjusted video. In this example, the luminance of a frame is adjusted every twenty frames. The luminance of each of these adjusted frames is either increased by 0.6, or decreased by 0.6. In this way, a binary “T” or “0” is encoded in the luminance of the frame, respectively. It should be understood that, in other examples, other encoding schemes may be used, including differing variation of the luminance, differing numbers of frames modified, differing spacing between modified frames, and differing parameters of the video being modified. In this example, the luminance of the frames has been adjusted in order to encode the binary “1000101000101010’, which can be interpreted as to the 16-bit integer 35370. This is the account ID of the user who receives the adjusted encoded video signal. If this user were to capture the received video in some way and share it, including byway of transcoding, the video would still be watermarked with this account ID. This makes it possible for the rights holder, when provided with the shared content, detect the deviation of the luminance, determine the corresponding binary signal, and therefore find the corresponding account ID. In Figure 5C, a graph of an example detected deviation per frame of the adjusted video of Figure 5B is shown, in this example, the detection process is a simple subtraction of the known average luminance per frame of the original video as shown in Figure 5A from the calculated average luminance per frame of the adjusted video of Figure 5B. This results in a series of “spikes” in the graph, / .e. non-zero values wherever the luminance of a frame in the adjusted video is deviant from the original video. In this example, the series of positive and negative deviations in the luminance can be interpreted as set out above, as binary “1” and “0”. In this example, therefore, the binary “1000101000101010” is recovered. Interpreted as a 16-bit integer, this gives 35370 — the account ID of the user. Figure 6 depicts a schematic example of a luminance enhancement layer for an example frame of an encoded video. This is depicted in grey in the example shown in Figure 6.. Such an enhancement layer could be adjusted to be deviant in any of the methods described herein. When the luminance is changed by, for example, positive two bits, the layer will be a quantity of shades of grey lighter, such as two shades lighter; and when the luminance is changed by, for example, negative two bits, the layer will be a quantity of shades of grey darker, such as two shades darker. As indicated above, concepts set out herein may be implemented at a client device, player on the device or decoder. Similarly, concepts may be embodied by modifications to an encoder, packager and / or content delivery network. At each of these entities, methods and processes described herein can be embodied as code (e.g., software code) and / or data. The functionality may be implemented in hardware or software as is well-known in the art of data compression and video streaming. For example, hardware acceleration using a specifically programmed Graphical Processing Unit (GPU) ora specifically designed Field Programmable Gate Array (FPGA) may provide certain efficiencies. For completeness, such code and data can be stored on one or more computer-readable media, which may include any device or medium that can store code and / or data for use by a computer system. When a computer system reads and executes the code and / or data stored on a computer-readabie medium, the computer system performs the methods and processes embodied as data structures and code stored within the computer-readable storage medium. In certain embodiments, one or more of the steps of the methods and processes described herein can be performed by a processor (eg., a processor of a computer system or data storage system). Generally, any of the functionality described in this text or illustrated in the figures can be implemented using software, firmware (e.g., fixed logic circuitry), programmable or nonprogrammable hardware, or a combination of these implementations. The terms “component” or “function” as used herein generally represents software, firmware, hardware or a combination of these. For instance, in the case of a software implementation, the terms “component” or “function” may refer to program code that performs specified tasks when executed on a processing device or devices. The illustrated separation of components and functions into distinct units may reflect any actual or conceptual physical grouping and allocation of such software and / or hardware and tasks. The above examples are to be understood as illustrative examples. Further examples are envisaged. It is to be understood that any feature described in relation to any one example may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the examples, or any combination of any other of the examples. Furthermore, equivalents and modifications not described above may also be employed within the scope of the accompanying claims.
Claims
1. An apparatus suitable for transferring a video signal to be output as an encoded video signal encoded including a base encoded stream and at least one enhancement stream, the enhancement stream being suitable for combining with the base encoded stream to reconstruct the input video signal, and each enhancement stream comprising respective frames, the apparatus being configured to:generate an adjusted enhancement stream by providing, in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream, a deviation, the deviation being identifiable in a video signal constructed using a base encoded stream and the adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of the output destination; andoutput the encoded video signal with the adjusted enhancement stream and the base encoded stream.
2. The apparatus according to claim 1, wherein the at least one frame is a plurality of frames.
3. The apparatus according to claim 2, wherein the plurality of frames is a part of a sequence of frames, the sequence of frames having at least one intervening frame between each frame of the plurality of frames.
4. The apparatus according to any one of the preceding claims, wherein the deviation is a variation of a luminance of a whole frame or one or more portions of a frame.
5. The apparatus according to any one of the preceding claims, wherein the deviation is provided in a value of at least two portions of the at least one frame.
6. The apparatus according to any one of the preceding claims, wherein the apparatus includes a library of deviation data, the apparatus being configured to generate an adjusted enhancement includes identifying, based on the intendedoutput destination of an encoded video signal, relevant deviation data in the library and writing the relevant deviation data from the libraiy into the adjusted enhancement stream.
7. The apparatus according to any one of the preceding claims, wherein the encoded video signal has an empty enhancement stream prior to outputting the encoded video signal, the apparatus being configured to output the encoded video signal by writing the adjusted enhancement stream into the empty enhancement stream.
8. The apparatus according to any one of the ciaims 1 to 6, wherein the apparatus is configured to output the encoded video signal by overwriting any existing enhancement stream in the encoded video signal9. The apparatus according to any one of the claims 1 to 6, wherein the apparatus is configured to generate the adjustment enhancement stream and output the encoded video signal by decoding at least one layer of an existing enhancement stream of the encoded video signal, applying the deviation to the decoded at least one layer and re-applying the decoded encoding.
10. The apparatus according to any one of the preceding claims, further configured to receive the encoded video signal.
11. The apparatus according to claim 10, further configured to receive a content request from an output destination and retrieve the encoded video signal in response to the content request.
12. The apparatus according to any one of the preceding claims, wherein the adjustment enhancement stream is an LCEVC Level 1 stream.
13. The apparatus according to any one of the preceding claims, further configured to encrypt or apply digital rights management to the output encoded video signal.
14. The apparatus according to any one of the preceding claims, wherein the deviation is readable as a plurality of bits.
15. The apparatus according to claim 15, wherein the plurality of bits are a representation of a combination of one or more of: a user ID, account ID, time of output of the encoded video signal, apparatus ID, IP address, and region of a user.
16. The apparatus according to any one of the preceding claims, further configured, during the generation of the adjustment enhancement stream, to apply a frame number offset to the at least one frame to which the deviation is provided, the frame number offset being varied based on output destination.
17. The apparatus according to any one of the preceding claims, wherein the apparatus is an edge device, the output being arranged to be output out of a network of which the edge device is a part.
18. The apparatus according to any one of the preceding claims, further configured, during the generation of the adjustment enhancement stream, to repeat the deviation at one or more predetermined intervals over the length of the adjustment enhancement stream.
19. A method of transferring a video signal to be output as an encoded video signal encoded including a base encoded stream and at ieast one enhancement stream, the enhancement stream being suitable for combining with the base encoded stream to reconstruct the input video signal, and each enhancement stream comprising respective frames, the method comprising:generating an adjusted enhancement stream by providing, in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream, a deviation, the deviation being identifiable in a video signal constructed using a base encoded stream and the adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of the output destination; andoutputting the encoded video signal with the adjusted enhancement stream and the base encoded stream.
20. A detector suitable for identifying an output destination of an encoded video signal represented in an independent video stream, the detector being arranged to:receive the encoded video signal;receive a frame range over which a deviation in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream is provided, the deviation being identifiable in a video signal constructed using a base encoded stream and an adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of the output destination;receive the representation of the encoded video signal in the independent video stream;compare, in the received frame range, the received representation of the encoded video signal and the received encoded video signal to identify differences;derive the deviation provided to the encoded video signal represented in the independent video stream based on the identified differences; andidentify the output destination of the encoded video signal of the representation of the encoded video signal based on the derived deviation.
21. A method of detecting an identity of an output destination of an encoded video signal represented in an independent video stream, the method comprising: receiving the encoded video signal;receiving a frame range over which a deviation in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream is provided, the deviation being identifiable in a video signal constructed using a base encoded stream and an adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of the output destination;receiving the representation of the encoded video signal in the independent video stream;comparing the received representation of the encoded video signal and the received encoded video signal to identify differences;deriving the deviation provided to the encoded video signal represented in the independent video stream based on the identified differences; andidentifying the output destination of the encoded video signal of the representation of the encoded video signal based on the derived deviation.
22. A system suitable for transferring a video signai to be output as an encoded video signal and detecting an identity of an output destination of an encoded video signal represented in an independent video stream, the system comprising the apparatus of any one of claims 1 to 18 and the detector of claim 20, the detector being arranged to receive the frame range from the apparatus.
23. Adata stream comprising:a base encoded stream; andan adjusted enhancement stream, the adjusted enhancement stream including a deviation in a value indicative of at least one constant of at least a portion of at least one frame of an enhancement stream the deviation being identifiable in a video signal constructed using a base encoded stream and the adjusted enhancement stream, the deviation being determined by an intended output destination of an encoded video signal including the adjusted enhancement stream, the deviation thereby being characteristic of the output destination, whereinthe base encoded stream and adjusted enhancement stream are combinable to generate a constructed signal.
24. A computer program comprising instructions which, when executed, cause an apparatus to perform the method according to claim 19 or claim 21 or to provide the apparatus of any one of claims 1 to 18 or the detector of claim 20 or the system of claim 22.
25. A non-transitory computer-readable medium comprising the computer program according to claim 24.
Citation Information
Patent Citations
Provision of marked data content to user devices of a communications network
WO2010025779A1