Systems and methods for adaptive video watermark application - Patents.com
By modulating pixel values in video frames to reduce perceptibility, data can be embedded in modern digital television signals, addressing the lack of a mechanism for watermarking and enabling dynamic content replacement and metadata provision in media devices.
Patent Information
- Application Number
- JP2023506336
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-07-29
- Filing Date
- 2021-07-30
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2041-07-30
AI Technical Summary
Modern digital television standards lack a mechanism for embedding data in video signals without visually impairing the content, as they do not utilize vertical blanking intervals, preventing the application of traditional digital watermarks.
A method and system for embedding data in video frames by modulating pixel values in predetermined regions, utilizing techniques such as adjusting color difference and luminance components, error correction, and inverse watermarks to reduce perceptibility, allowing media devices to detect and decode the watermark.
Enables the embedding of additional information in video signals that is imperceptible to the human eye, facilitating dynamic content replacement and providing metadata or triggers for media devices, while maintaining video quality.
Smart Images

Figure 0007798859000004 
Figure 0007798859000005 
Figure 0007798859000006
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 059,766, filed July 31, 2020, and also claims the benefit of U.S. Non-Provisional Application No. 17 / 389,147, filed July 29, 2021, both of which are incorporated by reference in their entirety for all purposes.
[0002]
[0002] This disclosure relates generally to embedding data in a video signal, and more particularly to embedding data in a video signal without visually impairing the video signal. [Background technology]
[0003]
[0003] Digital watermarking describes a technique for hiding certain data, such as identifying data about the origin of a digital media stream. Digital watermarks can be embedded in image files or video frames in a way that inhibits their removal and does not damage the underlying content. When such watermarked digital content is distributed online or on recorded media, data reflecting its origin travels with it, allowing the originator to claim the source of the content.
[0004] In cathode ray tube (CRT) televisions, watermarks can be embedded in the vertical blanking interval between frames. In CRT televisions, the displayed image is transmitted in rows of lines, which, when projected onto the phosphors lining the CRT, appear black-and-white (and later color). These lines are repeated as interlaced frames, with frames separated by several dozen lines that are not displayed. These non-displayed lines are called "vertical blanking intervals" (VBIs). The VBIs are used to allow the CRT to move its beam from bottom to top and stabilize before beginning to scan another video frame. Information can be embedded in the portion of the video signal that corresponds to the VBI. For example, data for subtitling for the hearing impaired or train timetables for videotext displays (particularly popular in Europe and Asia) is embedded in the VBI portion of the video signal.
[0005]
[0005] Digital television does not operate on cathode ray tubes and therefore does not require a VBI between video frames to operate. Modern digital television standards no longer implement data on the VBI, eliminating the VBI as a mechanism for embedding data. Modern digital television standards instead use separate data streams in which audio and video data are interwoven. This allows the video signal to contain additional data, but that data is inaccessible to media devices, preventing the application of watermarks.
[0006]
[0006] Therefore, there is a need for an alternative technique for inserting additional information into a video signal. Summary of the Invention
[0007]
[0007] A method for embedding data in video frames with reduced likelihood of being perceived by a user is described herein. The method includes: receiving a video frame; detecting a watermark in a first predetermined region of the video frame, where a first predetermined region of the video frame includes a first set of pixels, a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels in a third predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to the first pixel values and one or more contiguous subsets of pixels corresponding to the second pixel values; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to the first pixel values and a second symbol to the one or more contiguous subsets of pixels corresponding to the second pixel values; and generating a first sequence of symbols based on the first pixel values and the second pixel values.
[0008]
[0008] Described herein is a system for embedding data in video frames with reduced likelihood of being perceived by a user. The system includes one or more processors and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods as previously described.
[0009]
[0009] Described herein is a non-transitory computer-readable medium that stores instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods described above.
[0010]
[0010] These illustrative examples are provided to aid in understanding the present disclosure, not to limit or define it. Additional embodiments are described in the detailed description, and further description is provided therein.
[0011]
[0011] Exemplary embodiments of the present application, including its systems and methods, are summarized below in the following drawings: [Brief explanation of the drawings]
[0012] [Figure 1]
[0012] FIG. 1 illustrates an example of data embedded in the top row of a video frame using a two-level watermark, according to aspects of the present disclosure. [Figure 2]
[0013] 1A-1C illustrate exemplary binary watermarks embedded in the top row and top two rows of a video frame according to aspects of the present disclosure. [Figure 3]
[0014] 10A-10C illustrate example watermarks for enhanced lead-in data symbol sequences for enhanced detection of watermarked frames, in accordance with aspects of the present disclosure. [Figure 4]
[0015] 10A-10C illustrate example data symbols of a watermark in which the Euclidean distance between the 0 and 1 symbols is temporarily increased, according to aspects of the present disclosure. [Figure 5]
[0016] 1A-1C illustrate exemplary luminance / chrominance subsampling formats of 4:4:4, 4:2:2, 4:2:0, and 4:1:1, according to an embodiment of the present disclosure. [Figure 6]
[0017] FIG. 10 illustrates a graph showing RGB color space within Y′CbCr color space, according to an embodiment of the present disclosure. [Figure 7]
[0018] 10A-10C illustrate the effect of HSL mathematical expressions to convert a hexagon into a circle, according to aspects of the present disclosure. [Figure 8]
[0019] FIG. 1 illustrates the relationship between HSL color space and RGB color space, according to an embodiment of the present disclosure. [Figure 9]
[0020] FIG. 1 illustrates a color space representation of HSL in a biconic representation that reflects the available range of saturation relative to lightness, according to aspects of the present disclosure. [Figure 10A]
[0021] 10A-10C illustrate graphs of an example set of 3D watermark symbols and their corresponding slice points, according to aspects of the present disclosure. [Figure 10B]
[0022] FIG. 10 illustrates a graph of another example set of 3D watermark symbols and their corresponding slice points, according to aspects of the present disclosure. [Figure 11]
[0023] FIG. 10 illustrates a graph of symbol slice point determination according to an aspect of the present disclosure. [Figure 12]
[0024] FIG. 10 illustrates an example video frame in which a transition region proximate to a watermark is progressively darkened to increase the perceptual blurring of the watermark to the human eye, according to aspects of the present disclosure. [Figure 13A]
[0025] 1A-1C illustrate an example sequence of video frames in which values for the 0 and 1 symbols of a watermark alternate to cause perceptual blurring of the watermark to the human eye, according to aspects of the present disclosure. [Figure 13B]
[0026] 10A-10C illustrate example video inversion of data symbols every two frames to improve perceptual blending of data symbols, according to aspects of the present disclosure. [Figure 14]
[0027] 1 is a block diagram of an encoding flow diagram for applying a watermark to a video frame according to an aspect of the present disclosure. [Figure 15]
[0028] 1 is a block diagram of a decoding flow diagram for extracting watermarked data from video frames according to aspects of the present disclosure. [Figure 16]
[0029] 1 is a flowchart of an example process for decoding a code from a watermark embedded in a video frame, according to an aspect of the present disclosure. [Figure 17]
[0030] FIG. 1 illustrates an example computing device architecture of an example computing device capable of implementing various techniques described herein, in accordance with aspects of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013]
[0031] The present disclosure includes systems and methods for generating, embedding, and / or decoding digital data codes (e.g., watermarks) that result in digital images that are imperceptible to the human eye. The watermarks may be embedded in images or video frames to provide data and / or executable code to media devices such as smart televisions, set-top boxes, mobile devices, laptop computers, tablet computers, and desktop computers. The watermarks may be embedded in video frames by modulating pixel values in one or two rows above the pixels. For example, white pixels may correspond to a first symbol in a binary code and black pixels may correspond to a second symbol in the binary code. The media device may detect and extract symbols from the watermark. A processing component of the media device may process the symbols to decode the data and / or executable code in the watermark.
[0014]
[0032] The decoded symbols may provide additional information related to the displayed video and / or cause the media device to perform certain functions. The additional information may include, but is not limited to, information related to the content of the displayed video (e.g., actors, characters, setting, team, production staff, production features, or other aspects or characteristics of the content), metadata related to the displayed video (e.g., resolution, pixel values, broadcast origin, etc.), communications related to the displayed video, etc. The decoded symbols of the watermark may include a trigger signal at the beginning of the video or at a portion thereof (e.g., a commercial) to enable the media device to detect the trigger signal and substitute a video segment (e.g., an advertisement, commercial, other video segment, etc.) stored locally in the set-top or smart TV memory or display video information from a remote server. The decoded symbols may correspond to instructions that may be executed by a processing component of the media device to perform an operation such as replacing a video frame, or a portion thereof, with a locally stored video frame, a video frame decoded in the watermark, or a video frame streamed from a remote server.
[0015]
[0033] In some cases, modulating pixel values of a portion of a video frame may cause the watermark to be perceptible to the human eye. For example, if the modulation of pixel values creates pixels that have a strong contrast relative to pixels in the non-watermarked portion of the video frame, the watermark may be perceptible to the human eye. The watermark may be perceptible even when the watermark occupies a small portion of the video frame and / or is located on the edge of the video frame. A perceptible watermark may appear as a visual artifact that may cause a user to believe that there is an error in the media device or the video, or that it may distract the user from the non-watermarked portion of the video frame.
[0016]
[0034] A watermark may be modified using one or more processes to reduce the perceptibility of the watermark in a video frame. Note that each of the following watermark modification processes may be provided alone or in combination with other watermark modification processes. A first watermark modification process may take advantage of the insensitivity of human visual perception to small changes in pixel color. This modification process may adjust one or more color difference components (e.g., a chrominance-blue (Cb) difference signal and / or a chrominance-red (Cr) difference signal) of the watermark's pixels. The color difference components may be adjusted based on the color difference components of pixels adjacent to the watermark (e.g., one or more rows of pixels adjacent to the watermark). An average hue of nearby portions of the video frame may be determined. The watermark may be defined by shifting the average hue a first predetermined amount to represent a first symbol and a second predetermined amount to represent a second symbol. Because the watermark pixels have a hue that differs from nearby portions of the video frame by only the first and second predetermined amounts, the human eye is unlikely to notice the watermark.
[0017]
[0035] The second watermark modulation process may include adjusting the luminance component (Y) pixels of the watermark. The luminance component may be represented between 0 (for black) and 100 (for white). The watermark may be defined by modulating the luminance components of pixels that will represent a first symbol and pixels that will represent a second symbol of the watermark. To reduce the likelihood that the watermark is perceptible, the luminance variation between the first and second symbols may be less than 100. In one illustrative example, the first symbol may be represented by a pixel having a luminance of 10, and the second symbol may be represented by a pixel having a luminance of 50. Those skilled in the art will appreciate that other luminance values may be used. In some cases, the second watermark modulation process may be combined with the first watermark modulation process by varying both the chrominance and luminance components. For example, combining the first and second modulation processes may be implemented to increase the amount of symbols that can be embedded in a video frame (e.g., from a binary code to a base-four code) and / or to further reduce the likelihood that a user will perceive the watermark.
[0018]
[0036] The smaller the difference between the first and second symbols (e.g., 10 to 50 in the previous example), the greater the chance that noise in the video signal can cause errors (e.g., preventing detection of the watermark or corrupting the decoded watermark). To prevent noise from rendering the watermark unreadable, an error correction watermark may be inserted into the first video frame or the nth video frame that will contain the watermark. The error correction watermark contains the same coding as the previous watermark, but has a larger variation in luminance. Returning to the previous example of luminance that minimally varies between 10 and 50, the error correction watermark may contain a variation in luminance between 10 and 80. The error correction watermark may be inserted any number of times into any video frame of the set of video frames that will contain the watermark.
[0019]
[0037] The third watermark modulation process includes defining a first watermark and a second watermark, where the second watermark is the inverse of the first watermark. The first watermark may include a first pixel value corresponding to a first symbol and a second pixel value corresponding to a second symbol. The second watermark has the same sequence of symbols, but the first pixel value corresponds to the second symbol and the second pixel value corresponds to the first symbol. For example, the first watermark may include a first set of pixels corresponding to the first pixel value, a second set of pixels corresponding to the second pixel value, and a third set of pixels corresponding to the first pixel value. The symbol sequence of the first watermark may be first symbol, second symbol, first symbol (e.g., 010 in binary conversion). The second watermark may include a first set of pixels corresponding to the second pixel value, a second set of pixels corresponding to the first pixel value, and a third set of pixels corresponding to the second pixel value. Because the second watermark is intended to be the inverse of the first watermark, the media device decodes the second watermark the same as the first watermark, or into the first symbol, the second symbol, and the first symbol. The second watermark may be configured to be included in a video frame that immediately follows the video frame containing the first watermark. The media device may be configured to expect to store the watermarks of two adjacent video frames at a time.
[0020]
[0038] Inverting the watermark contained in the first video frame in subsequent frames causes the watermark to be perceived as a solid color (e.g., the average of pixel values representing the first symbol and pixel values representing the second symbol). This reduces the appearance of pixels near other pixels with contrasting chrominance and / or luminance (e.g., black pixels next to white pixels). As a result, the watermark may be less perceptible to a user of the media device.
[0021]
[0039] The fourth watermark modulation process involves modifying a set of pixels adjacent to the watermark according to pixel values of the watermark. If the watermark is placed in the top one or two rows of the video frame, the set of pixels adjacent to the watermark may be the next row or rows after the video frame (e.g., the second, third, etc. rows from the top of the video frame). The set of pixels may be referred to as the boundary between the watermark and the remaining pixels of the video frame. In some cases, each row of the boundary may be of the same pixel value. In other cases, the rows may be a gradient from the row closest to the watermark, which has pixel values based on the first and second pixel values of the watermark, to the row closest to the remainder of the video frame, which has pixel values based on adjacent pixels of the video frame. The gradient may mitigate the difference in pixel values between the video frame and the watermark to reduce the perceptibility of the watermark. Note that the pixels in the boundary row may be of the same pixel value, similar pixel values (e.g., allowing for slight variations to reduce perceptibility), or each pixel value may have a value based on the position of the pixel along the gradient (e.g., the boundary row, if more than one) and the pixel values of nearby pixels in the video frame.
[0022]
[0040] In one example, use of the systems and methods described herein includes embedding a code in a video to be used as a signal to a media device, such as a smart TV or set-top box. The code may cause the client television receiver to use a different video segment (e.g., one or more video frames) in place of the currently displayed video segment when the data code is received by the client television device. This process may be referred to as “dynamic insertion,” and when advertisements are involved, may be referred to as “dynamic ad insertion.” The dynamic replacement of a video segment with another video segment may occur in real time. For example, whenever one or more frames eligible for replacement are detected (e.g., based on detecting a watermark in one or more frames), the client device receiver may replace one or more frames with one or more other frames.
[0023]
[0041] In another exemplary use of the systems and methods described herein, a watermark may be used to trigger an on-screen pop-up window that provides additional information related to the displayed video frame. For example, if the displayed video frame includes a product, the pop-up window may include information related to the product. If the displayed video frame corresponds to a movie or television program, the pop-up window may include information related to actors in the video frame or in the movie or television program, television program production staff such as directors, information related to the movie or program involved, and / or any information related to the movie or television program.
[0024]
[0042] The pop-up window may include a uniform resource locator (URL) link to a website that contains additional information and / or provides the user with the ability to purchase the displayed product. The media device may include a web browser that can access the website and facilitate the purchase. Alternatively or additionally, the pop-up window may include a quick response (QR) code that the user can capture with the mobile device. Capturing the QR code can cause the mobile device to open a web browser and load the URL link in the QR code.
[0025]
[0043] 1 illustrates an example of data embedded in a top row of a video frame using a two-level watermark according to an embodiment of the present disclosure. Frame 104 represents a video frame that may be presented via a media device. A watermark 108 may be inserted by modulating pixel values of a set of pixels of frame 104. The set of pixels may be located at an edge of frame 104, such as one or more rows above the pixels, one or more rows below the pixels, two columns to the right of the pixels, or two columns to the left of the pixels (as shown). In some instances, watermark 108 may be split between two different locations in frame 108. For example, watermark 108 may be located in both the top row and the bottom row, the top row and the left column, etc.
[0026]
[0044] As shown, the watermark 108 may include a pixel with a first pixel value representing a first symbol (e.g., 0) of the binary code and a pixel with a second pixel value representing a second symbol (e.g., 1) of the binary code. The watermark 108 may include additional pixel values representing additional symbols of a non-binary code. The watermark 108 is represented by a discrete set of pixels representing the symbols of the binary code. If the source is lossless (e.g., signal data is not distorted or lost due to noise or other signal impedance), a single pixel may represent a single symbol. If the video source is lossy (e.g., broadcast television, cable television, etc., where portions of a frame may be distorted due to noise, distance, etc.), a set of pixels may represent a single symbol. As shown, eight pixels may be used to represent each symbol (e.g., two rows of four pixels). In some instances, for a particular video frame, such as the first video frame containing the watermark 108, each set of pixels may include additional pixels (e.g., two rows of eight pixels, etc.) to ensure that the media device can detect the watermark 108.
[0027]
[0045] The enlarged portion 112 of the watermark 108 shows the symbol represented in each set of pixels. In the illustrated example, sets of pixels having higher luminance (e.g., closer to white) are assigned a value of 1, and sets of pixels having lower luminance (e.g., closer to black) are assigned a value of 0. Luminance may vary between 0 (e.g., black) and 100 (e.g., white). In some instances, to reduce the perceptibility of the watermark, the difference in luminance between pixels representing 0 and pixels representing 1 may be minimized. For example, a pixel representing 1 may have a luminance of 50, and a pixel representing 0 may have a luminance of 10. The color components of a set of pixels of the watermark 108 may be selected based on the color components of nearby pixels (e.g., adjacent portions of the frame 104). Color components may be used for larger-base codes (e.g., codes using more than two symbols) and / or to further reduce the perceptibility of the watermark 108 by users of the media device.
[0028]
[0046] 2 illustrates an exemplary binary watermark embedded in the top row and the top two rows of a video frame according to aspects of the present disclosure. Watermark 201 illustrates a watermark comprising two rows of pixels representing a sequence of symbols. Each set of 16 pixels (two rows of 8 pixels) represents a symbol of the watermark, with darker pixels (e.g., having a luminance value Y′ of 16) representing a symbol of 0 and lighter pixels (e.g., having a luminance value Y′ of 50) representing a symbol of 1. The luminance values and / or chrominance components of the pixels may be selected based on non-watermarked portions of the frame to reduce the perceptibility of the watermark. Thus, black and white (e.g., luminance values Y′ of 0 and 100, respectively) may not be selected.
[0029]
[0047] The number of rows and / or pixels representing a single symbol of the watermark may be selected based on the signal quality of the video. For example, a high signal quality (e.g., little noise and / or loss, etc.) may use a single row. Watermark 210 illustrates a watermark comprising a single row of pixels representing the same sequence of symbols as watermark 201. Alternatively or additionally, a high signal quality may use fewer pixels per row (e.g., four pixels in a row, two pixels in a row, one pixel, etc.) to represent a single symbol. Similarly, a signal of poor quality may use additional pixels per row or additional rows. Using additional pixels and / or rows per symbol may reduce the amount of symbols that may be contained in a single video frame, but increase the likelihood that the watermark can be detected and correctly decoded. The media device may send an indication of the current signal quality to a remote server. The remote server may then modulate the watermark in each frame to increase the likelihood that the watermark can be detected and reduce the likelihood that watermark noise or other artifacts will affect the watermark.
[0030]
[0048] FIG. 3 illustrates an example watermark for an enhanced lead-in data symbol sequence for enhancing detection of watermarked frames according to aspects of the present disclosure. The watermark may include a set of pixels, each representing a binary code symbol. The amount of pixels in the set of pixels representing a symbol may be referred to as the pixel size of the symbol. In some instances, the pixel size may be four pixels (if the watermark is one row) or eight pixels (if the watermark is two rows). The pixel values of the set of pixels may correspond to approximately the same pixel value. Alternatively, the pixel values of the set of pixels may correspond to similar pixel values (e.g., within a range).
[0031]
[0049] A watermark may begin with a predetermined pattern of data signaling the start of the watermark (known as a lead-in pattern). The predetermined pattern may be located in the first 8 or 16 symbols of the watermark. The media device may first determine whether the predetermined pattern is detected in the first x pixels of the video frame (e.g., amount of pixels per symbol * number of symbols in the predetermined pattern of data). If the predetermined pattern is detected, the media device may continue decoding the remaining pixels in that row.
[0032]
[0050] The lead-in pattern in a watermark may be adjusted to increase the likelihood that the watermark will be detected by a media device. For example, the pixel size of each symbol may be increased. By increasing the pixel size of each symbol, the lead-in pattern may be more reliably decoded by a media device. In some examples, the pixel size of each symbol in the lead-in pattern may be doubled (e.g., as shown). The pixel size of the remainder of the symbols in the watermark may not be adjusted. For example, if the lead-in pattern is eight symbols with a pixel size of four, the lead-in pattern alone may occupy the first 64 pixels (e.g., a pixel size of eight per symbol for eight symbols), and the pixel size of the symbols after the lead-in pattern may remain four. A lead-in pattern may be particularly useful when watermarked frames occur periodically among a large number of video frames that do not contain the watermark.
[0033]
[0051] In this example, the illustrated watermark 301 includes a lead-in pattern with an increased pixel size per symbol. The symbols of the lead-in pattern are represented by double the pixel size (e.g., from two rows of four to two rows of eight). For example, symbols 302 and 303 are represented by 16 pixels instead of eight. The increased pixel size per symbol continues for the length of the lead-in pattern (e.g., 8 to 16 symbols or up to 256 pixels). At the end of the lead-in pattern and for the remainder of the watermark in that video frame, each symbol retains its unincreased pixel size (e.g., two rows of eight). The lead-in pattern may include a pixel size per symbol that is increased by any amount (e.g., without limitation, twice the normal pixel size per symbol, three times the normal pixel size per symbol, a fraction of the normal pixel size per symbol, a multiple of the normal pixel size per symbol, etc.).
[0034]
[0052] FIG. 4 illustrates an example of data symbols for a watermark in which the Euclidean distance between the 0 and 1 symbols is temporarily increased, according to an embodiment of the present disclosure. The watermark may be encoded in a video frame by modulating pixel values of a set of pixels, such as pixels in the top two rows of the video frame. Modulating the pixel values may include selecting two luminance values to represent two symbols of a binary code. The two luminance values may be selected based on the likelihood that a media device can detect each symbol in the presence of noise, etc., and the likelihood that the watermark will not be perceived by a user of the media device. If the luminance values for the first and second symbols are too close to each other, signal noise may prevent the media device from accurately determining whether the set of pixels represents the first or second symbol. The closer the luminance values for the first and second symbols are, the less likely a user will perceive the watermark. The luminance values may be selected as the smallest difference that provides a threshold probability that the media device will be able to detect and accurately decode the symbol sequence of the watermark. In some cases, the difference may be approximately 40, so that pixels representing the first symbol may have a luminance value of approximately 5-15, and pixels representing the second symbol may have a luminance value of approximately 45-55.
[0035]
[0053] Signal noise may induce errors in a decoded symbol sequence when the luminance difference between symbols is minimal (e.g., approximately 40). To reduce the likelihood of errors in a decoded symbol sequence, an error correction watermark may be inserted into one or more video frames of the set of video frames that will include the watermark. The error correction watermark may include a higher difference between the luminance values representing the first symbol and the luminance values representing the second symbol. In some cases, the difference between the high and low luminance values for the error correction watermark may be approximately 80, such that the pixel representing the first symbol may have a luminance value of approximately 10-20 and the pixel representing the second symbol may have a luminance value of approximately 75-85. When a media device receives the error correction watermark, the larger difference in luminance values between symbols increases the likelihood that the media device can detect and correctly decode the symbol sequence. The next video frame may include a normal watermark with a normal (smaller) difference in luminance values between symbols.
[0036]
[0054] An error correction watermark may be embedded in multiple frames of the set of frames. For example, an error correction watermark may be inserted every "n" frames. Alternatively or additionally, an error correction watermark may be inserted into one or more adjacent video frames each time an error correction watermark is inserted. For example, each time an error correction watermark is inserted into a video frame, the error correction watermark may also be inserted into one or more subsequent frames (e.g., for m-1 frames). That is, each time an error correction watermark is inserted, it may be inserted into "m" video frames.
[0037]
[0055] Alternatively or additionally, when the average luminance of the video frames is high (e.g., greater than a first threshold), the pixels of the watermark may be modulated such that the difference in luminance values between the pixels representing the first symbol and the pixels representing the second symbol is approximately 80 (e.g., using a luminance value of approximately 10-20 to represent the first symbol and a luminance value of approximately 70-80 to represent the second symbol, or any luminance value with a difference of approximately 80 therebetween). When the average luminance of the video frames is low (e.g., less than a second threshold), the pixels of the watermark may be modulated such that the difference in luminance values between the pixels representing the first symbol and the pixels representing the second symbol is approximately 40 (e.g., using a luminance value of approximately 10-20 to represent the first symbol and a luminance value of approximately 45-55 to represent the second symbol, or any luminance value with a difference of approximately 40 therebetween). Furthermore, when the average luminance of the video frame is low, the pixels of the watermark may have a color channel, such as Cr, adjusted between an extreme value for the 0 symbol color and an extreme value for the 1 symbol color. The first and second thresholds may be predetermined or dynamically determined based on pixel values of the video frame. In some cases, the first threshold may be equal to the second threshold. In other cases, the first threshold may be a difference from the second threshold.
[0038]
[0056] Alternative error correction processes may include embedding the same watermark in multiple adjacent video frames or multiple instances of the same video frame (each containing the same watermark). By transmitting the same watermark more than once, a media device may be able to better recover data that may be distorted by the video distribution path (e.g., from the source to the media device). If a lead-in pattern is detected but the remainder of the video is not reliably decodable, averaging the video values of subsequent video frames of the group can increase the signal-to-noise ratio to provide decodable data. In some instances, a media device may average the pixel values of each instance of the same watermark before decoding the watermark into symbols.
[0039]
[0057] A media device can identify related video frames (whether there are two or more in the group) by a unique lead-in pattern in the first watermark of a group of video frames that will contain the same watermark. The unique lead-in pattern may indicate the amount of frames contained in the video frame based on the unique lead-in pattern relating to a known amount of video frames, or based on the symbols in the unique lead-in pattern indicating how many of the following video frames will contain the same watermark. Alternatively, the first lead-in pattern may be used to indicate the start of a group of video frames and the second lead-in pattern may be used to indicate the last frame in the group of video frames. Alternative error correction processes may be combined with other processes as described herein, including error correction watermarks as previously described.
[0040]
[0058] FIG. 5 illustrates the effect of an HSL mathematical representation transforming a hexagon 501 into a circle 502, according to an embodiment of the present disclosure. Pixels may be represented in several different formats, including, but not limited to, the RGB (red, green, blue) color model, HLS (hue, saturation, lightness), and the luminance / chrominance (Y / C) system. Representing pixel information in any one format and converting it in another format may be utilized based on the media device, the nature of human perception, and the video source. For example, because human visual perception is relatively insensitive to changes in hue, hue information may be compressed to a greater extent than lightness. If some hue information is lost due to the degree of compression, it is less likely to result in human-perceptible artifacts.
[0041]
[0059] A pixel can be represented by three values in the HSL color space: hue H, saturation S, and lightness L. HSL is an alternative representation to the RGB color model. In the HSL representation, colors of each hue can be arranged in a radial slice around a central axis of neutral colors ranging from black at the bottom to white at the top. The HSL color space can model the way different colors of physical paint mix, with the saturation dimension resembling various shades of brightly pigmented paint and the lightness dimension resembling a mixture of those paints with varying amounts of black or white paint. The HSL model can resemble more perceptual color models, such as the Natural Color System (NCS) or Munsell color system, which center fully saturated colors on a circle at a lightness value of 1 / 2, where a lightness value of 0 corresponds to black and a lightness value of 1 represents white.
[0042]
[0060] Hue and chroma (e.g., attributes indicating colorfulness versus brightness) in the HLS representation may be defined with respect to the hexagon representation 501 (e.g., a projection of a three-dimensional RGB color space onto a two-dimensional plane). Chroma C is the fraction of the distance from the origin to the edge (e.g., C=range(R,G,B)=max(R,G,B)-min(R,G,B)). Hue H is the fraction of the distance around the edge of the hexagon 501 that passes through the projection point. Because hue can be undefined for a projection point therein that projects onto the origin, hue is mathematically defined piecewise (e.g., H=60°-H'). H' may have four definitions depending on the chroma and / or RGB values, thus when C=0, H' is undefined and max(R,G,B)=R.
[0043]
number
[0044] and max(R,G,B)=G,
[0045]
number
[0046] and max(R,G,B)=B,
[0047]
number
[0048] The definition of chroma and hue becomes like a geometric warping from a hexagonal representation 501 to a circular representation 502.
[0049]
[0061] 6A and 6B illustrate the relationship between the HSL color space and the RGB color space according to embodiments of the present disclosure. As shown in FIG. 6A, HSL may be represented by a cylinder 701. Hue 603 (the angular dimension in both color spaces represented by the "hue" arrow) may start at the red primary at 0°, pass through the green primary at 120°, the blue primary at 240°, and then wrap back to red at 360°. In each geometric shape, the central vertical axis may represent lightness. The central vertical axis may comprise a neutral, achromatic, or gray color ranging from black (value 0) at 0% lightness 605 at the bottom of the cylinder to white (value 1) at 100% lightness 605 at the top of the cylinder.
[0050]
[0062] Additive primary and secondary colors (red, yellow, green, cyan, blue, and magenta) and linear blends between adjacent pairs of those colors (sometimes called pure colors) may be arranged around the outer edge of the cylinder with a saturation of 1 (saturation 604 represented by the "saturation" arrow). Saturated colors may have a lightness 605 of 50% in HSL. Mixing these pure colors with black, creating so-called shades, may leave the saturation 804 unchanged. In HSL, the saturation 604 may also be unchanged by tinting with white. Blends with both black and white (called tones) may have a saturation 604 less than 100%.
[0051]
[0063] 6B shows an example diagram of a cubic representation 602 of the Red-Green-Blue (RGB) color space. The mathematical relationship between the RGB color space and the HSL color space is H=cos -1 ((0.5(RG))+(RB)) / ((((RG)2)+((RB)(GB))) 0.5 ), S=1-(3 / (R+G+B))*min(R,G,B), and L=R+G+B / 3. Because these definitions of saturation conflict with the intuitive notion of color purity, a biconic representation (also called a conic representation) is often used instead.
[0052]
[0064] FIG. 7 illustrates exemplary 4:4:4, 4:2:2, 4:2:0, and 4:1:1 luma / chrominance subsampling formats according to embodiments of the present disclosure. Luma / chrominance or Y / C systems (such as Y / C or Y'CbCr) may be variations of the HSL color space. These variations include "color difference" components or values derived from blue minus luma (V and Cr) and red minus luma (U and Cb). Because the human eye may have less spatial sensitivity to color than the luma, U, and V color signals of the Y'UV system, the primary component of the C signal may be substantially compressed through chroma subsampling. For example, in some cases, only half the horizontal resolution may be included in the video signal compared to the luma information.
[0053]
[0065] Chroma subsampling format 701 indicates a frame with a full 4:4:4 ratio, chroma subsampling format 702 indicates a 4:2:2 ratio, and chroma subsampling format 703 indicates a 4:2:0 ratio with half the vertical resolution. The 4:x:x representation can convey the ratio of luma and chrominance components. For example, chroma subsampling format 704 indicates 4:1:1 chroma with a quartered horizontal color resolution (as indicated by the empty dots), but full vertical color resolution (as indicated by the black dots). In this example, a video frame may contain a quarter color resolution compared to the luma resolution. Chroma subsampling format 701 with a 4:4:4 ratio may provide equal resolution for both luma and color information, equivalent to the RGB values of the raw video.
[0054]
[0066] The Y / C system may be a way of encoding RGB information, and the actual colors displayed may depend on the original RGB color space used to define the system. That is, the color space may be defined by how deep the red, green, and blue primaries are (e.g., called the color gamut). Values expressed as Y'UV or Y'CbCr may be directly converted to values of the original set of red, green, and blue primaries. The range of RGB colors and intensities may be much smaller than the range of colors and intensities encoded by Y'UV. This may be determined because when converting from Y'UV or Y'CbCr to RGB, the conversion may result in "invalid" RGB values. The systems and methods described herein may detect and correct invalid RGB values in video frames that contain invalid RGB values in a watermark.
[0055]
[0067] 8 shows a graph illustrating RGB color space within Y'CbCr color space, according to an embodiment of the present disclosure. The Y'CbCr color space (represented as block 802 for all possible Y'CbCr values) can be formed by balancing the RGB color space (represented as RGB color block 801) on its black point 804 with a white point 803 directly above the black point. The Y / C system can be converted directly to RGB, but RGB can be converted to Y'Cb'Cr by Y'=0.257*R'+0.504*G'+0.098*B'+16, Cb'=-0.148*R'-0.291*G'+0.439*B'+128, and Cr'=0.439*R'-0.368*G'-0.071*B'+128. Y'Cb'Cr' can be converted to RGB by R'=1.164*Y'-16)+1.596*(Cr'-128), G'=1.164*(Y'-16)-0.813*(Cr'-128)-0.392*(Cb'-128), and B'=1.164*(Y'-16)+2.017*(Cb'-128).
[0056]
[0068] In 8-bit encoding, the R, G, B, and Y channels may have a nominal range of [16...235], and the Cb and Cr channels may have a nominal range of [16...240] with 128 as the intermediate value. In RGB, reference black is (16,16,16) and reference white is (235,235,235). In Y'CBCR, reference black is (16,128,128) and reference white is (235,128,128), as shown in Figure 6. Values outside the nominal range are allowed, but they will generally be clamped for broadcast or display. Values 0 and 255 may be reserved as timing references and may not contain color data. In 10-bit encoding, the nominal values may be four times those of 8-bit encoding.
[0057]
[0069] The hue H, saturation S, and lightness L parameters can be manipulated to obscure a watermark embedded in a video frame. A watermark can be embedded in a video frame by modulating the H, S, and / or L values in a way that is detected by the media device but not by a user of the media device. For example, with reference to FIG. 6A , it can be shown that for any hue 603 at low or high lightness 605, the saturation 604 can be changed over a small area of the display screen without being noticeable to the human eye. Similarly, for low or high lightness 605 for some colors, the hue 603 can be shifted without being noticeable to the human eye. At near full lightness 605 (white) or near minimum lightness (black), the hue 603 and saturation 804 can be shifted without creating visible artifacts.
[0058]
[0070] FIG. 9 illustrates a color space representation of HSL in a bicone representation that reflects the available range of saturation relative to lightness, according to an embodiment of the present disclosure. The bicone representation 901 can be generated using the cylindrical representation 801 of FIG. 8A and generating two adjacent cones in the bicone shape 901 with white at the peak of the upper cone and black at the peak of the lower cone. The top, white 902, can be a single data point with maximum lightness, saturation 0, and an undefined hue (when saturation is 0). A change in saturation S or hue H will not change the corresponding value in RGB (which would be 255, 255, and 255, respectively, in an 8-bit system). When lightness is minimum, black 903, saturation S will be 0 and hue H will be undefined at the corresponding value in RGB (e.g., 0, 0, 0 in an 8-bit system). The "Y" luminance value may have a range 905 of approximately 16 to 235 out of 0 to 255, optionally excluding ranges 904 and 906 at the margins of the luminance scale. Cr and Cb values may have a range of 16 to 240 out of 0 to 255. Because color difference encoding consists of overlapping luminance and chroma values, not all values of Y'CbCr can convert to valid RGB values. For example, some combinations of Y'CbCr values can convert to negative RGB values. In some cases, a media device may correct invalid values using one or more error correction processes. One such error correction process may include setting values below the minimum allowable value (e.g., negative values) to the minimum allowable value (e.g., 0 in an 8-bit color space) and setting values greater than the maximum allowable value (e.g., greater than 255 in an 8-bit color space) to the maximum allowable value (e.g., 255).
[0059]
[0071] FIG. 10A shows a graph of an example set of 3D watermark symbols and their corresponding slice points according to an embodiment of the present disclosure. The watermark may be inserted into a video frame by modulating pixel values of pixels located at an edge portion of the video frame (e.g., the top one or two rows as described in FIG. 2). The modulation of pixel values may be based on the luminance and chroma of nearby pixels to limit the perceptibility of the watermark. In some cases, the luminance value of a pixel may be used to represent a symbol of a code. For example, a luminance value of 10 may be used to represent a first symbol, and a luminance value of 50 may be used to represent a second symbol. The chrominance value may be determined based on surrounding pixels (e.g., one or more rows of pixels adjacent to the watermark) and using the techniques described in FIGS. 5-9. Alternatively, the chrominance value may be used to represent the symbol of the watermark. In that case, the luminance value may be similar to the luminance value of nearby pixels.
[0060]
[0072] The watermark can be generated by modulating data in the three-dimensional space of the Y'CbCr color space 1001. The distance between a 0 symbol and a 1 symbol is the Euclidean distance in the three-dimensional color space, D=((Y1-Y0) 2 +(Cb1-Cb0) 2 +(Cr1-Cr0) 2 ) 1 / 2 A three-dimensional means of representing the data may provide greater distance between pixel values representing a first symbol and pixel values representing a second symbol. The use of a three-dimensional color space may provide a means for finding pixel values for a symbol that match pixel values of surrounding pixels to reduce the perceptibility of the pixel value selected for the symbol.
[0061]
[0073] A first pixel value 1002 may be selected to represent a first symbol 1002 based on nearby pixels, and a second pixel 1004 may be selected by shifting the luminance value of the first pixel value and / or by selecting a pixel value with a known value (e.g., black). The pixel values of the watermark may vary based on surrounding pixels, signal noise, a compression algorithm, etc. (e.g., to reduce the perceptibility of the watermark). The media device may use a symbol slice point 1003 to determine whether a pixel of the watermark corresponds to the first symbol or the second symbol. This ensures that the watermark can still be decoded while allowing the pixel values to vary (ensuring reduced perceptibility). The symbol slice point 1003 may be selected as a midpoint between the first pixel value 1002 and the second pixel value 1004. The symbol slice point 1003 represents the point at which the media device identifies the pixel value of a pixel of the watermark as corresponding to the first symbol or the second symbol. If a pixel has a value between the symbol slice point 1003 and the first pixel value 1002, the media device determines that the pixel represents a first symbol. If a pixel has a value between the symbol slice point 1003 and the second pixel value 1004, the media device determines that the pixel represents a second symbol.
[0062]
[0074] 10B shows a graph of another example set of 3D watermark symbols and their corresponding slice points according to aspects of the present disclosure. In the illustrated example, a first pixel value 1005 is selected as purple (Y'=87, Cb=186, Cr=201) and a second pixel value 1007 is selected as red (Y'=16, Cb=94, Cr=218). A symbol slice 1006 may be selected as the midpoint between the first pixel value 1005 and the second pixel value 1007. If a pixel has a value between the symbol slice point 1006 and the first pixel value 1005, the media device determines that the pixel represents the first symbol. If a pixel has a value between the symbol slice point 1006 and the second pixel value 1007, the media device determines that the pixel represents the second symbol.
[0063]
[0075] 11 shows a graph of symbol slice point determination according to an aspect of the present disclosure. The graph shows the percentage of symbols detected at various luma values. A first pixel value (e.g., representing a first symbol) may have a luma value of approximately 0. A second pixel value (e.g., representing a second symbol) may have a luma value of approximately 42. A slice point is selected as the midpoint between 0 and 42, or 21. If the pixel has a luma value less than 21, the media device determines that the pixel represents the first symbol. If the pixel has a luma value greater than 21, the media device determines that the pixel represents the second symbol. The media device may still detect symbols when the luma values do not exactly correspond to the first pixel value or the second pixel value.
[0064]
[0076] When the distance between the first pixel value and the second pixel value is greater than a threshold, the likelihood of a decoding error is reduced, for example, when very few, if any, symbols are detected near the slice point, where it may be difficult to determine whether the pixel value corresponds to the first symbol or the second symbol.
[0065]
[0077] FIG. 12 shows an example video frame in which a transition region proximate a watermark is progressively darkened to increase the perceived blur of the watermark to the human eye, according to aspects of the present disclosure. A boundary region may be created between the watermark and the remainder of the video frame. The boundary region may include one or more rows of pixels adjacent to the watermark. These rows are modified to create a visual blending effect that reduces the perceptibility of the watermark. For example, rows 1201 and 1202 may be adjusted to the same value (e.g., proportional to the pixel values representing the watermark symbols). In some cases, each row may be adjusted independently to create a gradient. For example, the luminance value Y of the pixels in row 1201 may be reduced based on the pixel values representing the watermark symbols, and the luminance value Y of the pixels in row 1201 may also be reduced, but less than that of row 1201.
[0066]
[0078] The boundary region may include any number of rows. A luminance gradient may be defined that is equal to the average luminance of the video frame divided by the number of rows in the boundary region. It is then determined whether the average luminance of the video frame is higher or lower than the average luminance of the watermark. If the average luminance of the video frame is higher than the average luminance of the watermark, the boundary region may shift from dark closest to the watermark to light closest to the video frame (e.g., from lower luminance to higher luminance). If the average luminance of the video frame is lower than the watermark, the boundary region may shift from light closest to the watermark to dark closest to the video frame (e.g., from higher luminance to lower luminance).
[0067]
[0079] For example, if the average luminance of the video frame is higher than the average luminance of the watermark, the luminance values of the pixels in the first row of the boundary region (e.g., the row adjacent to the watermark) may be reduced based on the average luminance of the video frame (e.g., a value proportional to the average luminance of the video frame, etc.). The luminance of the next row of the boundary region (the next row further from the watermark) may be reduced by the amount the previous row was reduced minus the luminance gradient. The luminance of each subsequent row further from the watermark may be reduced based on the amount the immediately previous row was reduced minus the luminance gradient.
[0068]
[0080] In another example, if the average luminance of the video frame is lower than the average luminance of the watermark, the luminance values of the pixels in the first row of the boundary region (e.g., the row adjacent to the watermark) may be increased based on the average luminance of the video frame (e.g., a value proportional to the average luminance of the video frame, etc.). The luminance of the next row of the boundary region (the next row further from the watermark) may be increased by the amount the previous row was increased by minus the luminance gradient. The luminance of each subsequent row further from the watermark may be increased based on the amount the immediately previous row was increased by minus the luminance gradient.
[0069]
[0081] Alternatively, if the average luminance of the video frame is lower than the first threshold, the boundary region may have a gradient from light closest to the watermark to dark furthest from the watermark. If the average luminance of the video frame is higher than the second threshold, the boundary region may have a gradient from dark closest to the watermark to light furthest from the watermark. The difference in luminance values for each row may be proportional to the average luminance value of the watermark or video frame. Note that the first threshold may be equal to or different from the second threshold.
[0070]
[0082] FIG. 13A shows an example sequence of video frames in which values for the 0 and 1 symbols of a watermark alternate to cause perceptual blurring of the watermark to the human eye, according to an embodiment of the present disclosure. Because the watermark involves modulation of pixel values to indicate two or more types of symbols (depending on the base code), the alternating pixel values may appear as flickering to a user of the media device. The flickering may be reduced or eliminated by presenting an inverted form of the video frame in subsequent video frames. A first version of a watermark may be embedded in frame 1. A second version of the same watermark may be embedded in frame 2. The second version of the watermark may be an inverted form of the first watermark. For example, each set of pixels having a luminance value representing a first symbol (e.g., a low luminance between 10 and 16) may be given a luminance value representing a second symbol (e.g., a high luminance between 45 and 55). Each set of pixels having a luminance value representing the second symbol may be given a luminance value representing the first symbol.
[0071]
[0083] After receiving a first version of a watermark in frame 1, a media device may expect that watermark in frame 2 to be inverted. When decoding the watermark in frame 2, the media device may invert the decoded symbols (e.g., each first symbol may be replaced with a second symbol and each second symbol may be replaced with a first symbol). The media device may receive an indication of how many frames will contain the same watermark (e.g., in alternating inverted form) to ensure that the watermark is detected and decoded correctly.
[0072]
[0084] In some cases, the next frame (e.g., frame 3) may contain an inverted watermark of the watermark contained in the previous frame (e.g., frame 2) that is equal to the original watermark (e.g., in frame 1). A watermark may be inverted one or more times using two or more video frames. By inverting pixels in alternating frames, a user may perceive the watermark as the average pixel value between the two frames. For example, if the first pixel is white and the inverted pixel is black, the two pixels appear gray when displayed in succession (as shown). This may cause the watermark to appear as a solid color rather than a flickering pixel. Increasing the number of times the watermark is inverted may reduce the likelihood that the watermark can be perceived, but may also reduce the amount of data that can be transmitted with a given set of video frames. The amount of frames over which the watermark is inverted is based on the amount of data to be embedded in the watermark and the likelihood that the watermark can be detected.
[0073]
[0085] FIG. 13B illustrates an example video inversion of data symbols every two frames to improve perceptual mixing of data symbols according to aspects of the present disclosure. In some instances, a sequence of watermarks may be inserted into a set of frames to transmit larger amounts of data to a media device. For example, frame 1 may include a first watermark, and frame 2 may include an inverted version of the first watermark. Frames 1 and 2 may be referred to as inversion pair A. The next watermark in the watermark sequence may be embedded in frame 3, and the inverse of that watermark is embedded in frame 4 (e.g., inversion pair B). Each odd video frame may include a new watermark in the watermark sequence, and each even frame may include an inverted version of the watermark embedded in the immediately previous video frame.
[0074]
[0086] When a sequence of watermarks is embedded in a set of video frames, it may be more perceptible to a user. The modulation of pixels between the video frames and the watermark may be visible as flicker. By embedding the watermark and an inverted version of the watermark in successive video frames, the flicker may be reduced or eliminated.
[0075]
[0087] 14 shows a block diagram of an encoding flow process for applying a watermark to a video frame according to an aspect of the present disclosure. The encoding flow process may be performed by a device that provides video services to a media device. An encoded video source 1401 (e.g., video data) is passed to a video decoder 1402. Decoded video from the decoder 1402 and watermark data 1403 are received by error correction coding 1404. The error correction coding 1404 includes error correction codes in the video data and / or data representing the watermark. This ensures that the media device (e.g., a television, a set-top box, etc.) can apply error correction to the data in the video frame and ensure that the data can be decoded.
[0076]
[0088] Next, data pixel value calculations 1405 are derived for the watermark data 1403. To reduce the likelihood that the watermark will be perceived by a user, pixel values may be selected based on the color and / or luminance of surrounding pixels. Data pixel value calculations 1405 derive approximate pixel values that represent the symbols 0 and 1. Once pixel value calculations 1405 are complete, the watermark may be embedded in the video frame by modulating the pixels of the video frame. In some cases, one or two rows of pixels may be modulated to form the watermark in the video frame. The pixels may be modulated between a first pixel value to represent 0 and a second pixel value to represent 1, as determined by data pixel value calculations 1405. The process continues with MPEG compression 1407, where the video frame is compressed, and then culminates in MPEG stitching 1408, where the compressed video frame is stitched with other video data from the encoded video source 1401. Because the watermark is located on a portion of the video frame, blocks 1401-1406 may operate using only a portion of the video frame (e.g., a portion larger than the size of the watermark). MPEG stitching 1408 combines that portion of video frame 1408 that contains the watermark with the rest of the video frame.
[0077]
[0089] Once the video frame is complete, the video frame is encoded for transmission 1409. The video frame may be encoded based on a transmission medium and / or communication protocol (e.g., broadcast television, streaming protocol, etc.). Once encoded, the video frame including the watermark is output at 1410.
[0078]
[0090] FIG. 15 shows a block diagram of a decoding flow process for extracting watermarked data from video frames according to an aspect of the present disclosure. The decoding flow process may be performed by a media device (e.g., a television, a set-top box, etc.). The decoding flow process may begin when a video signal is received 1501. Video frame analysis 1502 is performed on the video signal to identify the presence of a watermark in video frames of the video signal. In some instances, the video frame analysis 1502 identifies a lead-in pattern that indicates the presence of a watermark. A pixel value decoder 1503 extracts a set of pixels corresponding to symbols of a base-n symbol set. For example, in a base-2 symbol set (e.g., binary code), the pixel value decoder identifies pixels corresponding to symbol 0 and pixels corresponding to symbol 1.
[0079]
[0091] Error detection / correction 1504 detects errors in the pixel values of the watermark and attempts to correct them. Errors may be introduced by compression, noise, etc., causing pixel values to be altered. Error correction may be performed according to one or more error correction processes. One error correction process involves receiving duplicate video frames and / or watermarks and averaging pixel values of the two video frames. The average pixel value reduces the effect of pixel values altered by compression, noise, etc., by reducing the difference between the altered pixel values and the true pixel values. Increasing the amount of duplicate frames / watermarks further reduces the effect of errors. Other error correction processes may be performed in addition to or instead of the duplicate frame error correction process.
[0080]
[0092] The watermark code can be decoded by Message Decode 1505, which defines a sequence of symbols from each set of pixels that represent a symbol. Message Decode can remove the lead-in pattern from the sequence of symbols. The lead-in pattern is a subsequence of symbols that indicates the presence of a watermark in a video frame but does not contain any other data. The lead-in pattern can be removed from the sequence of symbols without reducing the loss of any of the watermark data.
[0081]
[0093] Data framing 1506 organizes the sequence of symbols in each video frame. If a set of watermarks representing a large message or large data set is inserted into the set of video frames, data framing 1506 determines the order in which each sequence of symbols in each frame will be placed relative to other sequences of symbols. The decoded message is then output for further processing. For example, if the watermark included executable code, the processor of the media device may then execute the executable code. The sequence of symbols in one or more watermarks may, for example, cause additional information related to the video content to be displayed, swap a video segment with another video segment (e.g., stored in cache memory, retrieved from a remote source, etc.), display a pop-up window, or a combination thereof.
[0082]
[0094] 16 shows a flowchart of an example process for decoding a code from a watermark embedded in a video frame according to an aspect of the present disclosure. At block 1604, a media device receives the video frame. The media device may be a device configured to process or display video, such as, but not limited to, a monitor, a television, a set-top box, or a mobile device such as a smartphone. A first predetermined region of the video frame may include a first set of pixels corresponding to the watermark. For example, the first predetermined region of the video frame may be the top one or two rows of the video frame.
[0083]
[0095] The frame may also include a second predetermined region adjacent to the first predetermined region and including a second set of pixels. The second set of pixels may have pixel values based on pixel values of pixels in a third predetermined region of the video frame (e.g., the remainder of the video frame). For example, the second predetermined region of the video frame may be a boundary region. The boundary region may include one or more rows of pixels having pixel values selected to reduce the perceptibility of the first predetermined region, as described in FIG. 12. For example, if the boundary region includes two or more rows, the pixel values of each row are selected according to a gradient with each row of the boundary region (starting with the row closest to the first predetermined region) having a higher luminance value.
[0084]
[0096] One or more watermark modification processes, as previously described, may be performed on a video frame to reduce the likelihood that a user will be able to perceive data embedded in a first predetermined region of the video frame. For example, the first predetermined region may include pixel values that are modulated (e.g., to convey data) and within a predetermined range to reduce perceptibility. In that example, the difference in pixel values (e.g., luminance values) representing a first symbol and a second symbol may be limited to 40. One or more video frames may include a larger difference (e.g., 80) between the pixel values representing the first symbol and the second symbol as error correction. In another example, a subsequent video frame may include the first predetermined region inverted so that, when displayed consecutively, the first predetermined region may appear as a solid color (e.g., an average of the first and second pixel values). Any combination of one or more watermark modifications may be applied to a video frame before it is received by a media device (or by a media device prior to being displayed to a user).
[0085]
[0097] At block 1608, the media device detects a watermark in a first predetermined region of the video frame. The media device may detect the watermark by detecting modulation of pixels in the first predetermined region. For example, the media device detects one or more pixels representing a first symbol (of a base-n code) and one or more pixels representing a second symbol. In some cases, the media device may detect a lead-in pattern in the predetermined region. The lead-in pattern includes a predetermined sequence of pixel values indicative of the watermark. The watermark may be located after the lead-in pattern along the same row. When the media device detects the lead-in pattern, the media device identifies the watermark.
[0086]
[0098] At block 1612, the media device identifies, in a first predetermined region, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value. A symbol may be represented by a contiguous subset of pixels (e.g., a predetermined amount of adjacent pixels, such as one row of four pixels, two rows of four pixels, etc.). The pixels of the contiguous subset may have similar pixel values (e.g., similar hue, chroma, luminance, RGB values, combinations thereof, etc.) that may correspond to the first pixel value or the second pixel value. For example, if the first pixel value and the second pixel value correspond to luminance values, the media device may identify each contiguous subset of pixels having a first luminance value (e.g., approximately 10-16) and each contiguous subset of pixels having a second luminance value (e.g., approximately 45-55).
[0087]
[0099] The first and second values may be selected to reduce the likelihood that a user will perceive the watermark. A user may perceive the contrast between a low luminance pixel next to a high luminance pixel and detect the presence of the watermark as an artifact in the video frame. When modulating pixels using luminance, the difference between the first luminance value and the second luminance may be approximately 40.
[0088]
[0100] While limiting the difference in luminance values may reduce the likelihood that a user may perceive a watermark (e.g., by adjacent, contrasting luminance values being perceived as flicker or artifacts), it may also increase the error rate of the media device. To reduce the likelihood of errors, a video frame may increase the difference between a first pixel value (representing a first symbol) and a second pixel value (representing a second pixel). For example, every n video frames may receive a video frame in which one or more contiguous subsets of pixels correspond to a first pixel value and one or more contiguous subsets of pixels correspond to a third pixel value. The first pixel value may continue to relate to a luminance value between 10 and 16. The third pixel value may relate to a luminance value of approximately 80. The difference between the third pixel value and the first pixel value may be approximately twice the difference between the second pixel value and the first pixel value. In some cases, m frames may be received having a watermark with an increased difference between a first pixel value and a second pixel value, and after m frames, the difference between the first pixel value and the second pixel value may be reduced (e.g., to about 40).
[0089]
[0101] In block 1616, the media device assigns a first symbol to one or more contiguous subsets of pixels corresponding to the first pixel value and assigns a second symbol to one or more contiguous subsets of pixels corresponding to the second pixel value. The first pixel value and the second pixel value may represent a first symbol and a second symbol of a base-2 symbol set (e.g., a binary code). Additional pixel values may be used to define a base-n symbol set. For example, modulating pixels based on luminance and chroma may create a base-4 symbol set.
[0090]
[0102] At block 1620, the media device generates a sequence of symbols based on the symbols assigned to one or more contiguous subsets of pixels. The sequence of symbols may provide additional information related to the video frame and / or cause the media device to perform several functions. The additional information may include, but is not limited to, information related to the content of the displayed video (e.g., actors, characters, setting, team, production staff, production features, or other aspects or characteristics of the content), metadata related to the displayed video (e.g., resolution, pixel values, broadcast origin, etc.), communications related to the displayed video, etc. The functions may include, but are not limited to, opening a web browser to a specific URL, opening a pop-up window on the video frame, substituting a locally stored or retrieved video segment from a remote server, etc.
[0091]
[0103] The watermark in the next video frame may have the same sequence of symbols but may be inverted. In the next video frame, each contiguous subset of pixels that corresponded to a first pixel value in the previous frame may now correspond to a second pixel value, and each contiguous subset of pixels that corresponded to a second pixel value in the previous frame may now correspond to a second first value. For example, as shown in FIG. 13A, frame 1 includes a first contiguous subset of pixels that correspond to a first pixel value (e.g., a luminance of 10-16) and a subsequent second contiguous subset of pixels that correspond to a second pixel value (e.g., a luminance of 40). In a subsequent frame (e.g., frame 2), the first contiguous subset of pixels corresponds to the second pixel value, and the second contiguous subset of pixels corresponds to the first pixel value. By inverting the pixel values in the subsequent frame, the pixels in the first predetermined region appear as an average of the first and second pixel values (e.g., as a solid color). As a result, a user may not perceive the watermark.
[0092]
[0104] 17 illustrates an exemplary computing device according to aspects of the present disclosure. For example, the computing device 1700 may implement any of the systems or methods described herein. In some instances, the computing device 1700 may be a component of or included within a media device. The components of the computing device 1700 are shown in electrical communication with each other using a connection 1705, such as a bus. The exemplary computing device architecture 1700 includes a processing unit (e.g., a CPU, a processor, etc.) 1710 and connections 1705 (e.g., a bus, etc.) configured to couple components of the computing device 1700, such as, but not limited to, memory 1715, read-only memory (ROM) 1720, random access memory (RAM) 1725, and / or storage devices 1730, to the processing unit 1710.
[0093]
[0105] The computing device 1700 may include a cache 1712 of high-speed memory directly connected to, in close proximity to, or incorporated within the processing unit 1710. The computing device 1700 may copy data from the memory 1715 and / or the storage device 1730 to the cache 1712 for faster access by the processing unit 1710. In this way, the cache 1712 may provide a performance boost that avoids processor 1710 delays while waiting for data. Alternatively, the processing unit 1701 may access data directly from the memory 1715 and / or the storage device 1730. The memory 1715 may include multiple types of memory (e.g., magnetic, optical, solid-state, etc.).
[0094]
[0106] The storage device 1730 may include one or more non-transitory computer-readable media, such as volatile and / or non-volatile memory. The non-transitory computer-readable media may store instructions and / or data accessible by the computing device 1700. The non-transitory computer-readable media may include, but are not limited to, a magnetic cassette, a hard disk drive (HDD), a flash memory, a solid-state memory device, a digital versatile disk, a cartridge, a compact disk, a random access memory (RAM) 1725, a read-only memory (ROM) 1720, combinations thereof, and the like.
[0095]
[0107] A storage device 1730 that stores one or more services, such as service 1 1732, service 2 1734, and service 3 1736, executable by the processing unit 1710 and / or other electronic hardware. The one or more services include instructions executable by the processing unit 1710 to perform operations, such as any of the techniques described herein, controlling the operation of devices in communication with the computing device 1700, controlling the operation of the processing unit 1710 and / or any dedicated processors, combinations therefor, etc. The processing unit 1710 may be a system-on-chip (SOC) that includes one or more cores or processors, buses, memory, clocks, memory controllers, caches, other processor components, etc. Multi-core processors may be symmetric or asymmetric.
[0096]
[0108] Computing device 1700 may include one or more input devices 1745, which may represent any number of input mechanisms, such as a microphone, a touch-sensitive screen for graphical input, a keyboard, a mouse, motion input, voice, a media device, a sensor, combinations thereof, etc. Computing device 1700 may include one or more output devices 1735 that output data to a user. Such output devices 1735 may include, but are not limited to, a media device, a projector, a television, speakers, combinations thereof, etc. In some cases, a multimodal computing device may allow a user to provide multiple types of input to communicate with computing device 1700. Communication interface 1740 may be configured to manage user input and computing device output. Communication interface 1740 may also be configured to manage communications with remote devices (e.g., establish connections, receive / send communications, etc.) over one or more communication protocols and / or over one or more communication media (e.g., wired, wireless, etc.). Computing device 1700 is not limited to the components as shown. Other components may be added, and illustrated components may be omitted.
[0097]
[0109] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instruction(s) and / or data. Computer-readable media may include non-transitory media in which data may be stored in a form excluding carrier waves and / or electronic signals. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. A computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0098]
[0110] While the present subject matter has been described in detail with reference to particular embodiments thereof, it will be appreciated that those skilled in the art, upon achieving an understanding of the foregoing, may readily create modifications to, variations of, and equivalents of, such embodiments. Numerous specific details are described herein to provide a thorough understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter may be practiced without these specific details. In other instances, methods, apparatuses, or systems that would be known to those skilled in the art have not been described in detail so as not to obscure the claimed subject matter. Accordingly, the present disclosure has been presented by way of example and not limitation and does not preclude the inclusion of such modifications, variations, and / or additions to the present subject matter as would be readily apparent to those skilled in the art.
[0099]
[0111] For clarity of explanation, in some instances, the disclosure may be presented as including individual functional blocks, including devices, device components, steps or routines in a method implemented in software, or functional blocks comprising a combination of hardware and software. Additional functional blocks other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.
[0100]
[0112] Individual embodiments may be described above as a process or method, which may be depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. While a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process is terminated when its operations are completed, but may have additional steps that are not shown. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.
[0101]
[0113] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored on or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. A portion of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc.
[0102]
[0114] Devices implementing the methods and systems described herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) to perform the necessary tasks may be stored on a computer-readable or machine-readable medium. The program code may be executed by a processor, which may include one or more processors, such as, but not limited to, one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated circuits or discrete logic circuitry. Such processors may be configured to perform any of the techniques described in this disclosure. The processor may be a microprocessor, a conventional processor, a controller, a microcontroller, a state machine, or the like. A processor may also be implemented as a combination of computing components (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration). Accordingly, the term "processor" as used herein may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein. The functionality described herein may also be embodied in a peripheral device or add-in card. Such functionality may also be implemented on a circuit board among different chips or different processes executing in a single device, as further examples.
[0103]
[0115] In the foregoing description, aspects of the present disclosure have been described with reference to particular examples thereof, but those skilled in the art will recognize that the present disclosure is not limited thereto. Accordingly, while illustrative examples of the present disclosure have been described in detail herein, it should be understood that the inventive concepts may, in some cases, be embodied and employed in various ways, and that the appended claims are intended to be construed to include such variations. The various features and aspects of the disclosure described above may be used individually or in any combination. Furthermore, the examples may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the present disclosure. Accordingly, the present disclosure and the figures are to be considered illustrative and not restrictive.
[0104]
[0116] The various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0105]
[0117] Unless otherwise specified, it should be appreciated that throughout this specification, descriptions utilizing terms such as “processing,” “calculating,” “computing,” “determining,” and “identifying” refer to the actions or processes of a computing device, such as one or more computers or one or more similar electronic computing devices, that manipulate or transform data represented as physical electronic or magnetic quantities in a memory, register, or other information storage device, transmission device, or media device of a computing platform. The use of “adapted to” or “configured to” herein is intended as open and inclusive phraseology that does not exclude devices adapted or configured to perform additional tasks or steps. Furthermore, the use of “based on” is intended to be open and inclusive in that a process, step, calculation, or other action “based on” one or more recited conditions or values may, in fact, be based on additional conditions or values other than those recited. The headings, lists, and numbering contained herein are for ease of description only and are not limiting. The following is a summary of the claims as originally filed: [C1] receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, and a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame; detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel values and the second pixel values; A method for providing the above. [C2] The method of C1, wherein the first predetermined region includes the first two rows of the video frame. [C3] The method of C1, wherein the second predetermined region includes a third row of the video frame. [C4] receiving a subsequent video frame; generating a second sequence of symbols from the subsequent video frame, wherein the second sequence of symbols is an inverted version of the first sequence of symbols; The method of C1, further comprising: [C5] receiving a subsequent video frame, wherein the first predetermined region of the subsequent video frame includes a third set of pixels; identifying, within the third set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a third pixel value, wherein a difference between the third pixel value and the first pixel value is approximately twice the difference between the second pixel value and the first pixel value; generating a second sequence of symbols from said third set of pixels; and The method of C1, further comprising: [C6] The method of C1, wherein the first pixel value corresponds to a luminance value. [C7] Detecting the watermark in the first predetermined region of the video frame comprises: detecting a lead-in pattern of pixels within the first predetermined region, the lead-in pattern of pixels being distinct from the first sequence of symbols; The method according to C1, comprising: [C8] one or more processors; a non-transitory computer-readable medium storing instructions; wherein the instructions, when executed by the one or more processors, cause the one or more processors to: receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, and a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame; detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel values and the second pixel values; A system that performs an operation including: [C9] The system of C8, wherein the first predetermined region includes the first two rows of the video frame. [C10] The system of C8, wherein the second predetermined region includes a third row of the video frame. [C11] The operation is receiving a subsequent video frame; generating a second sequence of symbols from the subsequent video frame, wherein the second sequence of symbols is an inverted version of the first sequence of symbols; The system of C8, further comprising: [C12] The operation is receiving a subsequent video frame, wherein the first predetermined region of the subsequent video frame includes a third set of pixels; identifying, within the third set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a third pixel value, wherein a difference between the third pixel value and the first pixel value is approximately twice the difference between the second pixel value and the first pixel value; generating a second sequence of symbols from said third set of pixels; and The system of C8, further comprising: [C13] The system of C8, wherein the first pixel value corresponds to a luminance value. [C14] Detecting the watermark in the first predetermined region of the video frame comprises: detecting a lead-in pattern of pixels within the first predetermined region, the lead-in pattern of pixels being separated from the first sequence of symbols; 10. The system of claim 8, comprising: [C15] A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, and a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame; detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel values and the second pixel values; A non-transitory computer-readable medium for performing operations including: [C16] The non-transitory computer-readable medium of C15, wherein the first predetermined region includes the first two rows of the video frame. [C17] The non-transitory computer-readable medium of C15, wherein the second predetermined region includes a third row of the video frame. [C18] The operation receiving a subsequent video frame; generating a second sequence of symbols from the subsequent video frame, wherein the second sequence of symbols is an inverted version of the first sequence of symbols; 15. The non-transitory computer-readable medium of claim 14, further comprising: [C19] The operation receiving a subsequent video frame, wherein the first predetermined region of the subsequent video frame includes a third set of pixels; identifying, within the third set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a third pixel value, wherein a difference between the third pixel value and the first pixel value is approximately twice the difference between the second pixel value and the first pixel value; generating a second sequence of symbols from said third set of pixels; and 15. The non-transitory computer-readable medium of claim 14, further comprising: [C20] Detecting the watermark in the first predetermined region of the video frame comprises: detecting a lead-in pattern of pixels within the first predetermined region, the lead-in pattern of pixels being distinct from the first sequence of symbols; 15. A non-transitory computer-readable medium as described in C15, comprising:
Claims
1. receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, and a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame; detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel value and the second pixel value; receiving a subsequent video frame; generating a second sequence of symbols from the subsequent video frame, wherein the second sequence of symbols is an inverted version of the first sequence of symbols; A method for providing the above.
2. Receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame. detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel value and the second pixel value; receiving a subsequent video frame, wherein the first predetermined region of the subsequent video frame includes a third set of pixels; identifying, within the third set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a third pixel value, wherein a difference between the third pixel value and the first pixel value is approximately twice the difference between the second pixel value and the first pixel value; generating a second sequence of symbols from said third set of pixels; and A method for providing the above.
3. The method of claim 1 or 2, wherein the first predetermined region comprises the first two rows of the video frame.
4. The method of claim 1 or 2, wherein the second predetermined region comprises the third row of the video frame.
5. The method of claim 1 or 2, wherein the first pixel value corresponds to a luminance value.
6. Detecting the watermark in the first predetermined region of the video frame comprises: detecting a lead-in pattern of pixels within the first predetermined region, the lead-in pattern of pixels being distinct from the first sequence of symbols; 3. The method of claim 1 or 2, comprising:
7. one or more processors; a non-transitory computer-readable medium storing instructions; wherein the instructions, when executed by the one or more processors, cause the one or more processors to: receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, and a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame; detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel value and the second pixel value; receiving a subsequent video frame; generating a second sequence of symbols from the subsequent video frame, wherein the second sequence of symbols is an inverted version of the first sequence of symbols; A system that performs an operation including:
8. One or more processors; a non-transitory computer-readable medium storing instructions; wherein the instructions, when executed by the one or more processors, cause the one or more processors to: receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, and a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame; detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel value and the second pixel value; receiving a subsequent video frame, wherein the first predetermined region of the subsequent video frame includes a third set of pixels; identifying, within the third set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a third pixel value, wherein a difference between the third pixel value and the first pixel value is approximately twice the difference between the second pixel value and the first pixel value; generating a second sequence of symbols from said third set of pixels; and A system that performs an operation including:
9. 9. The system of claim 7 or 8, wherein the first predetermined region comprises the first two rows of the video frame.
10. 9. The system of claim 7 or 8, wherein the second predetermined region comprises a third row of the video frame.
11. 9. The system of claim 7 or 8, wherein the first pixel value corresponds to a luminance value.
12. Detecting the watermark in the first predetermined region of the video frame comprises: detecting a lead-in pattern of pixels within the first predetermined region, the lead-in pattern of pixels being separated from the first sequence of symbols; 9. The system of claim 7 or 8, comprising:
13. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, and a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame; detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel value and the second pixel value; receiving a subsequent video frame; generating a second sequence of symbols from the subsequent video frame, wherein the second sequence of symbols is an inverted version of the first sequence of symbols; A non-transitory computer-readable medium for performing operations including:
14. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receiving a video frame, wherein a first predetermined region of the video frame includes a first set of pixels, and a second predetermined region of the video frame includes a second set of pixels, the second set of pixels having pixel values based on pixel values of pixels of a third predetermined region of the video frame; detecting a watermark in the first predetermined region of the video frame; identifying, within the first set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a second pixel value; assigning a first symbol to the one or more contiguous subsets of pixels corresponding to a first pixel value and assigning a second symbol to the one or more contiguous subsets of pixels corresponding to a second pixel value; generating a first sequence of symbols based on the first pixel value and the second pixel value; receiving a subsequent video frame, wherein the first predetermined region of the subsequent video frame includes a third set of pixels; identifying, within the third set of pixels, one or more contiguous subsets of pixels corresponding to a first pixel value and one or more contiguous subsets of pixels corresponding to a third pixel value, wherein a difference between the third pixel value and the first pixel value is approximately twice the difference between the second pixel value and the first pixel value; generating a second sequence of symbols from said third set of pixels; and A non-transitory computer-readable medium for performing operations including:
15. 15. The non-transitory computer-readable medium of claim 13 or 14, wherein the first predetermined region comprises the first two rows of the video frame.
16. 15. The non-transitory computer-readable medium of claim 13 or 14, wherein the second predetermined region comprises a third row of the video frame.
17. Detecting the watermark in the first predetermined region of the video frame comprises: detecting a lead-in pattern of pixels within the first predetermined region, the lead-in pattern of pixels being distinct from the first sequence of symbols; 15. The non-transitory computer-readable medium of claim 13 or 14, comprising:
Citation Information
Patent Citations
How to embed fingerprints for multimedia content identification
JP2005537731A
Content information acquisition device and program, and content distribution device
JP2015106862A
Method and system for synchronizing video and data
JP2019508991A
System for distributing metadata embedded in video
US20160277793A1
Embedding video watermarks without visible impairments
US20190261012A1