WATERMARKS FOR DIGITAL IMAGES
The computer system addresses the challenge of embedding and decoding watermarks in images by pre-processing and robustly embedding watermarks in images, ensuring effective protection against unauthorized use.
Patent Information
- Application Number
- DE112023002946
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-05
- Filing Date
- 2023-02-09
- Publication Date
- 2025-05-08
AI Technical Summary
Existing technologies face challenges in effectively embedding and decoding watermarks in images to prevent unauthorized use, especially when images are modified or distributed.
A computer system processes an image by pre-processing it and embedding a watermark through channel-coding and selective block embedding, ensuring the watermark is robust and imperceptible.
The solution enables the embedding of identifying information in images that can be reliably decoded even after modifications, effectively linking the image to its original creator and preventing unauthorized use.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND TECHNICAL AREA
[0001] This application relates to embedding a watermark in an image and decoding an image to recover an embedded watermark. STATE OF THE ART
[0002] Images are often original works that are subject to protection. Images are increasingly being shared and publicly disseminated to a wider audience. Sharing and distributing images can lead to unauthorized uses. For example, an image may be used in a way that is not authorized by the image creator or a rights holder. Preventing unauthorized uses often requires establishing that the image used is actually derived from or copied from the protected image. SUMMARY OF THE INVENTION
[0003] In one embodiment, a computer system comprises a processor and a memory having stored thereon executable instructions that, when executed by the processor, cause the processor to receive a first image and receive information for embedding in the first image, preprocess the first image to generate a preprocessed first image, preprocess the information by at least channel coding the information to generate preprocessed information including a plurality of sets of bits, embed the preprocessed information in the preprocessed first image by at least selecting a plurality of blocks of the preprocessed first image and embedding a respective set of bits of the plurality of sets of bits into each block of the plurality of blocks.In one embodiment, each set of bits of the plurality of sets of bits is embedded in a minimum number of blocks of the plurality of blocks, wherein the minimum number of blocks is greater than one.
[0004] In one embodiment, the executable instructions cause the processor to preprocess the first image by rotating the first image by 90 degrees, 180 degrees, or 270 degrees. In one embodiment, the executable instructions cause the processor to preprocess the first image by performing feature detection on the first image to generate a plurality of feature vectors, determining an average feature vector of the plurality of feature vectors, and rotating the first image to a quadrant that minimizes an angle between the average feature vector and a central axis of the first image.
[0005] In one embodiment, the executable instructions cause the processor to preprocess the first image by obtaining the preprocessed first image as a luma channel component of the first image. In one embodiment, the executable instructions cause the processor to frequency transform the plurality of blocks to generate a plurality of frequency-transformed blocks. In one embodiment, the executable instructions cause the processor to embed a respective set of bits of the plurality of sets of bits into each block of the plurality of blocks by embedding the respective set of bits of the plurality of sets of bits into each frequency-transformed block of the plurality of frequency-transformed blocks, and in response to embedding the plurality of sets of bits, inversely transforming the plurality of frequency-transformed blocks into a spatial domain.
[0006] In one embodiment, the executable instructions cause the processor to embed a first set of bits of the plurality of sets of bits in a first block of the plurality of blocks by at least selecting a first bit of the first set of bits, selecting a pixel position in the first block, identifying a level of the pixel position, and quantizing the level of the pixel position in dependence on a logical state of the first bit.In one embodiment, the executable instructions that cause the processor to quantize the level of the pixel position depending on the logical state of the first bit cause the processor, when it is determined that the logical state of the first bit is a logical zero, to quantize the level to one of an odd value and an even value, and when it is determined that the logical state of the first bit is a logical one, to quantize the level to the other of the odd value and the even value.
[0007] In one embodiment, the executable instructions cause the processor to divide the quantized level of the pixel position by a robustness factor. In one embodiment, the executable instructions cause the processor to generate a watermarked image based on embedding the preprocessed information in the preprocessed first image and cause the watermarked image to be output for distribution. In one embodiment, the information is a decentralized identifier defined by World Wide Web Consortium (W3C) Proposed Recommendation 1.0.
[0008] In one embodiment, the executable instructions cause the processor to preprocess the first image by establishing an image processing pipeline that includes a plurality of image preprocessing stages, selecting one or more of the plurality of image preprocessing stages depending on contextual information associated with the first image, a property of the first image, an image type associated with the first image, or a preprocessing configuration, and preprocessing the first image by the selected one or more of the plurality of image preprocessing stages to generate the preprocessed first image.
[0009] In one embodiment, the executable instructions cause the processor to preprocess the information by establishing an information processing pipeline including a plurality of information preprocessing stages, selecting one or more of the plurality of information preprocessing stages depending on the type of information or a preprocessing configuration, and preprocessing the information by the selected one or more of the plurality of information preprocessing stages to generate the preprocessed information.
[0010] In one embodiment, a method comprises receiving a first image and receiving information for embedding in the first image, preprocessing the first image to generate a preprocessed first image, preprocessing the information by at least channel coding the information to generate preprocessed information including a plurality of sets of bits, embedding the preprocessed information in the preprocessed first image by at least selecting a plurality of blocks of the preprocessed first image and embedding a respective set of bits of the plurality of sets of bits in each block of the plurality of blocks. In one embodiment, each set of bits of the plurality of sets of bits is embedded in a minimum number of blocks of the plurality of blocks, wherein the minimum number of blocks is greater than one.
[0011] In one embodiment, preprocessing the first image comprises performing feature detection on the first image to generate a plurality of feature vectors, determining an average feature vector of the plurality of feature vectors, and rotating the first image to a quadrant that minimizes an angle between the average feature vector and a central axis of the first image. In one embodiment, preprocessing the first image comprises obtaining the preprocessed first image as a luma channel component of the first image.In one embodiment, the method comprises embedding a first set of bits of the plurality of sets of bits into a first block of the plurality of blocks by at least selecting a first bit of the first set of bits, selecting a pixel position in the first block, identifying a level of the pixel position, and quantizing the level of the pixel position depending on a logical state of the first bit.
[0012] In one embodiment, quantizing the level of the pixel position depending on the logical state of the first bit comprises: if the logical state of the first bit is determined to be a logical zero, quantizing the level to one of an odd value and an even value, and if the logical state of the first bit is determined to be a logical one, quantizing the level to the other of the odd value and the even value. In one embodiment, the method comprises dividing the quantized level of the pixel position by a robustness factor.
[0013] In one embodiment, preprocessing the first image comprises establishing an image processing pipeline including a plurality of image preprocessing stages, selecting one or more of the plurality of image preprocessing stages depending on contextual information associated with the first image, a property of the first image, an image type associated with the first image, or a preprocessing configuration, and preprocessing the first image by the selected one or more of the plurality of image preprocessing stages to generate the preprocessed first image.
[0014] In one embodiment, preprocessing the information comprises establishing an information processing pipeline including a plurality of information preprocessing stages, selecting one or more of the plurality of information preprocessing stages depending on the type of information or a preprocessing configuration, and preprocessing the information by the selected one or more of the plurality of information preprocessing stages to generate the preprocessed information. SHORT DESCRIPTION OF THE CHARACTERS Fig. Figure 1 shows an example environment for digitally encoding and retrieving a watermark. Fig. Figure 2A shows the steps by which a first image is watermarked. Fig. Figure 2B shows sub-stages of a coding stage. Fig. 3A - Fig. 3D shows a flow diagram of a method for preprocessing the first image by an image processing pipeline, preprocessing information by an information processing pipeline, and operating the encoding stage. Fig. Figure 4A shows an example of the preprocessing of the first image by the image processing pipeline. Fig. Figure 4B illustrates the generation of a watermarked image. Fig. Figure 5 shows an example of a three-stage discrete wavelet transform. Fig. Figure 6 shows the steps by which watermark information is extracted from a second image. Fig. 7A and Fig. 7B shows a flowchart of a method for extracting information from the second image. Fig. Figure 8 shows an example of a branch decoding trellis. Fig. Figure 9 shows a block diagram of a computer system. DETAILED DESCRIPTION
[0015] The following provides the use of a digital watermark to embed identifying information into an image, which may be an original work or a work of art. The identifying information can be retrieved even if the image has been modified, allowing a link to be created between the image and a record of the identifying information. These techniques allow artists and creators of original works, among other things, to share their work on a public platform. In the event that an image is used in an unauthorized manner (e.g., as part of a media campaign without permission), a link can be created between the image and misuse.
[0016] Fig. 1 shows an example of an environment 100 for digitally watermarking an image. The environment 100 includes first and second computer systems 102, 104. The first computer system 102 receives a first image 106 and processes the first image 106 to generate a watermarked image 108. The first image 106 may, for example, be a work of art or an original work. The first image 106 may be intended to be disseminated, distributed, or shared. However, the owner or copyright holder of the first image 106 may wish to prevent unauthorized uses of the first image 106. For example, the owner may wish to prevent unauthorized use of the first image 106 in a media campaign, among other uses.
[0017] The first computer system 102 receives the first image 106 and processes the first image 106 to generate a watermarked image 108. Embedded or encoded in the watermarked image 108 is information or data (a watermark) that is retrievable to identify the watermarked image 108. The watermarked image 108 may then be disseminated, distributed, or shared. When the watermarked image 108 is shared, the watermarked image 108 may, among other things, be edited, cropped, compressed, or formatted, resulting in a second image 110. Alternatively, the watermarked image 108 may not be altered during distribution, and thus the second image 110 may be the same image as the watermarked image 108.The second computer system 104 receives the second image 110. The second computer system 104 processes the second image 110 to determine whether the watermark is encoded or embedded in the second image 110. For example, the second computer system 104 may decode the second image 110 and determine whether the watermark is present in the second image 110.
[0018] If the watermark is present in the second image 110, it can be determined that the second image 110 originates from the first image 108.
[0019] For example, the second image 110 may be a copy of the watermarked image 108. Instead, the watermarked image 108, or a portion thereof, may have been cropped or its format converted, among other things, to produce the second image 110, or a portion thereof.
[0020] The digital watermark is preferably incorporated into the watermarked image 108 in such a way that it is imperceptible to the human eye. Visual artifacts associated with the watermark are preferably embedded in (or blended with) the watermarked image 108 in such a way that they are imperceptible to the human eye. Although the first image 106 is modified by the watermark to produce the watermarked image 108, the modification is performed with respect to features of the watermarked image 108 that are different from those on which the human eye concentrates or focuses.Accordingly, the watermarked image 108 may be perceived by a human viewer as a copy of the first image 106, whereas the watermarked image 108 actually contains identifying information that is not part of the first image 106.
[0021] Furthermore, a watermark should be robust. Thus, the watermark should survive cropping and format conversion, among other measures that may be taken to generate the second image 110 from the watermarked image 108. Furthermore, the watermark may also be created using frequency-domain or time-domain processing or preprocessing.
[0022] Fig. Figure 2A shows the steps by which the first image 106 is watermarked. The steps for watermarking the first image 106 may be performed by the first computer system 102 as described below. The steps include preprocessing steps performed as part of the preprocessing prior to encoding the first image 106. The first computer system 102 receives the first image 106 and information 112 to be embedded in the first image 106 using a watermark.
[0023] The first computer system 102 performs various preprocessing stages of an image processing pipeline 114 on the first image 106. The first computer system 102 also performs various preprocessing stages of an information processing pipeline 116 on the information 112. The image processing pipeline 114 includes an image rotation adjustment stage 118, a color space conversion stage 120, and a channel extraction stage 122. The information processing pipeline 116 includes a validation stage 124 and a compression stage 126. The compression stage 126 described herein includes channel coding (e.g., convolutional coding) of the information 112.
[0024] The preprocessed image output from image processing pipeline 114 and the preprocessed information output from information processing pipeline 116 are fed to an encoding stage 128. Encoding stage 128 embeds the preprocessed information into the preprocessed image. Encoding stage 128 generates image data with the preprocessed image embedded with the preprocessed information. Encoding stage 128 outputs the image data to a post-processing stage 130. Post-processing stage 130 may perform "inverse" operations that are the reverse of those performed in image processing pipeline 114. Post-processing stage 130 may include channel combining, color space reconversion, and rotation adjustment. For example, the rotation adjustment may reverse the image rotation adjustment performed as part of image rotation adjustment stage 118.The post-processing stage 130 outputs the watermarked image 108.
[0025] It should be noted that various stages of the image and information processing pipelines 114, 116 may be omitted depending on the image and the desired watermarking. For example, in one embodiment, the image rotation adjustment stage 118 may be omitted. Instead, the first image 106 may be processed by the image processing pipeline 114 without changing the orientation of the first image 106. For example, the first image 106 may also be processed without changing its orientation or rotating it so that it is upside down.
[0026] Furthermore, stages may also be added to the image or information processing pipelines 114, 116. For example, a feature detection stage or an image normalization stage may also be added to the image processing pipeline 114. For example, such a feature detection stage may be used by the first computer system 102 to identify a category of the first image 106. The category may be, among other things, one or more categories such as photograph or line drawing. The coding technique used in the coding stage 128 may be changed or adapted depending on the identified category. Furthermore, the stages included in the image processing pipeline 114 may be changed depending on the image category. Furthermore, contextual information may influence the stages used as part of the image processing pipeline 114.Contextual information may, for example, be the type of software used to generate the first image 106. For example, if the image is generated using a specific type of software, color space conversion may be omitted.
[0027] The first computer system 102 may form the image processing pipeline 114. The first computer system 102 may selectively add a plurality of image preprocessing stages to the image processing pipeline 114. The first computer system 102 may also determine the plurality of image preprocessing stages to be added to the image processing pipeline 114 depending on contextual information associated with the first image (such as the software used to generate the first image), properties of the first image (such as the color space of the first image, a size of the first image, or a dimension of the first image), an image type associated with the first image (such as whether the first image represents a line drawing or a photograph, or even the type of photograph), or a preprocessing configuration (such as a user setting of the level of computational intensity to be used to perform image preprocessing).For example, if it is selected that the computational intensity should be limited, the first computer system 102 may minimize the number of preprocessing stages, and vice versa. The first computer system 102 may then preprocess the first image through the multiple image preprocessing stages to generate a preprocessed first image.
[0028] In the information processing pipeline 116, the validation stage 124 may also be omitted if the information 112 does not have a specific format to be validated. As described below, the validation stage 124 validates that the information is in the form of a decentralized identifier (DID). However, if the first image 106 is to be watermarked with information of any format, then the validation stage 124 does not necessarily need to be used to ensure that the information is in the form of a DID. Furthermore, the compression stage 126 may also be omitted if the information is to be embedded without compression.
[0029] The first computer system 102 may configure the information processing pipeline 116 according to a modular approach. The first computer system 102 may selectively add a plurality of information preprocessing stages to the information processing pipeline 116. The first computer system 102 may determine the plurality of information preprocessing stages to be added to the information processing pipeline 116 depending on the type of information or the preprocessing configuration. The first computer system 102 may then preprocess the information through the plurality of information preprocessing stages to generate the first information image. The type of information may be, for example, the format of the information. If the information is to have a specific format, the first computer system 102 may configure the information processing pipeline 116 with a preprocessing stage to validate the information.
[0030] The use of a pipeline-based technique is advantageous in that watermark encoding and decoding can be adapted to changing requirements and environments. Various preprocessing and postprocessing stages can be added or omitted to accommodate different image types, information types, and / or encoding or decoding requirements.
[0031] Fig. 2B shows sub-stages of the coding stage 128. The coding stage 128 includes a block selection sub-stage 132, a frequency transformation sub-stage 134, an embedding sub-stage 136, and an inverse frequency transformation sub-stage 138 (or sub-stage for inverse frequency transformation).
[0032] The block selection substage 132 may select a block or blocks of the preprocessed image into which the preprocessed information, or a portion thereof, is to be embedded. The frequency transformation substage 134 may transform the selected block from a spatial domain to a frequency domain. The embedding substage 136 may embed the preprocessed information, or a portion thereof, into the frequency domain representation of the selected block. The inverse frequency transformation substage 138 may transform the frequency domain representation of the selected block back to the spatial domain. As described herein, the spatial domain representation is provided as image data to the post-processing stage 130 for generating the watermarked image 108.
[0033] Fig. 3A - Fig. 3D shows a flow diagram of a method 300 for preprocessing the first image 106 by the image processing pipeline 114, for preprocessing the information 112 by the information processing pipeline 116, and for operating the encoding stage 128 to encode the preprocessed first image with the preprocessed information.
[0034] The first computer system 102 receives the first image 106 and the information 112 at step 302. The first computer system 102 passes the first image 106 to the image processing pipeline 114 at step 304 and passes the information 112 to the information processing pipeline 116. In the image processing pipeline 114, the first computer system 102 preprocesses the first image 106 through the image rotation adjustment stage 118. In the image rotation adjustment stage 118, the first computer system 102 rotates the first image 106.
[0035] Preprocessing of the first image 106 in the image rotation adjustment stage 118 is performed by identifying a local vertical axis associated with the image and rotating the image such that the orientation of the image is reversed (or upside down, or such that the image is rotated 180 degrees). The image rotation adjustment stage 118 advantageously standardizes the image rotation or orientation for watermark encoding and detection. The image rotation adjustment stage 118 increases resilience against rotation attacks, which use an adjustment or change in rotation to evade watermark detection. The image rotation adjustment stage 118 uses features of the first image 106 to identify a local vertical axis inherent in the image.The local vertical axis is used to specify the orientation of the first image 106 during subsequent preprocessing and encoding.
[0036] As part of the image rotation adjustment stage 118, the first computer system 102 loads the first image 106 in the red-green-blue (RGB) color space in step 306. The first computer system 102 subjects the first image to feature detection in step 308 to obtain feature vectors of the image. Feature detection associates each feature or element of a number of features or elements with a feature vector having a magnitude and a direction. The first computer system 102 may use any feature detection technique (to determine the feature vectors), such as accelerated robust features (SURF), scale-invariant feature transform (SIFT), or features from accelerated segment test (FAST).
[0037] The first computer system 102 determines an average vector of the feature vectors of the image at step 310. The first computer system 102 rotates the image at step 312 to cause the average vector to point along a central axis of the image. The rotation may also be performed to a next quadrant. The quadrant may minimize an angle between the average vector and the central axis of the first image 106. The rotation to the next quadrant preserves the rectangular properties of the image.
[0038] Fig. Figure 4A shows an example of the preprocessing of the first image 106 by the image processing pipeline 114. After determining the feature vectors of the first image 106, the first computer system 102 determines the average vector of the feature vectors. Then, the first computer system 102 rotates the first image 106 so that the average vector points in the direction of a central axis of the first image 106. The first computer system 102 approximates the rotation to the nearest quadrant to keep the first image 106 square or rectangular. Fig. 4A, the first computer system 102 rotates the first image 106 by 180°. However, other rotation options include 90° and 270°. Furthermore, if the average vector is already aligned to the central axis, the rotation is 0° and the image is not rotated. As described herein, feature detection sets a convention for the orientation of the first image 106 when the first image undergoes the processing described herein. Rotating the first image 106 reduces the perceptibility of the watermark because the human eye is less likely to notice the watermark if the watermark is embedded in an image where the image orientation is reversed with respect to the natural orientation of the image.
[0039] With reference again to the Fig. 3A to 3D, the color space conversion stage 120 will now be described. In the color space conversion stage 120, the first computer system 102 converts the image from the RGB color space to the YCrCb color space at step 314. In the YCrCb color space, Y is a luma channel component, Cr is a red-difference chroma channel component, and Cb is a blue-difference chroma channel component. The information 112 can be embedded in the luma channel component (luminance channel) with a higher degree of modification and without increasing the perceptibility of the changes to the human eye.
[0040] It should be noted that the information can also be embedded in channels of the RGB color space. In particular, due to the fact that the green channel is associated with lower perceptibility than the red or blue channels, the information can be encoded in the green channel. In this case, an RGB-to-YCrCb color space conversion for the purpose of embedding the information can be omitted.
[0041] In the channel extraction stage 122, the first computer system 102 extracts the Y channel from the YCrCb color space at step 316. The first computer system 102 splits the Y channel into n×n bit blocks at step 318, where n can be any number of bits. For example, the block size can be 4×4 or 12×12 bits or larger. The block size can also be 32×32 or 64×64. If the block size is smaller than 4×4 or 12×12 bits, embedding the information 112 in the block can change the properties of the blocks such that they become perceptible to the human eye. Instead, selecting a relatively large block size (such as 1024×1024) may not be sufficiently robust against attacks to allow decoding or retrieval of the embedded information. Furthermore, although an n×n block is described here, the blocks into which the image is divided can also be rectangular (e.g.,mxn) instead of being quadratic.
[0042] Fig. Figure 4B illustrates the generation of a watermarked image 108. As in Fig. As shown in Figure 4B, the first computer system 102 converts the rotated first image from the RGB color space to the YCrCb color space in the color space conversion stage 120. The first computer system 102 extracts the Y channel from the YCrCb color space and splits the Y channel into n×n bit blocks in the channel extraction stage 122.
[0043] With further reference to the Fig. 3A- Fig. 3D, the validation and compression stages 124, 126 of the information processing pipeline 116 are set up to preprocess a decentralized identifier (DID). A DID is a globally unique persistent identifier. An example DID is defined by Proposed Recommendation 1.0 of the World Wide Web Consortium (W3C). A DID may have the form "did:example:12345abde." The "did" in the first part of the DID is a Uniform Resource Identifier (URI) scheme identifier. The "example" in the second part of the DID identifies a "DID method," which is a specific method used to generate a specific identifier. The third part of the DID is a specific identifier associated with the "DID method." A DID may be licensed under the Creative Rights Initiative of Wacom Co., Ltd. (known as the “DID method”) and can, for example, have the form “did:cri:12345abde”.
[0044] In the validation stage 124, the first computer system 102 determines in step 320 whether the information represents a valid DID. Determining whether the information represents a valid DID may include determining whether the information is in the format of a DID and / or determining whether the third part (the specific identifier) is a valid identifier (e.g., of a creative work to be watermarked). If the result of the determination is "no," the first computer system 102 outputs in step 322 that the information does not have a supported format, which terminates the preprocessing.
[0045] If the result of the determination is "yes," the first computer system 102 preprocesses the information using compression stage 126. In compression stage 126, the first computer system 102 removes the first and second portions of the information in step 324. However, the first computer system 102 retains the third portion (the specific identifier) of the information. The first computer system 102 compresses the information in step 326. For example, if the third portion (the specific identifier) is in Unicode Transformation Format 8 (UTF-8) format, the first computer system 102 may convert each character into four bits. Each character can be represented by one to four bytes in UTF-8. Converting to four bits results in the representation being reduced by one-eighth to one-half. The third portion may consist of 32 characters, which, when converted to four bits per character, results in an output of 128 bits.
[0046] In step 328, the first computer system 102 encodes the information using a channel encoder. The channel encoder may be a convolutional encoder. The channel encoding may also add parity bits or checksums to the information to enable error recovery. Because the convolutional encoder is sliding in nature, a time-invariant trellis decoder may be used at the decoder end. The trellis decoder enables maximum likelihood decoding.
[0047] In encoding stage 128, first computer system 102 performs block selection on the vector of n×n-bit Y-channel blocks as part of block selection substage 132. At 330, first computer system 102 starts a pseudorandom number generator (PRNG) with a known seed value. The seed value may be known to second computer system 104. Accordingly, when decoding to recover the information, second computer system 104 may use the same seed value to identify blocks in which the information, or a portion thereof, is watermarked. When seeded with the same value, the PRNG generates the same series of numbers across respective iterations.
[0048] The first computer system 102 uses the PRNG at step 332 to identify a next n×n block from the vector to embed information therein. In a first iteration, the next n×n block is a first n×n block of the vector to be selected for embedding information. Using the PRNG distributes and randomizes the watermark across the Y channel of the first image 106.
[0049] It should be noted, however, that the use of a PRNG can be omitted. Alternatively, a block at (or near) a center of the Y-channel can be identified and used to embed information into it. Subsequent blocks can be identified based on a pattern, such as a spiral pattern, that starts at the center and continues outward to an edge or periphery of the Y-channel. Prioritizing blocks at or near a spatial center of the image is advantageous in that the blocks are less likely to be subjected to clipping. The blocks are more likely to be retained when the image is distributed or split, and thus more likely to be subsequently available for watermark detection.
[0050] Furthermore, feature-based identification can also be used to identify blocks in which information is to be encoded. The n×n blocks of the Y channel can, for example, be ordered depending on the strength of features within a block. The block with the highest feature strength can be selected first. Subsequent blocks can be selected in descending order of feature strength.
[0051] Instead, feature-based identification can be used to identify the block with the greatest feature strength (as determined by a feature detector). The block with the greatest feature strength can be considered the center of the image and can be selected for encoding. The selection of subsequent blocks can then be performed in a spiral or circular fashion away from the first block and toward the boundaries of the Y channel of the image.
[0052] In the frequency transformation substage 134, the first computer system 102 performs a frequency transformation on the n×n block to generate a frequency-domain representation of the block. Although a discrete wavelet transform (DWT) is described here, the techniques described herein can be used for any type of frequency-domain transform. The first computer system 102 performs a three-stage discrete wavelet transform on the block in step 334. The first computer system 102 performs the DWT in three stages sequentially.
[0053] Fig. 5 shows an example of a three-stage DWT. In a first operation, the first computer system 102 performs a DWT on a block 502. The first computer system 102 then selects a first region 504 of the DWT-transformed block. The first region 504 may be a region comprising low-frequency portions of the two dimensions of the DWT-transformed block (referred to as "LL1"). The DWT-transformed block may include, in addition to the LL region, an HH region comprising high-frequency portions of the two dimensions, an LH region comprising a low-frequency portion of a first dimension and a high-frequency portion of a second dimension, and an HL region comprising a high-frequency portion of the first dimension and a low-frequency portion of the second dimension. In a second operation, the first computer system 102 performs a DWT on the first region 504.The first computer system 102 then selects a second region 506 ("LL2") of the twice-transformed block. The second region 506 may again be the LL region. In a third operation, the first computer system 102 performs a DWT on the second region 506. The first computer system 102 selects a third region 508 ("LL3"), which is the LL region of the transformed block. Selecting the third region 508 for information encoding is more robust against compression attacks. This is due to the fact that compression algorithms use the third region 508 to compress image data. Accordingly, placing the encoded information or watermark in the third region 508 makes it more likely that the watermark will survive compression.
[0054] With further reference to the Fig. 3A- Fig. 3D, in step 336, the first computer system 102 selects a portion of the DWT-transformed block. In step 338, the first computer system 102 selects the next m data bits from the information. For each n×n block, the first computer system 102 may select m data bits to encode in that block, where m may be eight, for example. The information may be provided by the information processing pipeline 116 or its compression stage 126, as described herein. In the encoding stage 128, the first computer system 102 encodes the portion with the m data bits.
[0055] It is noted that each set of m data bits may be encoded a minimum number of times (z) (i.e., at least z times) as described herein. The first computer system 102 may, in step 338, track the number of times the set of m bits has been selected and determine whether the set of m bits has already been selected the minimum number of times (z). The first computer system 102 may select a different set of m bits if it is determined that a previously selected set of m bits has already been selected and encoded the minimum number of times (z).
[0056] In the encoding stage 128, the first computer system 102 identifies a pixel of the selected region (of the DWT-transformed block) and the value (Q) of that pixel in step 340. The first computer system 102 obtains a next bit (b) of the m data bits in step 342. The positions of the pixels used to encode the m data bits may be predetermined or predefined. The first computer system 102 may be configured to select the pixel positions for encoding the m data bits for each region. The first computer system 102 determines in step 344 whether bit (b) of the m data bits is a logical zero or a logical one. The first computer system 102 then modifies the value (Q) of the pixel depending on the state of bit (b).
[0057] The first computer system 102 is provided with a robustness factor. The robustness factor establishes a relationship between robustness and imperceptibility. The robustness factor is negatively correlated with imperceptibility. A higher robustness factor results in a more robust watermark but makes a watermark more susceptible to perceptibility.
[0058] If it is determined that bit (b) of the m data bits is a logical one, in step 346, the first computer system 102 quantizes the value (Q) of the selected range to an even value (e.g., the nearest even value or the highest or lowest even value). The first computer system 102 divides the even value by the robustness factor to obtain a replacement value (Q'). The first computer system 102 replaces the value (Q) of the selected range with the replacement value (Q').
[0059] If it is determined that bit (b) of the m data bits is a logical zero, in step 346, the first computer system 102 quantizes the value (Q) of the selected range to an odd value (e.g., the next odd value or the highest or lowest odd value). The first computer system 102 divides the odd value by the robustness factor to obtain a replacement value (Q'). The first computer system 102 replaces the value (Q) of the selected range with the replacement value (Q').
[0060] Bit (b) is encoded by changing a pixel of the transformed region according to different levels. A logical one is specified by setting the pixel of the transformed region to a ratio of an even quantization and the robustness factor. A logical zero is specified by setting the pixel of the transformed region to a ratio of an odd quantization and the robustness factor. Note that this convention can also be reversed, and even quantization can be used for logical zero and odd quantization for logical one. Quantization to odd and even values is robust and also imperceptible in that it has a small footprint and does not dramatically change the properties of the block.
[0061] The first computer system 102 determines in step 350 whether the m data bits selected from the information were encoded using robustness factor quantization. If not, the first computer system 102 returns to obtain a next pixel and its next value (Q) from the selected region of the DWT-transformed block and obtain a next bit (b) for encoding using the next value (Q). The first computer system 102 may determine "yes" in step 350 if the first computer system 102 has encoded all m selected bits in corresponding m pixel positions using the even and odd quantization described herein.
[0062] If "yes," then in step 352, the first computer system 102 performs an inverse three-stage DWT on the selected region to convert the selected region back to an n×n block in spatial format. In step 352, the first computer system 102 replaces the n×n block (identified in step 332) in the vector with the watermarked version of the n×n block generated in step 352.
[0063] The first computer system 102 determines in step 356 whether all bits of the information have been encoded the minimum number of times (z). If the result of this determination is "no," then the first computer system 102 returns to identify the next n×n block in the vector for embedding the information. The minimum number of times (z) may be an odd number. Encoding bits of the information (or each bit of the information) a minimum number of z times allows for error correction or a checksum to be performed. Using an odd number of times allows for a simple majority voting technique to be used at a decoder to decode the information. The embedding substage 136 and the inverse frequency transform substage 138 are performed for each set of m bits the minimum number of times (z).For example, if the information output by the information processing pipeline 116 is 160 bits (or twenty sets of m = 8 bits) and the minimum number of times is three (z = 3), then the embedding substage 136 and the inverse frequency transform substage 138 are operated sixty (or 20 * 3) times. In this case, the first computer system 102 identifies sixty n×n blocks (in step 332). The first computer system 102 thus encodes the same set of eight bits in three different n×n blocks.
[0064] If the result of the determination is "Yes," in step 356, the first computer system 102 replaces the Y channel of the first image 106 with the watermarked Y channel in step 358. In the watermarked Y channel, the n×n blocks identified by the first computer system 102 (in step 332) are each replaced by the n×n blocks generated by the first computer system 102 (in step 352).
[0065] In the post-processing stage 130, the first computer system 102 combines the watermarked Y channel with the Cr and Cb channels of the first image 106 in step 360. The first computer system 102 replaces the Y channel of the first image 106 with the watermarked Y channel. The first computer system 102 then converts the YCrCb image to the RGB color space in step 362. The first computer system 102 then rotates the RGB image in step 364 to generate the watermarked image 108. The rotation can be performed at the same angle and in the opposite direction as the rotation performed during the image rotation adjustment stage 118. In the post-processing stage, the first computer system 102 thus reverses the pre-processing performed in the image processing pipeline 114.In the post-processing stage, the first computer system 102 causes the watermarked image 108 to have the same orientation and color space as the first image 106. Referring again to . Fig. 4B shows the combination of the watermarked Y channel with the Cr and Cb channels of the first image 106 and the image rotation performed after the coding stage 128.
[0066] Fig. 6 shows the steps for extracting watermark information from the second image 110. As described herein, the watermarked image 108 may be disseminated or distributed. During dissemination or distribution, the watermarked image 108 may be modified (for example, by cropping or compressing). Such a modification may result in a second image 110. Alternatively, the second image 110 may be the same image as the watermarked image 108 that was shared or transmitted.
[0067] The stages for extracting watermark information may be performed by the second computer system 104. The stages include an image processing pipeline 202 and a decoding stage 204. The image processing pipeline 202 includes an image rotation adjustment stage 206, a color space conversion stage 208, and a channel extraction stage 210. The decoding stage 204 includes a block selection substage 212, a frequency transformation substage 214, an extraction substage 216, and a decoding substage 218.
[0068] The image processing pipeline 202 is similar to the image processing pipeline 114 described herein with reference to the Fig. 2 and Fig. 3A - Fig. 3D. The second computer system 104 processes the second image 110 through the image processing pipeline 202 to obtain a vector of n×n bit blocks for a Y channel of the second image 110.
[0069] In the decoding stage 204, the second computer system 104 processes the vector of n×n bit blocks for the Y channel. The block selection substage 212 and the frequency transformation substage 214 of the decoding stage 204 are similar to the block selection substage 132 and the frequency transformation substage 134 of the encoding stage 128, respectively, which are described herein with reference to Fig. 2 and Fig. 3A - Fig. 3D. The extraction substage 216 and the decoding substage 218 are used by the second computer system 104 to decode and extract encoded information.
[0070] Fig. 7A and Fig. 7B shows a flow diagram of a method 700 for extracting information from the second image 110. In the block selection substage 212, the second computer system 104 receives the vector of n×n bit blocks for the Y channel of the second image 110 at step 702. The second computer system 104 starts a PRNG with a known seed value at 704. The seed value may be the same as the seed value used by the first computer system 102 to perform block selection. If a different block selection technique is used by the first computer system 102 instead, the second computer system 104 may use the same technique.
[0071] The second computer system 104 uses the PRNG at step 706 to identify a next n×n block of the vector and extract information embedded therein. In a first iteration, the next n×n block is a first n×n block of the vector to be selected for information extraction. As described herein, instead of randomly embedding information, a block at (or near) a center of the Y-channel may be identified. Subsequent blocks may be identified based on a pattern, such as a spiral pattern, that begins at the center and continues outward to an edge or periphery of the Y-channel.
[0072] In the frequency transformation substage 214, the second computer system 104 performs a frequency transformation on the n×n block to generate a frequency-domain representation of the block. The frequency transformation is the same as used by the first computer system 102 in transforming the block for information embedding. The second computer system 104 performs a three-stage discrete wavelet transformation on the block in step 708, as described herein. The second computer system 104 selects a region of the DWT-transformed block in step 710, where the region may be the LL3 region, as described herein.
[0073] In the extraction substage 216, the second computer system 104 identifies a next pixel in step 708 and determines whether the pixel is encoded with a logical zero or a logical one. The position of the pixels of the n×n block and the order of encoding the pixels performed by the first computer system 102 are known to the second computer system 104 (and determined, for example, by convention). The second computer system 104 identifies a level associated with the pixel and multiplies the level by the robustness factor. If the product of the multiplication is an odd value, then the encoded bit is a logical zero, and if the product of the multiplication is an even value, then the encoded bit is a logical one.
[0074] In response to identifying the encoded bit, the second computer system 104 determines in step 714 whether all m bits have been extracted from the block. If the result of this determination is "no," then the second computer system 104 returns to identify another pixel of the block to extract an encoded bit from. The next pixel may be at a next position in the order of positions according to which the first computer system 102 encodes the data into n×n blocks. If the result of the determination is "yes" and all m bits have been retrieved from the block, then the second computer system 104 adds the m bits to an extracted data vector in step 716.
[0075] Then, in step 718, the second computer system 104 determines whether the data has been extracted the minimum number of times (z). The extracted data can be expected to correspond to the embedded information. As described herein, the size of the information embedded in the watermarked image can be fixed and known to the second computer system 104. The information can include a number of bits (e.g., 128 bits) corresponding to an identifier and a certain number of parity bits for error detection or correction. The information is embedded multiple times by the first computer system (the minimum number of times (z)). The second computer system 104 can then retrieve data corresponding to the product of the number of bits of the information and the minimum number of times (z).
[0076] If at step 718 the result of the determination is "no," then the second computer system 104 returns to use the PRNG to identify a next block from which m more bits are to be retrieved. If the result of the determination is "yes," then the method 700 continues with the decoding substage 218. In the decoding substage, at step 720 the second computer system 104 splits the data vector into a number of vectors (w), where the number of vectors (w) corresponds to the minimum number of times (z). The vectors may each be of equal size. The size corresponds to the size of the encoded information.
[0077] At step 722, the second computer system 104 performs a voting strategy on the vectors to extract a data stream. The voting strategy may be a majority voting strategy. For each index position of the vectors, the second computer system 104 determines the most frequent value of the vectors. The number of vectors (w) and the minimum number of occurrences (z) may be odd numbers (e.g., three or five). The most frequent value is determined by majority voting. For example, if w = z = 3 and the tenth position of a first vector is 0, the tenth position of a second vector is 0, and the tenth position of a third vector is 1, then the most frequent value for the tenth position is 0. Accordingly, the second computer system 104 determines that the tenth position of the data stream is 0.The voting strategy, as well as the use of parity bits and channel coding, are used to correct errors introduced into the encoded information by splitting the watermarked image 108 or converting the watermarked image 108 into the second image 110.
[0078] As described herein, channel coding is performed on the first image 106. The second computer system 104 decodes the data stream at step 724. The second computer system 104 may use a branch decoding trellis (e.g., a maximum likelihood decoder or a Viterbi channel decoder).
[0079] Fig. Figure 8 shows an example of a branch decoding trellis. The trellis is a time-indexed state diagram. Each state transition is associated with an input bit (x i) and corresponds to a forward step in the trellis. A path through the trellis is shown in bold, and solid lines indicate transitions where the input bit (x i ) is 0, and dashed lines indicate transitions where the input bit (x i ) 1. Each branch of the trellis has a branch metric, and a path metric represents the squared Euclidean distance. Paths that diverge and rejoin another path with a smaller path metric are systematically eliminated (or pruned). The second computer system 104 maintains a minimal path.
[0080] When a next bit is sampled, path metrics are calculated for the two paths leaving each state at a previous sample time by adding branch metrics to previous state metrics. Two path metrics entering a respective state at the current sample time are compared, and the path with the minimum metric is selected as a survivor path. The second computer system 104 processes all bits of the data stream as input bits to determine a most likely (maximum likelihood) path corresponding to a most likely binary string.
[0081] With further reference to the Fig. 7A and Fig. 7B, the second computer system 104 decompresses the binary string resulting from the decoding of the data stream at step 726. As described herein with reference to Fig. 3A - Fig. 3D, the first computer system 102 compresses the information prior to channel encoding. Accordingly, the second computer system 104 reverses the compression. The second computer system 104 may generate a binary string having a UTF-8 format. The second computer system 104 prepends a DID header to the binary string at step 728. The DID header may be the same DID header that the first computer system 102 trimmed during encoding. The second computer system 104 determines a DID embedded in the second image 110. Consequently, the second computer system 104 determines whether the second image 110 contains a watermark corresponding to a work or asset. The determination may be used to prevent misuse of the work or asset and to link the second image 110 to the work or asset.An artist or creator can thus obtain strong evidence that the second image 110 is their work. This determination can be used by the artist or creator, among other things, for an infringement lawsuit. It is noted that the techniques described herein can also be used to provide or detect watermarks in videos and individual images therein.
[0082] Fig. Figure 9 shows a block diagram of a computer system 900. The first computer system 102 and the second computer system 104 described herein may be configured similarly to the computer system 800. In various embodiments, the computer system 900 may include one or more server computer systems, cloud computing platforms or virtual machines, desktop computer systems, laptop computer systems, netbooks, mobile phones, personal digital assistants, televisions, cameras, automotive computers, electronic media players, etc.
[0083] In various embodiments, computer system 900 includes a processor 901 or central processing unit ("CPU") for executing computer programs or executable instructions. First computer system 102 may use processor 901 to perform processing through image and information processing pipelines 114, 116 and encoding stage 128 described herein. Second computer system 104 may use processor 901 to perform processing through image processing pipeline 202 and decoding stage 204 described herein.
[0084] Computer system 900 includes computer memory 902 for storing the programs or executable instructions and data. The first and second computer systems 102, 104 may each store executable instructions representing the pipelined processing described herein. The first and second computer systems 102, 104 may each store executable instructions representing the encoding and decoding stages described herein. Processor 901 of first computer system 102 may execute the executable instructions to perform the operations described herein with respect to first computer system 102. Processor 901 of second computer system 104 may execute the executable instructions to perform the operations described herein with respect to second computer system 104.
[0085] Computer memory 902 stores an operating system including a kernel and device drivers. Computer system 900 includes a persistent storage device 903, such as a hard disk or flash drive, for persistently storing programs and data, a computer-readable media drive 904, such as a floppy disk, CD-ROM, or DVD drive, for reading programs and data stored on a computer-readable medium, and a network connection 905 for connecting computer system 900 to other computer systems to send and / or receive data, such as via the Internet or other network and its network hardware, such as switches, routers, repeaters, electrical cables and optical fibers, light transmitters and receivers, radio transmitters and receivers, and the like. Network connection 905 may be wired or wireless.The first computer system 102 may output the watermarked image 108 for distribution via the network connection 905. The second computer system 104 may receive the second image 110 via the network connection 905 and determine whether information (or a specific DID) is embedded in the second image 110.
[0086] The various embodiments described above may be combined to provide further embodiments. These and other changes may be made to the embodiments in light of the above detailed description. In general, in the following claims, the terms used should not be construed to limit the claims to the specific embodiments disclosed in the specification and claims, but should be construed to include all possible embodiments, along with the full scope of equivalents appropriate to those claims. Accordingly, the claims are not limited by the disclosure.
[0087] This application claims priority to U.S. Non-Provisional Patent Application No. 17 / 857,875, filed July 5, 2022, the entirety of which is incorporated herein by reference.
Claims
[1] Computer system comprising: a processor; and a memory having stored thereon executable instructions which, when executed by the processor, cause the processor to: Receiving a first image and receiving information for embedding in the first image; Preprocessing the first image to generate a preprocessed first image; Preprocessing the information by at least channel coding the information to generate preprocessed information including a plurality of sets of bits; Embedding the preprocessed information into the preprocessed first image by at least: Selecting multiple blocks of the preprocessed first image; and Embedding a respective set of bits of the plurality of sets of bits in each block of the plurality of blocks, wherein each set of bits of the plurality of sets of bits is embedded in a minimum number of blocks of the plurality of blocks, wherein the minimum number of blocks is greater than one. [2] The computer system of claim 1, wherein the executable instructions cause the processor to preprocess the first image by: Rotate the first image by 90 degrees, 180 degrees or 270 degrees. [3] The computer system of claim 1, wherein the executable instructions cause the processor to preprocess the first image by: performing feature detection on the first image to generate a plurality of feature vectors; Determining an average feature vector of the plurality of feature vectors; and Rotating the first image to a quadrant that minimizes an angle between the average feature vector and a central axis of the first image. [4] The computer system of claim 1, wherein the executable instructions cause the processor to preprocess the first image by: Obtaining the preprocessed first image as a luma channel component of the first image. [5] Computer system according to claim 1, where the executable instructions cause the processor to: Frequency transforming the plurality of blocks to generate a plurality of frequency transformed blocks, and wherein the executable instructions cause the processor to embed in each block of the plurality of blocks a respective set of bits of the plurality of sets of bits by: Embedding the respective set of bits of the plurality of sets of bits into each frequency-transformed block of the plurality of frequency-transformed blocks; and in response to embedding the plurality of sets of bits, inversely transforming the plurality of frequency transformed blocks into a spatial domain. [6] The computer system of claim 1, wherein the executable instructions cause the processor to: Embedding a first set of bits of the plurality of sets of bits into a first block of the plurality of blocks by at least: selecting a first bit of the first set of bits; Selecting a pixel position in the first block; identifying a level of pixel position; and Quantizing the level of the pixel position depending on a logical state of the first bit. [7] The computer system of claim 6, wherein the executable instructions that cause the processor to quantize the level of the pixel position in dependence on the logic state of the first bit cause the processor to: if it is determined that the logical state of the first bit is a logical zero, quantizing the level to one of an odd value and an even value; and if it is determined that the logical state of the first bit is a logical one, quantizing the level to the other of the odd value and the even value. [8] The computer system of claim 7, wherein the executable instructions cause the processor to: Dividing the quantized level of the pixel position by a robustness factor. [9] The computer system of claim 1, wherein the executable instructions cause the processor to: Generating a watermarked image based on embedding the preprocessed information in the preprocessed first image; and Arrange for the watermarked image to be output for distribution. [10] The computer system of claim 1, wherein the information is a decentralized identifier defined by the World Wide Web Consortium (W3C) Proposed Recommendation 1.
0. [11] The computer system of claim 1, wherein the executable instructions cause the processor to preprocess the first image by: Setting up an image processing pipeline that includes multiple image preprocessing stages; Selecting one or more of the plurality of image preprocessing stages depending on contextual information associated with the first image, a property of the first image, an image type associated with the first image, or a preprocessing configuration; and Preprocessing the first image by the selected one or more of the plurality of image preprocessing stages to generate the preprocessed first image. [12] The computer system of claim 1, wherein the executable instructions cause the processor to preprocess the information by: Setting up an information processing pipeline that includes multiple information preprocessing stages; Selecting one or more of the plurality of information preprocessing stages depending on the type of information or a preprocessing configuration; and Preprocessing the information by the selected one or more of the plurality of information preprocessing stages to produce the preprocessed information. [13] Method comprising: Receiving a first image and receiving information for embedding in the first image; Preprocessing the first image to generate a preprocessed first image; Preprocessing the information by at least channel coding the information to generate preprocessed information including a plurality of sets of bits; Embedding the preprocessed information into the preprocessed first image by at least: Selecting multiple blocks of the preprocessed first image; and Embedding a respective set of bits of the plurality of sets of bits in each block of the plurality of blocks, wherein each set of bits of the plurality of sets of bits is embedded in a minimum number of blocks of the plurality of blocks, wherein the minimum number of blocks is greater than one. [14] The method of claim 13, wherein preprocessing the first image includes: performing feature detection on the first image to generate a plurality of feature vectors; Determining an average feature vector of the plurality of feature vectors; and Rotating the first image to a quadrant that minimizes an angle between the average feature vector and a central axis of the first image. [15] The method of claim 13, wherein preprocessing the first image includes: Obtaining the preprocessed first image as a luma channel component of the first image. [16] Method according to claim 13, comprising: Embedding a first set of bits of the plurality of sets of bits into a first block of the plurality of blocks by at least: selecting a first bit of the first set of bits; Selecting a pixel position in the first block; identifying a level of pixel position; and Quantizing the level of the pixel position depending on a logical state of the first bit. [17] The method of claim 16, wherein quantizing the level of the pixel position in dependence on the logic state of the first bit comprises: if it is determined that the logical state of the first bit is a logical zero, quantizing the level to one of an odd value and an even value; and if it is determined that the logical state of the first bit is a logical one, quantizing the level to the other of the odd value and the even value. [18] Method according to claim 17, comprising: Dividing the quantized level of the pixel position by a robustness factor. [19] The method of claim 13, wherein preprocessing the first image includes: Setting up an image processing pipeline that includes multiple image preprocessing stages; Selecting one or more of the plurality of image preprocessing stages depending on contextual information associated with the first image, a property of the first image, an image type associated with the first image, or a preprocessing configuration; and Preprocessing the first image by the selected one or more of the plurality of image preprocessing stages to generate the preprocessed first image. [20] The method of claim 13, wherein preprocessing the information includes: Setting up an information processing pipeline that includes multiple information preprocessing stages; Selecting one or more of the plurality of information preprocessing stages depending on the type of information or a preprocessing configuration; and Preprocessing the information by the selected one or more of the plurality of information preprocessing stages to produce the preprocessed information.