Digital Image Watermarking Processing
The described method embeds and detects watermarks in images using rotation, channel encoding, and frequency conversion to ensure robustness and imperceptibility, addressing unauthorized use and proving image authenticity.
Patent Information
- Application Number
- JP2024571879
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-05
- Filing Date
- 2023-02-09
- Publication Date
- 2025-07-10
AI Technical Summary
Images are often subject to unauthorized use and require protection to prove derivation or copying, especially in digital environments where sharing and publication are frequent, necessitating a method to embed and detect watermarks that are robust and imperceptible.
A computer system processes an image through pipelines to embed a watermark by rotating, channel encoding, and frequency converting it, ensuring it remains imperceptible and robust against modifications like cropping and format conversion, using techniques like discrete wavelet transform and channel encoding.
The embedded watermark allows for proving image authenticity and preventing unauthorized use by remaining visually undetectable and resilient to common image alterations, enabling efficient detection and linking back to the original work.
Smart Images

Figure 2025521435000001_ABST
Abstract
Description
Technical Field
[0001] This application is directed to embedding a watermark in an image and restoring the embedded watermark by decoding the image.
Background Art
[0002] Images are often original works that need to be protected. Images are being shared and published to a wider range of viewers more and more frequently. During sharing and publication, unauthorized use of images occurs. For example, the image may be used in a form not permitted by the producer of the image or the rights owner with respect to the benefits of the image. To prevent unauthorized use, it is often necessary to prove whether the image used was actually derived from or copied from an image to which protection applies.
Summary of the Invention
Means for Solving the Problems
[0003] In one embodiment, a computer system includes a processor and a memory storing executable instructions. When the executable instructions are executed by the processor, the processor is caused to receive a first image, receive information to be embedded in the first image, preprocess the first image to generate a preprocessed first image, preprocess the information by at least channel encoding the information to generate preprocessed information including a plurality of bit sets, and embed the preprocessed information in the preprocessed first image by at least selecting a plurality of blocks of the preprocessed first image and embedding each bit set of the plurality of bit sets in each block of the plurality of blocks. In one embodiment, each bit set of the plurality of bit sets is embedded in a minimum number of blocks of the plurality of blocks, and the minimum number of blocks is two or more.
[0004] In one embodiment, the executable instructions cause the processor to preprocess the first image by rotating the first image by 90 degrees, 180 degrees, or 270 degrees. In one embodiment, the executable instructions cause the processor to perform feature detection on the first image to generate a plurality of feature vectors, determine an average feature vector of the plurality of feature vectors, and rotate the first image to a quadrant that minimizes the angle between the average feature vector and the central axis of the first image, thereby preprocessing the first image.
[0005] In one embodiment, the executable instructions cause the processor to preprocess the first image by obtaining the preprocessed first image as a luminance channel component of the first image. In one embodiment, the executable instructions cause the processor to perform frequency conversion on a plurality of blocks to generate a plurality of frequency-converted blocks. In one embodiment, the executable instructions cause the processor to embed each of a plurality of bit sets into each of the frequency-converted blocks of the plurality of frequency-converted blocks, and inverse-transform the plurality of frequency-converted blocks into the spatial domain in response to embedding the plurality of bit sets, thereby embedding each of the plurality of bit sets into each of the plurality of blocks.
[0006] In one embodiment, the executable instructions cause the processor to embed in a first block of a plurality of blocks a first bit set of a plurality of bit sets by at least selecting a first bit of the first bit set, selecting a pixel position in the first block, identifying a level of the pixel position, and quantizing the level of the pixel position according to a logical state of the first bit. In one embodiment, the executable instructions that cause the processor to quantize the level of the pixel position according to the logical state of the first bit cause the processor to quantize the level to one of an odd value or an even value in response to determining that the logical state of the first bit is logical 0, and to quantize the level to the other of the odd value or the even value in response to determining that the logical state of the first bit is logical 1.
[0007] In one embodiment, the executable instructions cause the processor to divide the quantized level of the pixel position by a robustness factor. In one embodiment, the executable instructions cause the processor to generate a watermarked image based on embedding preprocessed information in a preprocessed first image and to output the watermarked image for distribution. In one embodiment, the information is a decentralized identifier defined by World Wide Web Consortium (W3C) Recommendation 1.0.
[0008] In one embodiment, the executable instructions cause the processor to preprocess a first image by configuring an image processing pipeline that includes a plurality of image preprocessing stages, selecting one or more of the plurality of image preprocessing stages according to context information associated with the first image, characteristics of the first image, an image type associated with the first image, or a preprocessing configuration, and preprocessing the first image through the selected one or more of the plurality of image preprocessing stages to generate a preprocessed first image.
[0009] In one embodiment, the executable instructions cause the processor to configure an information processing pipeline including a plurality of information preprocessing stages, select one or more of the plurality of information preprocessing stages according to the type of information or the preprocessing configuration, and preprocess the information by preprocessing the information through the selected one or more of the plurality of information preprocessing stages to generate preprocessed information.
[0010] In one embodiment, the method includes receiving a first image and receiving information to be embedded in the first image, preprocessing the first image to generate a preprocessed first image, preprocessing the information by at least channel encoding the information to generate preprocessed information including a plurality of bit sets, and embedding the preprocessed information in the preprocessed first image by at least selecting a plurality of blocks of the preprocessed first image and embedding each bit set of the plurality of bit sets in each block of the plurality of blocks. In one embodiment, each bit set of the plurality of bit sets is embedded in a minimum number of blocks of the plurality of blocks, and the minimum number of blocks is two or more.
[0011] In one embodiment, preprocessing the first image includes performing feature detection on the first image to generate a plurality of feature vectors, determining an average feature vector of the plurality of feature vectors, and rotating the first image to a quadrant that minimizes the angle between the average feature vector and the central axis of the first image. In one embodiment, preprocessing the first image includes obtaining the preprocessed first image as a luminance channel component of the first image. In one embodiment, the method includes embedding a first bit set of the plurality of bit sets in a first block of the plurality of blocks by at least selecting a first bit of the first bit set, selecting a pixel position in the first block, specifying a level of the pixel position, and quantizing the level of the pixel position according to the logical state of the first bit.
[0012] In one embodiment, quantizing the level of a pixel position according to the logical state of a first bit includes quantizing the level to one of an odd value or an even value in response to determining that the logical state of the first bit is logical 0, and quantizing the level to the other of the odd value or the even value in response to determining that the logical state of the first bit is logical 1. In one embodiment, the method includes dividing the quantized level of the pixel position by a robustness factor.
[0013] In one embodiment, preprocessing a first image includes constructing an image processing pipeline including a plurality of image preprocessing stages, selecting one or more of the plurality of image preprocessing stages according to context information associated with the first image, characteristics of the first image, an image type associated with the first image, or a preprocessing configuration, and preprocessing the first image through the selected one or more of the plurality of image preprocessing stages to generate a preprocessed first image.
[0014] In one embodiment, preprocessing information includes constructing an information processing pipeline including a plurality of information preprocessing stages, selecting one or more of the plurality of information preprocessing stages according to the type of information or a preprocessing configuration, and preprocessing the information through the selected one or more of the plurality of information preprocessing stages to generate preprocessed information.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7A
Figure 7B
Figure 8
Figure 9
BEST MODE FOR CARRYING OUT THE INVENTION
[0016] In this specification, it is provided to embed identification information into an image that can be an original work or an art work using digital watermarking processing. The identification information can be retrieved so that a link can be created between the image and the record of the identification information even if the image is modified. This technique enables, among other things, artists and the authors of original works to share their works on public platforms. In a situation where an image is used in an improper way (for example, as part of a media campaign without permission), a link can be created between the image and the abuse.
[0017] FIG. 1 shows an example of an environment 100 for digital watermarking an image. The environment 100 includes a first computer system 102 and a second computer system 104. The first computer system 102 receives a first image 106 and generates a watermarked image 108 by processing the first image 106. The first image 106 can be, among other things, an art work or an original work. The first image 106 can be subject to dissemination, distribution, or sharing. However, the owner or copyright holder of the first image 106 may attempt to prevent improper use of the first image 106. For example, the owner may attempt to prevent improper use of the first image 106 in, among other things, media campaigns.
[0018] The first computer system 102 receives a first image 106, processes the first image 106, and acts on the first image 106 to generate a watermarked image 108 from the first image 106. The watermarked image 108 has information or data (a watermark) that can be retrieved to identify the watermarked image 108 embedded therein or is encoded using this information or data. The watermarked image 108 can be disseminated, distributed, or shared. When the watermarked image 108 is shared, the watermarked image 108 can become a second image 110, among other things, through editing, cropping, compression, or format conversion. Also, the watermarked image 108 may not be modified during distribution, and thus the second image 110 may be the same as the watermarked image 108. The second computer system 104 receives the second image 110. The second computer system 104 processes the second image 110 to determine whether the second image 110 is encoded using a watermark, i.e., whether a watermark is embedded therein. For example, the second computer system 104 can decrypt the second image 110 and determine whether a watermark exists in the second image 110.
[0019] If a watermark exists in the second image 110, it can be proven that the second image 110 is based on the first image 108. For example, the second image 110 can be a copy of the watermarked image 108. Alternatively, the watermarked image 108 or a portion thereof may be cropped or undergo a format conversion to generate the second image 110 or a portion of the second image 110.
[0020] The digital watermark is preferably inserted into the watermarked image 108 such that it is not visually perceptible to humans. Visual artifacts associated with the watermark are buried (or mixed with the watermarked image 108) in the watermarked image 108 such that they cannot be visually recognized by humans. The watermark modifies the first image 106 to produce the watermarked image 108, and this modification is required to constitute a feature of the watermarked image 108 other than features on which a human eye focuses or pays attention. Thus, while the watermarked image 108 is required to be perceived by a human viewer as a copy of the first image 106, it actually contains identification information that is not part of the first image 106.
[0021] In addition, the watermark is required to be robust. Therefore, the watermark withstands, among other actions performed to generate the second image 110 from the watermarked image 108, cropping and format conversion. Further, the watermark may be applied using frequency domain or time domain processing or preprocessing.
[0022] Figure 2A shows the stage of watermarking the first image 106. The stage of watermarking the first image 106 can be executed by the first computer system 102 as described herein. The stage includes a preprocessing stage that is executed as part of the preprocessing before encoding the first image 106. The first computer system 102 receives the first image 106 and the information 112 to be embedded in the first image 106 using the watermarking technique.
[0023] The first computer system 102 executes various preprocessing stages that constitute the image processing pipeline 114 on the first image 106. The first computer system 102 also executes various preprocessing stages that constitute the information processing pipeline 116 on the information 112. The image processing pipeline 114 includes an image rotation adjustment stage 118, a color space conversion stage 120, and a channel extraction stage 122. The information processing pipeline 116 includes a verification stage 124 and a compression stage 126. The compression stage 126 described herein includes channel encoding (e.g., convolutional encoding) of the information 112.
[0024] The preprocessed image output by the image processing pipeline 114 and the preprocessed information output by the information processing pipeline 116 are supplied to the encoding stage 128. The encoding stage 128 embeds the preprocessed information into the preprocessed image. The encoding stage 128 generates image data including the preprocessed image with the preprocessed information embedded therein. The encoding stage 128 outputs the image data to the post-processing stage 130. The post-processing stage 130 may execute operations that are "inverse" to the operations executed in the image processing pipeline 114. The post-processing stage 130 may include channel combination, color space reconversion, and rotation adjustment. For example, the rotation adjustment may reverse the image rotation adjustment executed as part of the image rotation adjustment stage 118. The post-processing stage 130 outputs the watermarked image 108.
[0025] Note that the various stages of the image processing pipeline 114 and the information processing pipeline 116 may be removed according to the image and the desired watermark processing. For example, in one embodiment, the image rotation adjustment stage 118 may be omitted, and in that case, the first image 106 may be processed through the image processing pipeline 114 without changing its orientation. The first image 106 may be processed without changing its orientation, i.e., without rotating, for example, upside down.
[0026] Furthermore, stages may be added to the image processing pipeline 114 or the information processing pipeline 116. For example, a feature detection stage or an image normalization stage may be added to the image processing pipeline 114. For example, a feature detection stage may be used by the first computer system 102 to identify the category of the first image 106. The category may be one or more of a plurality of categories including, for example, photographs and line drawings. The encoding technique used in the encoding stage 128 may be changed or adjusted according to the identified category. Furthermore, the stages included in the image processing pipeline 114 may be changed according to the category of the image. In addition, the context information may notify the stages used as part of the image processing pipeline 114. The context information may include the type of software used to generate the first image 106. For example, if the image is generated using a certain type of software, the color space conversion may be omitted.
[0027] The first computer system 102 may configure the image processing pipeline 114. The first computer system 102 may selectively add a plurality of image preprocessing stages to the image processing pipeline 114. The first computer system 102 may determine to add a plurality of image preprocessing stages to the image processing pipeline 114 according to context information associated with the first image (such as the software used to generate the first image), characteristics of the first image (such as the color space of the first image, the size of the first image, or the dimensions of the first image), the image type associated with the first image (whether the first image represents a line drawing or a photograph, or the type of photograph, etc.), or a preprocessing configuration (such as a user configuration of the level of arithmetic density used to perform image preprocessing). For example, when the arithmetic density is selected to be limited, the first computer system 102 can minimize the number of preprocessing stages, and vice versa. The first computer system 102 may then preprocess the first image through a plurality of image preprocessing stages to generate a preprocessed first image.
[0028] In the information processing pipeline 116, if the information 112 does not have a specific format to be verified, the verification stage 124 may not be necessary. As described herein, the verification stage 124 verifies that the information is in the form of a decentralized identifier (DID). However, if information of any format can be added as a watermark to the first image 106, the verification stage 124 may not be used to confirm that the information is a DID. In addition, the compression stage 126 may not be necessary if it is required to embed the information uncompressed.
[0029] The first computer system 102 may configure the information processing pipeline 116 according to a modular approach. The first computer system 102 may selectively add a plurality of information preprocessing stages to the information processing pipeline 116. The first computer system 102 may determine a plurality of information preprocessing stages to add to the information processing pipeline 116 according to the type of information or the configuration of the preprocessing. The first computer system 102 may then generate information for embedding in the first image by preprocessing the information through the plurality of information preprocessing stages. The type of this information may be the format of the information. If the information needs to have a specific format, the first computer system 102 may configure the information processing pipeline 116 with a preprocessing stage for verifying the information.
[0030] Pipeline-based techniques are advantageous in that watermark encoding and decoding can adapt to changing requirements and environments. Various preprocessing stages and postprocessing stages may be added or removed to accommodate various image types, information types, and / or encoding or decoding requirements.
[0031] FIG. 2B shows the sub-stages of the encoding stage 128. The encoding stage 128 includes a block selection sub-stage 132, a frequency conversion sub-stage 134, an embedding sub-stage 136, and an inverse frequency conversion sub-stage 138.
[0032] The block selection sub-stage 132 may select a block of the pre-processed image into which the pre-processed information or a portion of the pre-processed information is to be embedded. The frequency conversion sub-stage 134 may convert the selected block from the spatial domain to the frequency domain. The embedding sub-stage 136 may embed the pre-processed information or a portion of the pre-processed information into the frequency domain representation of the selected block. The inverse frequency conversion sub-stage 138 may convert the frequency domain representation of the selected block back to the spatial domain. As described herein, the spatial domain representation is supplied as image data to a post-processing stage 130 for generating the watermarked image 108.
[0033] Referring to FIGS. 3A-3D, a flowchart of a method 300 is shown that includes pre-processing of a first image 106 through an image processing pipeline 114, pre-processing of information 112 through an information processing pipeline 116, and operations of an encoding stage 128 that encodes the first pre-processed image using the pre-processed information.
[0034] The first computer system 102 receives, at 302, a first image 106 and information 112. The first computer system 102 transfers, at 304, the first image 106 to the image processing pipeline 114 and the information 112 to the information processing pipeline 116. In the image processing pipeline 114, the first computer system 102 pre-processes the first image 106 through an image rotation adjustment stage 118. In the image rotation adjustment stage 118, the first computer system 102 performs a rotation on the first image 106.
[0035] The preprocessing of the first image 106 in the image rotation adjustment stage 118 is performed to identify a local vertical axis associated with the image and rotate the image so that the orientation of the image is reversed (or upside down, or the image is rotated 180 degrees). The image rotation adjustment stage 118 advantageously standardizes the rotation of the image for watermark encoding and detection. The image rotation adjustment stage 118 adds resistance to rotation attacks that can be used for the adjustment of rotation to slip through watermark detection. The image rotation adjustment stage 118 uses the features of the first image 106 to identify a local vertical axis that is unique with respect to the image. This local vertical axis is used to specify the orientation of the first image 106 in subsequent preprocessing and encoding.
[0036] As part of the image rotation adjustment stage 118, the first computer system 102 loads the first image 106 in the red, green, and blue (RGB) color space at 306. The first computer system 102 performs feature detection on the first image at 308 to obtain a feature vector of the image. Feature detection associates each feature or element of several features or elements with a feature vector having a magnitude and a direction. The first computer system 102 may utilize any feature detection technique (for determining the feature vector), such as SURF (speeded up robust features), SIFT (scale - invariant feature transform), FAST (features from accelerated segment test), for example.
[0037] The first computer system 102 determines, at 310, the average vector of the feature vectors of the image. The first computer system 102 rotates, at 312, the image so that the average vector is along the central axis of the image. The rotation can be performed in 90° increments until the average vector is at the position closest to the central axis of the image. Such rotation can minimize the angle between the average vector and the central axis of the first image 106, and also maintain the rectangular characteristics of the image.
[0038] FIG. 4A shows an example of preprocessing the first image 106 through the image processing pipeline 114. After determining the feature vector of the first image 106, the first computer system 102 determines the average vector of the feature vectors. Next, the first computer system 102 rotates the first image 106 so that the average vector is in the direction of the central axis of the first image 106. The first computer system 102 approximates the rotation with a multiple of 90° that makes the average vector closest to the central axis of the image in order to keep the first image 106 square or rectangular. In FIG. 4A, the first computer system 102 rotates the first image 106 by 180°. However, other options for rotation include 90° and 270°. Further, when the average vector is aligned with the central axis, the rotation is 0° and the image is not rotated. As described here, feature detection sets the rules regarding the orientation of the first image 106 when the first image undergoes the processing described here. By rotating the first image 106, if the watermark is embedded in a state where the orientation of the image is opposite to the natural orientation of the image, it becomes difficult for the human eye to notice the watermark, thus reducing the perceptibility of the watermark.
[0039] Referring back to FIGS. 3A - 3D, the color space conversion stage 120 will be described. In the color space conversion stage 120, the first computer system 102 converts the image from the RGB color space to the YCrCb color space at 314. In the YCrCb color space, Y is the luminance channel component, Cr is the red color difference channel component, and Cb is the blue color difference channel component. The information 112 can be embedded in the luminance channel component (luminance channel) while making a higher degree of change without enhancing the perceptibility of the change to the human eye.
[0040] Note that the information may be embedded in the channels of the RGB color space. In particular, since the green channel has limited perceptibility compared to the red and blue channels, the information may be encoded in the green channel. In this example, for the purpose of embedding information, the RGB - YCrCb color space conversion may not be necessary.
[0041] In the channel extraction stage 122, the first computer system 102 extracts the Y channel from the YCrCb color space at 316. The first computer system 102 divides the Y channel into n×n bit blocks at 318. Here, n can be any number of bits. For example, the block size may be 4×4 bits or 12×12 bits, or even larger. The block size may be 32×32 or 64×64. If the block size is smaller than 4×4 bits or 12×12 bits, embedding the information 112 in the block may modify the characteristics of the block to be perceptible to the human eye. Instead, selecting a relatively large block size (such as 1024×1024) may result in insufficient attack resistance required to enable decoding or extraction of the embedded information. Additionally, although n×n blocks are described herein, the blocks into which the image is divided may be rectangular (e.g., m×n) instead of square.
[0042] FIG. 4B schematically illustrates the generation of the watermarked image 108. As shown in FIG. 4B, the first computer system 102 converts the rotated first image from the RGB color space to the YCrCb color space in the color space conversion stage 120. The first computer system 102 extracts the Y channel from the YCrCb color space and divides the Y channel into n×n bit blocks in the channel extraction stage 122.
[0043] Referring back to FIGS. 3A-3D, the verification stage 124 and the compression stage 126 of the information processing pipeline 116 are configured to preprocess a decentralized identifier (DID). A DID is a globally unique persistent identifier. An example of a DID is defined by the World Wide Web Consortium (W3C) Recommendation 1.0. A DID may have the form did:example:12345abde. The "did" in the first part of the DID is an identifier in the Uniform Resource Identifier (URI) scheme. The "example" in the second part of the DID identifies the "DID method", which is a particular scheme used to generate a particular identifier. The third part of the DID is a particular identifier associated with the "DID method". A DID may be generated under Wacom's Creative Rights Initiative (as the "DID method") and may have the form did:cri:12345abde.
[0044] At 320 within the verification stage 124, the first computer system 102 determines whether the information is a valid DID. Determining whether the information is a valid DID may include determining whether the information has the form of a DID and / or whether the third part (the particular identifier) is a valid identifier (e.g., of a creative work that is to be watermarked). If a negative determination is made, the first computer system 102 declares at 322 that the information does not have a supported form. Thereby, the preprocessing ends.
[0045] When a positive determination is made, the first computer system 102 preprocesses the information by the compression stage 126. In 324 within the compression stage 126, the first computer system 102 deletes the first and second portions of the information. The first computer system 102 maintains the third portion (a specific identifier) of the information. In 326, the first computer system 102 compresses the information. For example, if the third portion (the specific identifier) has the form of UTF-8 (Unicode Transformation Format - 8), the first computer system 102 may convert each character to 4 bits. In UTF-8, each character can be represented by 1 to 4 bytes. The conversion to 4 bits results in reducing the representation to 1 / 8 to 1 / 2. The third portion may be 32 characters, which, when converted to 4 bits per character, produces an output of 128 bits.
[0046] In 328, the first computer system 102 encodes the information using a channel encoder. The channel encoder can be a convolutional encoder. Channel encoding can add parity bits or a checksum to the information to enable error recovery. Since the convolutional encoder slides, a time-invariant trellis decoder can be used on the decoding side. The trellis decoder enables maximum likelihood decoding.
[0047] In the symbolization stage 128, the first computer system 102 performs block selection on a vector of n×n-bit Y-channel blocks as part of the block selection sub-stage 132. The first computer system 102 starts a pseudo random number generator (PRNG) using a known seed at 330. The seed may be known to the second computer system 104. Thus, at the time of decoding, the second computer system 104 uses the same seed to identify blocks in which information or a portion of the information has been watermarked for information recovery. When initialized with the same value, the PRNG generates the same sequence of numbers in each iteration.
[0048] The first computer system 102 uses the PRNG at 332 to identify the next n×n block from the vector for embedding information therein. In the first iteration, the next n×n block is the first n×n block of the vector and is selected for embedding information. The use of the PRNG distributes and randomizes the watermark across the Y-channel of the first image 106.
[0049] Note that the use of the PRNG may not be necessary. Instead, a block at the center of (or near) the Y-channel is identified and used for embedding information therein. Subsequent blocks may be identified based on a pattern such as a spiral pattern that moves outward from the center towards the edges or periphery of the Y-channel. Prioritizing blocks that are at or near the spatial center of the image is advantageous in that these blocks are less likely to be exposed to cropping. These blocks are more likely to be maintained when the image is distributed or shared and are thus more likely to be available for subsequent watermark detection.
[0050] Furthermore, feature-based identification may be used to identify the blocks to be encoded using the information. The n×n blocks of the Y channel may be arranged according to the strength of the features within the blocks. The block having the highest feature strength may be selected first. The selection of subsequent blocks may proceed in descending order of feature strength.
[0051] Alternatively, feature-based identification may be used to identify one block having the maximum feature strength (determined by feature detection). The block having the maximum feature strength is considered to be the center of the image and may be selected for encoding. The selection of subsequent blocks may be performed in a spiral or circular shape away from the first block and towards the boundary of the Y channel of the image.
[0052] In the frequency conversion sub-stage 134, the first computer system 102 performs frequency conversion on the n×n blocks to generate a frequency domain representation of the blocks. Although discrete wavelet transform (DWT) is described here, the techniques described here may be used for any type of frequency domain conversion. The first computer system 102 performs a 3-level discrete wavelet transform on the blocks in 334. The first computer system 102 continuously performs a 3-level DWT.
[0053] FIG. 5 shows an example of 3-level DWT. In a first operation, the first computer system 102 performs DWT on block 502. The first computer system 102 then selects a first region 504 of the block transformed by DWT. The first region 504 can be a region (denoted as "LL1") that includes the low-frequency parts of two dimensions of the block transformed by DWT. The block transformed by DWT can include an LL region, an HH region that includes high-frequency parts of two dimensions, an LH region that includes the low-frequency part of the first dimension and the high-frequency part of the second dimension, and an HL region that includes the high-frequency part of the first dimension and the low-frequency part of the second dimension. In a second operation, the first computer system 102 performs DWT on the first region 504. The first computer system 102 then selects a second region 506 ("LL2") of the block transformed twice. The second region 506 can also be an LL region. In a third operation, the first computer system 102 performs DWT on the second region 506. The first computer system 102 selects a third region 508 ("LL3") that is the LL region of the transformed block. The selection of the third region 508 for information encoding is more resistant to compression attacks. This is due to the fact that the compression algorithm utilizes the third region 508 to compress the image data. Therefore, the placement of the encoded information or watermark in the third region 508 makes the watermark more resistant to compression.
[0054] Referring back to FIGS. 3A - 3D, the first computer system 102 selects, at 336, a region of blocks transformed by the DWT. The first computer system 102 selects, at 338, the next m - bit data from the information. For each n×n block, the first computer system 102 may select, for example, if m can be 8, the m - bit data to be encoded in the block. The information may be supplied from the information processing pipeline 116 or the compression stage 126 of the information processing pipeline 116 as described herein. At the encoding stage 128, the first computer system 102 encodes the region using the m - bit data.
[0055] Note that, as described herein, each set of m - bit data may be encoded at least z times. The first computer system 102, at 338, tracks the number of times a set of m - bits is selected and may determine whether the set of m - bits has been selected the minimum number of times (z). The first computer system 102 may select a different set of m - bits in response to determining that a previously selected set of m - bits has been encoded for a specified minimum number of times (z).
[0056] At 340 within the encoding stage 128, the first computer system 102 identifies the pixels of the selected region (of the blocks transformed by the DWT) and the pixel value (Q). The first computer system 102, at 342, obtains the next bit (b) of the m data bits. The pixel positions used to encode the m data bits may be determined in advance or specified in advance. For any region, the first computer system 102 may be set with the pixel positions to select when encoding the m data bits. The first computer system 102, at 344, determines whether the logical value of the bit (b) among the m data bits is "0" or "1". The first computer system 102 then changes the pixel value (Q) according to the state of the bit (b).
[0057] The first computer system 102 is configured to have a robustness coefficient. The robustness coefficient sets the relationship between robustness and imperceptibility. The robustness coefficient is negatively correlated with imperceptibility. A higher robustness coefficient results in a more robust watermarking process but increases the perceptibility of the watermark.
[0058] In response to determining that the logical value of bit (b) among the m data bits is "1", the first computer system 102 quantizes the value (Q) of the selected region to an even value (e.g., the nearest even value, or the rounded-up or rounded-down even value) at 346. The first computer system 102 also divides this even value by the robustness coefficient to obtain a replacement value (Q'). The first computer system 102 replaces the value (Q) of the selected region with the replacement value (Q').
[0059] In response to determining that the logical value of bit (b) among the m data bits is "0", the first computer system 102 quantizes the value (Q) of the selected region to an odd value (e.g., the nearest odd value, or the rounded-up or rounded-down odd value) at 346. The first computer system 102 also divides this odd value by the robustness coefficient to obtain a replacement value (Q'). The first computer system 102 replaces the value (Q) of the selected region with the replacement value (Q').
[0060] The bit (b) is encoded by changing one pixel of the region converted into the frequency domain according to the logical state (0 or 1) of the bit (b). That a certain pixel of the converted region is set to the ratio of the quantization result to the robustness coefficient for an even value indicates that the logical value of the bit (b) is "1". That a certain pixel of the converted region is set to the ratio of the quantization result to the robustness coefficient for an odd value indicates that the logical value of the bit (b) is "0". Note that this rule may be reversed, and quantization to an even value may be used to indicate a logical value of "0", or quantization to an odd value may be used to indicate a logical value of "1". Quantization to odd and even values has durability and is difficult to notice because it has a small footprint and does not dramatically modify the characteristics of the block.
[0061] At 350, the first computer system 102 determines whether the encoding (encoding using robustness coefficient quantization) of m-bit data selected from the information is completed. In response to a negative determination, the first computer system 102 returns to obtaining the next pixel and the next value (Q) of the next pixel from the selected region of the block transformed by the DWT, and obtaining the next bit (b) to be encoded using the next value (Q). If the first computer system 102 encodes all of the m selected bits at their respective m pixel positions using the even and odd quantizations described herein, it can make an affirmative determination at 350.
[0062] In response to making an affirmative determination, at 352, the first computer system 102 performs an inverse 3-level DWT on the selected region to convert the selected region back into an n×n block in spatial form. At 352, the first computer system 102 replaces the n×n block (identified at 332) in the vector with the watermarked version of the n×n block generated at 352.
[0063] In 356, the first computer system 102 determines whether all bits of the information have been encoded for a minimum number of times (z). If a negative determination is made, the first computer system 102 returns to identifying the next n×n block in the vector in order to embed the information. The minimum number of times (z) may be odd. By encoding the bits of the information (or each bit of the information) at least z times, it becomes possible to perform error correction or checksum. By using an odd number of times, it becomes possible to decode the information using a simple majority voting technique on the decoding side. The embedding sub-stage 136 and the inverse frequency conversion sub-stage 138 are executed for a minimum number of times (z) for each set of m bits. For example, if the information output by the information processing pipeline 116 is 160 bits (i.e., 20 sets of m = 8 bits) and the minimum number of times is 3 (z = 3), the embedding sub-stage 136 and the inverse frequency conversion sub-stage 138 operate 60 times (or 20×3). In this case, the first computer system 102 identifies 60 n×n blocks (in 332). The first computer system 102 encodes the same set of 8 bits in three different n×n blocks.
[0064] If an affirmative determination is made in 356, the first computer system 102 replaces the Y channel of the first image 106 with the watermarked Y channel in 358. In the watermarked Y channel, the n×n blocks identified by the first computer system 102 in (332) are each replaced with the n×n blocks generated by the first computer system 102 in (352).
[0065] In 360 within the post - processing stage 130, the first computer system 102 combines the water - marked Y channel with the Cr channel and the Cb channel of the first image 106. The first computer system 102 replaces the Y channel of the first image 106 with the water - marked Y channel. In 362, the first computer system 102 converts the YCrCb image to the RGB color space. In 364, the first computer system 102 rotates the RGB image to generate the water - marked image 108. The rotation can have the same angle and the opposite direction as the rotation executed during the image rotation adjustment stage 118. In the post - processing stage, the first computer system 102 reverses the pre - processing executed during the image processing pipeline 114. In the post - processing stage, the first computer system 102 makes the water - marked image 108 have the same orientation and color space as the first image 106. Referring back to FIG. 4B, it is shown that the combination of the water - marked Y channel with the Cr channel and the Cb channel of the first image 106, and the image rotation are executed after the encoding stage 128.
[0066] FIG. 6 shows the stage of extracting water - mark information from the second image 110. As described herein, the water - marked image 108 can be disseminated or distributed. In the dissemination or distribution, the water - marked image 108 can be altered (e.g., by cropping or compression). This alteration may result in the second image 110. Or, the second image 110 may be the same as the shared or transmitted water - marked image 108.
[0067] The stage of extracting watermark information can be executed by the second computer system 104. This stage includes an image processing pipeline 202 and a decoding stage 204. The image processing pipeline 202 includes an image rotation adjustment stage 206, a color space conversion stage 208, and a channel extraction stage 210. The encoding stage 204 includes a block selection sub-stage 212, a frequency conversion sub-stage 214, an extraction sub-stage 216, and a decoding sub-stage 218.
[0068] The image processing pipeline 202 is similar to the image processing pipeline 114 described herein with reference to FIGS. 2 and 3A-3D. The second computer system 104 processes the second image 110 through the image processing pipeline 202 and, for the second image 110, obtains a vector of n×n-bit blocks for the Y channel of the second image 110.
[0069] In the decoding stage 204, the second computer system 104 processes the vector of n×n-bit blocks for the Y channel. The block selection sub-stage 212 and the frequency conversion sub-stage 214 of the decoding stage 204 are similar to the block selection sub-stage 132 and the frequency conversion sub-stage 134 of the encoding stage 128 described herein with reference to FIGS. 2 and 3A-3D, respectively. The extraction sub-stage 216 and the decoding sub-stage 218 are utilized by the second computer system 104 to decode and extract the encoded information.
[0070] Figures 7A and 7B show a flowchart of a method 700 for extracting information from a second image 110. At 702 within the block selection sub-stage 212, the second computer system 104 receives a vector of n×n-bit blocks for the Y channel of the second image 110. The second computer system 104, at 704, starts a PRNG using a known seed. The seed can be the same as the seed used by the first computer system 102 to perform block selection. Alternatively, if the first computer system 102 uses a different block selection technique, the second computer system 104 may also utilize the same technique.
[0071] The second computer system 104, at 706, uses the PRNG to identify the next n×n block of the vector and extract the information embedded in this block. In the first iteration, the next n×n block is the first n×n block of the vector and is selected for information extraction. As described herein, as an alternative to random information embedding, a block at the center (or near the center) of the Y channel may be identified. Subsequent blocks may be identified based on a pattern such as a spiral pattern that moves outward from the center towards the edge or periphery of the Y channel.
[0072] In the frequency conversion sub-stage 214, the second computer system 104 performs a frequency conversion on the n×n block to generate a frequency domain representation of the block. The frequency conversion is the same as that used by the first computer system 102 when converting the block for information embedding. The second computer system 104, at 708, performs a three-level discrete wavelet transform on the block as described herein. The second computer system 104, at 710, selects a region of the block transformed by the DWT. This region can be the LL3 region described above.
[0073] Within extraction sub-stage 216 at 712, the second computer system 104 identifies the next pixel and determines whether this pixel is encoded using a logical value of "1" or a logical value of "0". The position of the pixels in the n×n block and the order of pixel encoding executed by the first computer system 102 are known to the second computer system 104 (and are determined, for example, by rules). The second computer system 104 identifies the level associated with the pixel and multiplies this level by a robustness coefficient. If the product of the multiplication is an odd value, the encoded bit is a logical value of "0", and if the product of the multiplication is an even value, the encoded bit is a logical value of "1".
[0074] In response to identifying the encoded bits, at 714, the second computer system 104 determines whether all m bits have been extracted from the block. If a negative determination is made, the second computer system 104 returns to identifying another pixel in the block and extracting the encoded bits from this pixel. This other pixel can be the next position in the order of positions when the first computer system 102 encodes the data into the n×n block. If a positive determination is made and all m bits have been retrieved from the block, at 716, the second computer system 104 adds those m bits to the extracted data vector.
[0075] Next, at 718, the second computer system 104 determines whether the data has been extracted over the minimum number of times (z). The extracted data is expected to correspond to the embedded information. As described above, the size of the information embedded in the watermarked image is fixed and may be known to the second computer system 104. The information may include a number of bits corresponding to an identifier (e.g., 128 bits) and a certain number of parity bits for error detection or correction. The information is embedded a plurality of times (minimum number of times (z)) by the first computer system. The second computer system 104 may then retrieve data corresponding to the product of the number of bits of the information and the minimum number of times (z).
[0076] If a negative determination is made at 718, the second computer system 104 returns to using the PRNG to identify the next block from which additional m bits are to be retrieved. If an affirmative determination is made, method 700 proceeds to the decoding sub-stage 218. At 720 within the decoding sub-stage, the second computer system 104 divides the data vector into a plurality of vectors (w). Here, the number of vectors (w) corresponds to the minimum number of times (z). These vectors may each be of equal size. The size corresponds to the size of the encoded information.
[0077] At 722, the second computer system 104 extracts a data stream by executing a voting strategy on the vectors. The voting strategy can be a majority voting strategy. For each index position of the vectors, the second computer system 104 determines the most common vector value. The number of vectors (w) and the minimum number (z) can be odd numbers (e.g., 3 or 5). The most common value is determined by majority voting. For example, if w = z = 3, the 10th position of the first vector is 0, the 10th position of the second vector is 0, and the 10th position of the third vector is 1, then the most common value at the 10th position is 0. Thus, the second computer system 104 determines that the 10th position of the data stream is 0. The voting strategy is used to correct errors introduced into the encoded information through the sharing of the watermarked image 108 or the conversion of the watermarked image 108 to a second image 110, similar to the use of parity bits and channel coding.
[0078] As previously explained, channel coding is performed on the first image 106. At 724, the second computer system 104 decodes the data stream. The second computer system 104 can use a branch decoding trellis (e.g., a maximum likelihood decoder or a Viterbi channel decoder).
[0079] Figure 8 shows an example of a branch decoding trellis. The trellis is a state diagram with a time index. Each state transition is associated with an input bit (x i ) and corresponds to a forward step in the trellis. The paths in the trellis are shown thickly, with solid lines indicating transitions where the input bit (x i ) is 0, and dashed lines indicating transitions where the input bit (x i ) is 1. Each branch of the trellis has a branch metric, and the path metric represents the square of the Euclidean distance. Paths that branch and rejoin another path with a smaller path metric are systematically eliminated (or pruned). The second computer system 104 maintains the minimum path.
[0080] When the next bit is sampled, the path metrics of the two paths from each state at the previous sampling time are calculated by adding the branch metric to the previous state metric. At the current sampling time, the two path metrics entering each state are compared, and the path with the minimum metric is selected as the surviving path. The second computer system 104 processes all the bits of the data stream as input bits to determine the most likely (maximum likelihood) path corresponding to the most likely binary sequence.
[0081] Referring again to FIGS. 7A and 7B, the second computer system 104 unfolds the binary sequence obtained as a result of decoding the data stream at 726. As described with reference to FIGS. 3A to 3D, the first computer system 102 compresses information before channel encoding. Therefore, the second computer system 104 reverses that compression. The second computer system 104 may generate a binary sequence in UTF-8 format. At 728, the second computer system 104 adds a DID header to the beginning of the binary sequence. The DID header may be the same DID header that the first computer system 102 trimmed during encoding. The second computer system 104 determines the DID embedded in the second image 110. And based on the result, the second computer system 104 determines whether the second image 110 contains a watermark corresponding to a work or an asset. This determination is used to prevent abuse of the work or asset and to link the second image 110 to the work or asset. The artist or author can have strong evidence that the second image 110 is their work. This determination can be used by the artist or author, especially for infringement claims. Note that the techniques described in this specification can also be used to watermark video and its constituent images and detect the watermarks therein.
[0082] FIG. 9 shows a block diagram of a computer system 900. The first computer system 102 and the second computer system 104 described in this specification may be configured in the same manner as the computer system 900. In various embodiments, the computer system 900 may be one or more server computer systems, a cloud computing platform or virtual machine, a desktop computer system, a laptop computer system, a netbook, a mobile phone, a personal digital assistant, a television, a camera, an in-vehicle computer, an electronic media player, and the like.
[0083] In various embodiments, the computer system 900 includes a processor 901 or a central processing unit ("CPU") that executes a computer program or executable instructions. The first computer system 102 may use the processor 901 to execute the processing through the image processing pipeline 114 and the information processing pipeline 116, and the encoding stage 128 described in this specification. The second computer system 104 may use the processor 901 to execute the processing through the image processing pipeline 202 and the encoding stage 204 described in this specification.
[0084] The computer system 900 includes a computer memory 902 that stores programs or executable instructions and data. The first computer system 102 and the second computer system 104 may each store executable instructions that represent the pipeline processing described herein. The first computer system 102 and the second computer system 104 may each store executable instructions that represent the encoding stage and the decoding stage described herein. The processor 901 of the first computer system 102 may execute the executable instructions to perform the operations described herein associated with the first computer system 102. The processor 901 of the second computer system 104 may execute the executable instructions to perform the operations described herein associated with the second computer system 104.
[0085] The computer memory 902 stores an operating system including a kernel and a device driver. The computer system 900 includes a persistent storage device 903, such as a hard drive or a flash drive, that permanently stores programs and data, a computer-readable medium drive 904, such as a floppy drive, a CD-ROM drive, or a DVD drive, for reading programs and data stored on a computer-readable medium, and a network connection unit 905 that connects the computer system 900 to other computer systems for transmitting and receiving data via the Internet or another network, etc., and connects the computer system 900 to its network connection hardware, such as a switch, a router, a repeater, an electrical cable and an optical fiber, an optical emitter and an optical receiver, a wireless transmitter and a wireless receiver, etc. The network connection unit 905 may be a modem or wireless. The first computer system 102 may output the watermarked image 108 for distribution via the network connection unit 905. The second computer system 104 may receive a second image 110 via the network connection unit 905 and may determine whether information (i.e., a specific DID) is embedded in the second image 110.
[0086] Additional embodiments can be provided by combining the various embodiments described above. In light of the above detailed description, these and other changes can be made to these embodiments. In general, the terms used in the following claims are not to be construed as limiting the claims to the specific embodiments disclosed in the specification and claims, but rather are to be construed as including all possible embodiments together with the full scope of equivalents to which such claims are entitled. Accordingly, the claims are not limited by the present disclosure.
[0087] This application claims the benefit of priority to U.S. Non-Provisional Patent Application No. 17 / 857,875, filed Jul. 5, 2022, the entire disclosure of which is hereby incorporated by reference herein.
Claims
1. A computer system comprising: a processor; and a memory storing executable instructions, wherein when the executable instructions are executed by the processor, the processor is caused to: receive a first image and information for embedding into the first image; generate a preprocessed first image by preprocessing the first image; generate preprocessed information including a plurality of bit sets by preprocessing the information using at least channel encoding of the information; and embed the preprocessed information into the preprocessed first image, wherein embedding the preprocessed information into the preprocessed first image includes at least: selecting a plurality of blocks of the preprocessed first image; and embedding each bit set of the plurality of bit sets into each block of the plurality of blocks, wherein each bit set of the plurality of bit sets is embedded into a minimum number of blocks among the plurality of blocks, and the minimum number is greater than 1, a computer system.
2. The executable instructions cause the processor to preprocess the first image by rotating the first image by 90 degrees, 180 degrees, or 270 degrees. The computer system according to claim 1.
3. The executable instructions: generate a plurality of feature vectors by performing feature detection on the first image; determine an average feature vector of the plurality of feature vectors; and rotate the first image in 90° increments such that the angle between the average feature vector and the central axis of the first image is minimized, thereby causing the processor to preprocess the first image. The computer system according to claim 1.
4. The executable instructions cause the processor to preprocess the first image by obtaining the preprocessed first image as a luminance channel component of the first image. The computer system according to claim 1.
5. The executable instructions cause the processor to generate a plurality of frequency-converted blocks by frequency-converting the plurality of blocks, and further cause the processor to perform a process of embedding each bit set of the plurality of bit sets into each block of the plurality of blocks. The process of embedding each of the plurality of bit sets into each of the plurality of blocks is embedding each of the plurality of bit sets into each of the plurality of frequency-converted blocks, and in response to embedding the plurality of bit sets, inverse-transforming the plurality of frequency-converted blocks into the spatial domain, and is executed by The computer system according to claim 1.
6. The executable instructions cause the processor to execute a process of embedding a first bit set of the plurality of bit sets into a first block of the plurality of blocks. The process of embedding the first bit set of the plurality of bit sets into the first block of the plurality of blocks includes at least selecting a first bit of the first bit set, selecting a pixel position within the first block, identifying a level of the pixel position, and quantizing the level of the pixel position according to a logical state of the first bit, and is executed by The computer system according to claim 1.
7. The executable instructions that cause the processor to quantize the level of the pixel position according to the logical state of the first bit cause the processor to quantize the level to one of an odd value or an even value in response to determining that the logical state of the first bit is "0", and quantize the level to the other of the odd value or the even value in response to determining that the logical state of the first bit is "1", and execute The computer system according to claim 6.
8. The executable instructions cause the processor to divide the level of the quantized pixel position by a robustness coefficient. The computer system according to claim 7.
9. The executable instructions cause the processor to generate a watermarked image based on embedding the preprocessed information into the preprocessed first image, and output the watermarked image for distribution, and execute The computer system according to claim 1.
10. The information is a distributed identifier defined by the World Wide Web Consortium (W3C) Recommendation 1.
0. The computer system according to claim 1.
11. The executable instructions comprise configuring an image processing pipeline including a plurality of image preprocessing stages, selecting one or more of the plurality of image preprocessing stages according to the context information associated with the first image, the characteristics of the first image, the image type associated with the first image, or the preprocessing configuration, and generating the preprocessed first image by preprocessing the first image through the selected one or more of the plurality of image preprocessing stages, causing the processor to preprocess the first image. The computer system according to claim 1.
12. The executable instructions cause the processor to configure an information processing pipeline including a plurality of information preprocessing stages, select one or more of the plurality of information preprocessing stages according to the type of the information or the preprocessing configuration, and generate the preprocessed information by preprocessing the information through the selected one or more of the plurality of information preprocessing stages, causing the processor to preprocess the information. The computer system according to claim 1.
13. A method comprising: receiving a first image and information for embedding in the first image; generating a preprocessed first image by preprocessing the first image; generating preprocessed information including a plurality of bit sets by preprocessing the information using at least channel coding of the information; and embedding the preprocessed information in the preprocessed first image, wherein embedding the preprocessed information in the preprocessed first image comprises at least selecting a plurality of blocks of the preprocessed first image, and executing by embedding each bit set of the plurality of bit sets in each block of the plurality of blocks, each bit set of the plurality of bit sets being embedded in a minimum number of blocks of the plurality of blocks, the minimum number being greater than 1. A method.
14. Preprocessing the first image comprises generating a plurality of feature vectors by performing feature detection on the first image, determining an average feature vector of the plurality of feature vectors, and Rotating the first image in 90° increments such that the angle between the average feature vector and the central axis of the first image is minimized, The method according to claim 13.
15. Preprocessing the first image includes obtaining the preprocessed first image as a luminance channel component of the first image. The method according to claim 13.
16. Embedding a first bit set of the plurality of bit sets into a first block of the plurality of blocks, Embedding the first bit set of the plurality of bit sets into the first block of the plurality of blocks includes at least Selecting a first bit of the first bit set, Selecting a pixel position within the first block, Identifying the level of the pixel position, and Quantizing the level of the pixel position according to the logical state of the first bit. The method according to claim 13.
17. Quantizing the level of the pixel position according to the logical state of the first bit is Quantizing the level to one of an odd value or an even value in response to determining that the logical state of the first bit is "0", and Quantizing the level to the other of the odd value or the even value in response to determining that the logical state of the first bit is "1". The method according to claim 16.
18. Dividing the level of the quantized pixel position by a robustness coefficient The method according to claim 17.
19. Preprocessing the first image includes Constructing an image processing pipeline including a plurality of image preprocessing stages, Selecting one or more of the plurality of image preprocessing stages according to context information associated with the first image, characteristics of the first image, image type associated with the first image, or preprocessing configuration, and Generating the preprocessed first image by preprocessing the first image through the selected one or more of the plurality of image preprocessing stages. The method according to claim 13.
20. Preprocessing the information includes Constructing an information processing pipeline including a plurality of information preprocessing stages, selecting one or more of the plurality of information preprocessing stages according to the type of the information or the preprocessing configuration, and generating the preprocessed information by preprocessing the information through the selected one or more of the plurality of information preprocessing stages, The method according to claim 13.