Image watermarking
By redundant encoding and adversarial training of digital watermarks, the problem of watermark extraction in the image is solved, especially when there are different types of distortions, a watermark extraction effect with high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202080096987.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-01-13
AI Technical Summary
The prior art is difficult to effectively extract digital watermarks from images, especially when the images undergo distortion such as cropping, rotation, blurring or JPEG compression.
Digital watermarks are redundantly encoded using channel encoder and encoder models and adversarially trained in the decoder model to generate encoded images capable of recovering the original watermark in the presence of different types of distortions.
It is realized that digital watermarks are extracted with high accuracy when there are known and unknown types of distortions, improving the robustness and anti-distortion capabilities of the watermark system.
Smart Images

Figure CN115136183B_ABST
Abstract
Description
Technical Field
[0001] This specification generally relates to extracting digital watermarks embedded in images, regardless of distortions that may have been introduced into these images. Background Art
[0002] Image watermarking (also referred to as digital watermarking in this specification) is the process of embedding a digital watermark into an image - that is, embedding information into the image such that the image with the digital watermark is visually indistinguishable from the original image that does not include the digital watermark. Although image watermarking has several applications, traditionally it has been used to identify the ownership of copyright in an image or otherwise identify the source of the image. As an example, the source of an image may embed a digital watermark into the image before distributing the image. Subsequently, when the recipient receives the image, the recipient can extract the digital watermark from the image, and if the extracted digital watermark is the same as the digital watermark embedded into the image by the source, the recipient can confirm that the received image originated from the source.
[0003] However, from the time the image is distributed by the source until it is received by the target entity, one or more different types of distortions may be introduced into the image. Examples of such image distortions include, but are not limited to, cropping, rotation, blurring, and JPEG compression. Thus, when the recipient receives the image, the image may include one or more of such distortions. In some cases, the distortion may corrupt the image such that all or part of the digital watermark can no longer be extracted. As a result, the recipient of the image may not be able to confirm the source of the image. Summary of the Invention
[0004] In general, one innovative aspect of the subject matter described in this specification can be embodied in a method that includes the following operations: obtaining a first image and a first data item to be embedded in the first image; inputting the first data item into a channel encoder, where the channel encoder encodes an input data item of a first length into redundant data that implicitly or explicitly includes (1) the input data item and (2) new data that is redundant as at least a portion of the input data item and has a second length greater than the first length, where the new data can recover the input data in the presence of channel distortion; obtaining a first encoded data item from the channel encoder and in response to inputting the first data item into the channel encoder; inputting the first encoded data item and the first image into an encoder model, where the encoder model encodes the input image and the input data item to obtain an encoded image in which the input data item has been embedded as a digital watermark; and obtaining a first encoded image in which the first encoded data has been embedded as a digital watermark from the encoder model and in response to inputting the first encoded data item and the first image into the encoder model. Other embodiments of this aspect include corresponding systems, devices, apparatuses, and computer programs configured to perform the actions of the method. The computer program (e.g., instructions) can be encoded on a computer storage device.
[0005] These and other embodiments may each optionally include one or more of the following features.
[0006] In some implementations, the method may include the following operations: inputting the first encoded image into a decoder model, where the decoder model decodes the input encoded image to obtain data that is predicted to be embedded as a digital watermark within the input encoded image; obtaining a second data that is predicted to be the first encoded data from the decoder model and in response to inputting the first encoded image into the decoder model; inputting the second data into a channel decoder, where the channel decoder decodes the input data to recover the original data that was previously encoded by the channel encoder to generate the input data; and obtaining a third data that is predicted to be the first data from the channel decoder and in response to inputting the second data into the channel decoder.
[0007] In some embodiments, the method may include the following operations: obtaining a set of input training images; obtaining a first set of training images, wherein each image in the first set of training images is generated by encoding an input training image and an encoded data item using an encoder model, and wherein the encoded data item is generated by encoding an original data item using a channel encoder; inputting the first set of training images into an attack network, wherein the attack network uses a set of input images to generate a corresponding set of images including different types of image distortions; and using the attack network and in response to inputting the first set of input training images into the attack network, generating a second set of training images, wherein the images in the second set of training images correspond to the images in the first set of training images.
[0008] In some embodiments, the method may include the following operations of training the attack network using the first set of training images and the second set of training images, wherein the training includes: for each training image in the first set of training images and the corresponding training image in the second set of training images: inputting the training image from the second set of training images into a decoder model; obtaining, from the decoder model and in response to inputting the training image from the second set of training images into the decoder model, a first predicted data item predicted to be embedded as a digital watermark within the training image; determining a first image loss representing the difference in image pixel values between the training image in the first set of training images and the corresponding training image in the second set of training images; determining a first message loss representing the difference between the first predicted data item and the encoded data item embedded within the training image in the first set of training images; and using the first image loss and the first message loss to train the attack network.
[0009] In some embodiments, the method may include the operations of training the encoder model and the decoder model, wherein the training includes: for each training image in the first set of training images: inputting the training image into a decoder model; obtaining, from the decoder model and in response to inputting the training image into the decoder model, a second predicted data item predicted to be embedded within the training image; determining a second image loss representing the difference in image pixel values between the training image and the corresponding input training image; determining a second message loss representing the difference between the second predicted data item and the encoded data embedded within the training image; and using the second image loss, the second message loss, and the first message loss to train each of the encoder model and the decoder model.
[0010] In some embodiments, each of the attack model, the encoder model, and the decoder model may be a convolutional neural network.
[0011] In some embodiments, the second image loss may include an L2 loss and a GAN loss; and the second message loss may include an L2 loss.
[0012] In some embodiments, each of the first message loss and the first image loss may include an L2 loss.
[0013] In some embodiments, the method may include operations of training a channel encoder and a channel decoder, where the training includes: obtaining a set of training data items; for each training data item in the set of training data items: generating an encoded training data item using the channel encoder; for the encoded training data item, generating a modified training data item using a channel distortion approximation model, where the encoded training data item is distorted using the channel distortion approximation model to generate the modified training data item; determining a channel loss representing a difference between the encoded training data item and the corresponding modified training data item; and using the channel loss to train each of the channel encoder and the channel decoder.
[0014] Particular embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. The innovations described in this specification are capable of extracting a watermark message embedded within an encoded image, regardless of the type of distortion that may be introduced between the time the digital watermark is embedded into the image and the time the digital watermark is extracted from the encoded image. Traditional watermarking systems are trained on a particular type of image distortion. While such traditional systems can extract digital watermarks from images that have suffered the distortion that the systems are trained on with a high level of accuracy, these systems generally cannot extract digital watermarks from images that have suffered a different type of distortion (i.e., a distortion that the systems are not trained on) with the same level of accuracy. In contrast, the techniques described in this specification can (1) extract digital watermarks from images that have suffered a known type of distortion (i.e., the distortion that the traditional systems are trained on) with the same accuracy as traditional watermarking systems, and (2) extract digital watermarks from images that have suffered an unknown distortion (i.e., a distortion that the traditional systems are not trained on) with a higher level of accuracy compared to traditional watermarking systems. Thus, the techniques described in this specification can allow for more reliable encoding of watermarks (or other hidden data) in images transmitted over a noisy / distorted channel.
[0015] Correspondingly, the innovations described in this specification do not require any prior knowledge or exposure to specific distortions to achieve high-accuracy extraction of digital watermarks from images with the same distortion. Traditional watermarking systems typically require exposure to specific types of distortions during training in order to be able to extract digital watermarks embedded within images suffering from such distortions with a high level of accuracy. In contrast, the innovations described in this specification do not require any prior knowledge or exposure to specific distortions during training or otherwise to achieve high-accuracy extraction of watermark messages embedded within images suffering from such distortions. For example, as described throughout this document, the innovations described in this specification utilize adversarial training to achieve distortion-agnostic digital watermark extraction from images. As part of this adversarial training, an adversarial model generates training images that implicitly incorporate a wide collection of image distortions co-adapted with the training.
[0016] Furthermore, the innovations described in this specification are more robust than traditional systems. This is because the innovations described in this specification do not simply embed a digital watermark into an image, but rather add redundancy to the digital watermark before embedding the digital watermark into the image, which in turn increases the likelihood of recovering the original digital watermark in the presence of a certain reasonable amount of channel distortion.
[0017] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the specification, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a block diagram of an example environment in which a system is trained to embed a digital watermark into an image and then extract the digital watermark from the image.
[0019] Figure 2 is a flowchart of an example process for training a watermarking system to embed a digital watermark into an image and then extract the digital watermark from the image.
[0020] Figure 3 is an example environment in which a Figure 1 trained system embeds a digital watermark into an image and then extracts the digital watermark from the image.
[0021] Figure 4 is a flowchart of an example process for embedding a digital watermark into an image and then extracting the digital watermark from the image.
[0022] Figure 5 is a block diagram of an example computer system.
[0023] Like reference numerals and names in the different drawings indicate like elements. DETAILED IMPLEMENTATION MANNER
[0024] This specification generally relates to extracting digital watermarks embedded in images, regardless of the distortions that may have been introduced into these images.
[0025] Figure 1 It is a block diagram of an example environment 100 where the system is trained to embed a digital watermark into an image and then extract the digital watermark from the image. Refer to Figure 2 The structure and operation of the components of environment 100 are described.
[0026] Figure 2 It is a flowchart of an example process 200 for training a watermark system to embed a digital watermark into an image and then extract the digital watermark from the image. The operations of process 200 are described below for illustrative purposes only. The operations of process 200 can be performed by any suitable device or system (e.g., Figure 1 the system shown or any other suitable data processing device). The operations of process 200 can also be implemented as instructions stored on a non-transitory computer-readable medium. The execution of the instructions causes one or more data processing devices to perform the operations of process 200.
[0027] Process 200 obtains a set of input training images and a corresponding set of data items to be embedded into the input training images (at 202). In some embodiments, the data items are messages or other information to be embedded as digital watermarks into the images. The data items (which can be of any length) can be a binary digit string (e.g., 0100101) or a string consisting of digits (i.e., binary and non-binary digits) and / or characters. For example, the data item can simply be a text message, such as "WATERMARK" or "COPYRIGHT OF XYZ CORP". As another example, the data can be a digital signature or other fingerprint of an entity that is the intended source of the image and / or can be used to verify the source of the image. The set of input training images can be obtained from any storage location containing images (such as local storage on a user device or a storage location accessible via a network (e.g., an image archive, a social media platform, or another content platform)).
[0028] The operations 204 - 220 described below describe operations used in the adversarial training of an encoder model and a decoder model (described further below). As part of the adversarial training, two sets of training images are used to train the encoder model and the decoder model. The first set of training images is generated by the encoder model, and the second set of training images is generated by an attack model (described further below), which uses the first set of training images to generate a corresponding second set of images distorted using different types of image distortions. The attack network is trained to generate training images that include a large number of different amounts of image distortion - i.e., much more than the specific types of distortion used in the training of traditional systems. The decoder model decodes the two sets of training images to obtain data items embedded within each image as digital watermarks. Then, the encoder model and the decoder model are trained based on the resulting "image loss" and "message loss" (described further below) for each image in the first set of training images and the corresponding images in the second set of training images. Thus, to perform this adversarial training, operations 204 - 220 are iteratively performed for each training image in the set of input training images and for each corresponding data item from the set of data items to be embedded into the input training images.
[0029] Process 200 generates an encoded data item based on a data item (at 204). Process 200 inputs data item 102 into channel encoder 104, and channel encoder 104 outputs an encoded data item. In some embodiments, channel encoder 104 is a machine learning model (e.g., a neural network model) that is trained to encode each input data item into a redundant data item (which is also referred to as an encoded data item) that implicitly or explicitly includes (1) the input data and (2) new data that is redundant to the input data item (e.g., the entire input data or a redundancy of at least a portion of the input data). The channel encoder is trained to generate the redundant data item / encoded data item such that the input data item can be recovered in the presence of channel distortion. As used in this specification, channel distortion refers to the systematic error between the point at which the encoded data item is generated and the point at which the decoder model outputs the data item predicted to be embedded within the image. Training the channel encoder to generate the redundant data / encoded data item is described below. The new data of the redundant data item can be added to the encoded data item in different ways. For example, for the data item {001100}, the encoded data item includes the data item and redundant data that replicates (one or more times) the input data item. In this example, the encoded data item can be replicated twice, resulting in {001100001100}, which is twice the length of the input data item. Another example technique for adding redundancy in the form of new data of the redundant data item includes, but is not limited to, Hamming codes (also referred to as block codes).
[0030] Process 200 generates a training image (at 206) by embedding an encoded data item (which is generated in operation 204) as a digital watermark into an input training image. In some embodiments, the encoder model 110 receives the encoded data item and the input training image as inputs, and outputs an encoded image in which the encoded data item is embedded as a digital watermark. The encoder model 110 is a convolutional neural network (CNN).
[0031] Process 200 inputs the training image into an attack network, and the attack network outputs a modified image including a specific type of image distortion (at 208). In some embodiments, the attack network 112 can be a two-layer convolutional neural network (CNN) that generates a modified image with a set of different image distortions based on the input training image. In other embodiments, the attack network can be the Fast Gradient Sign Method (FGSM).
[0032] For each of the modified image (generated in operation 208) and the training image (generated in operation 206), process 200 predicts the digital watermark embedded in the image (at 210). In some embodiments, the modified image is input into the decoder model 114, and the decoder model 114 outputs a first predicted data item that is predicted to be the digital watermark embedded within the modified image. Similarly, the training image is input into the decoder model 114, and the decoder model 114 outputs a second predicted data item that is predicted to be the digital watermark embedded within the training image. Like the encoder model 110, the decoder model 114 is a CNN.
[0033] Process 200 determines a first image loss that represents the difference in image pixel values between the training image and the modified image (at 212). In some embodiments, the first image loss includes the L2 loss between the image pixel values of the training image and the modified image. Thus, the first image loss can be represented using the following equation:
[0034]
[0035] As used in the above equation, "I adv " refers to the image pixel values of the modified image, "I enc " refers to the image pixel values of the encoded image / training image, and "α 1 " refers to the scalar weight. Alternatively, other losses that compare the image pixel values of the training image and the modified image, such as the L1 loss or p-norm, can be used. As another alternative, other image metrics (such as resolution, error rate, fidelity, and signal-to-noise ratio) can be used instead of image pixel values.
[0036] Process 200 determines a first message loss (at 214) that represents a difference between a first predicted data item and an encoded data item embedded in a training image. In some embodiments, the first message loss includes an L2 loss between the first predicted data item and the encoded data item. Thus, the first message loss can be represented using the following equation:
[0037]
[0038] As used in the above equation, "X' adv " refers to the first predicted data item, "X'" refers to the encoded data item, and "α 2 " refers to a scalar weight. Alternatively or additionally, other losses that compare the first predicted data item and the encoded data item can be used, such as an L1 loss or a p-norm.
[0039] Process 200 uses the first image loss and the first message loss to train the attack network 112 (at 216). In some embodiments, the attack network 112 is trained to minimize a training loss, which is represented by the difference between the first image loss and the first message loss. As an example, the attack network 112 can be trained to minimize the following training loss:
[0040]
[0041] As used in the above equation, "L adv " refers to the training loss of the attack network. All other parameters referenced in the above equation are described with reference to operations 212 and 214. The scalar weight "α 1 " controls the intensity of the distortion generated by the attack network 112, while the first message loss prompts the attack network 112 to generate a modified training image that reduces the bit accuracy. Additionally, the complexity of the attack network 112 (e.g., based on the number of layers of the CNN used) and the scalar weight "α 2 " provide a measure of the strength of the attack network. However, in some embodiments, other training losses that compare the first image loss and the first message loss can be used alternatively or additionally.
[0042] Process 200 determines a second image loss (at 218) that represents an image pixel value difference between a training image and a corresponding input training image. In some embodiments, the second image loss includes an L2 loss and a Wasserstein generative adversarial network (WGAN) loss from an evaluation network that is trained to distinguish the image pixel values of the training image from the corresponding input training image. Thus, the second image loss can be represented using the following equation:
[0043]
[0044] As used in the above equation, "L" I " refers to the second image loss, "α" 1 I " and "α" 2 I " are scalar weights, "I" co " refers to the image pixel values of the input training image, and "I" en " refers to the image pixel values of the training image, which is the encoded version of the input training image into which the encoded data item has been embedded. However, in some embodiments, other loss functions for comparing the difference in image pixel values between the training image and the corresponding input training image may alternatively or additionally be used. For example, the WGAN loss function may be replaced with the min-max loss function. As another alternative, other image metrics (such as resolution, error rate, fidelity, and signal-to-noise ratio) may be used instead of the image pixel values.
[0045] Process 200 determines a second message loss (at 220) representing the difference between the second predicted data item and the encoded data item embedded in the training image. In some embodiments, the second message loss is the L2 loss between the second predicted data item and the encoded data item embedded in the training image. Thus, the second message loss can be represented using the following equation:
[0046] L M = α M ||X' dec - X''|| 2
[0047] As used in the above equation, "L" M " refers to the second message loss, "X'" dec " refers to the second predicted data item, "X''" refers to the encoded data item, and "α" 1 M " is the scalar weight. Other losses for comparing the second predicted data item and the encoded data item, such as the L1 loss or the p-norm, may alternatively or additionally be used.
[0048] Process 200 uses the image loss, the message loss, and the first message loss to train each of the encoder model and the decoder model (at 222). In some embodiments, each of the encoder model 110 and the decoder model 114 is trained to minimize an overall model loss that is a combination (e.g., sum) of the second image loss, the second message loss, and the first message loss.
[0049] In some embodiments, Figure 1The system shown does not include channel encoder 104. In this implementation, data items are directly embedded into the input training image by encoder model 110 (as opposed to generating encoded data items and then embedding the encoded data items into the input training image).
[0050] Figure 3 is an example environment 300 in which a training system Figure 1 embeds a digital watermark into an image and then extracts the digital watermark from the image. Block diagram of the environment 300.
[0051] Except for the attack network 112 that is only used during training, the system of environment 300 includes all components of the system of environment 100. Thus, the environment system 300 includes channel encoder 104, encoder model 110, decoder model 114, and channel decoder 308. The structure and operation of these components of environment 300 have been described in the context of training with reference to Figure 1 and Figure 2 . The same operations occur when these components are used to embed a digital watermark into an image and then extract the digital watermark from the image. These operations are generally described with reference to Figure 4 .
[0052] In some implementations, the channel decoding models (i.e., channel encoder 104 and channel decoder 308) are trained as follows. First, a set of training data items is obtained. In some implementations, the training data items are the same set of data items referred to at operation 202. The training data items are input into channel encoder 104, which in turn generates a corresponding set of encoded training data items (as described above with reference to Figure 2 ). For each encoded training data item in the set of encoded training data items, a set of modified training data items (also referred to as noise samples) is generated using a binary symmetric channel (BSC) model, which is used to approximate the channel distortion (which refers to the systematic error between the point where the encoded data item is generated and the point where the decoder model outputs the data item predicted to be embedded within the image). The BSC is a standard channel model that assumes each bit independently randomly flips with some probability p. Alternatively, other channel models (such as a binary erasure channel (BEC) model) can be used instead of the BSC.
[0053] Each of the channel encoder 104 and the channel decoder 308 is trained to minimize a channel loss, which represents the loss between the encoded training data item and each modified training data item. In some embodiments, the channel loss can be the VIMCO loss, which represents a multi-sample variance lower bound objective for obtaining low variance gradients. The channel decoding model can be trained for a certain number of iterations, or until the loss of the channel decoding model reaches (e.g., equals or is below) a certain loss threshold.
[0054] Figure 4 is a flowchart of an example process 400 for embedding a digital watermark into an image and subsequently extracting the digital watermark from the image. The operations of process 400 are described below for illustrative purposes only. The operations of process 400 can be performed by any suitable device or system (e.g., Figure 3 the system shown or any other suitable data processing device). The operations of process 400 can also be implemented as instructions stored on a non-transitory computer-readable medium. The execution of the instructions causes one or more data processing devices to perform the operations of process 400.
[0055] Process 400 obtains a first image 304 and a first data item 302 to be embedded in the first image 304 (at 402). The first image 304 and the first data item 302 can be obtained from the same or similar image sources as described with reference to Figure 2 In addition, the first data item 302 is of the same type as the data item described with reference to Figure 2 description.
[0056] Process 400 inputs the first data item 302 into the channel encoder 104 (at 404).
[0057] Process 400 obtains a first encoded data item from the channel encoder 104 and in response to inputting the first data item 302 into the channel encoder (at 406). As described with reference to Figure 2 The channel encoder 104 encodes an input data item of a first length into redundant data that implicitly (e.g., only some input data items or representations of the input data item) or explicitly (e.g., a copy of the entire input data item) includes (1) the input data item and (2) new data that is redundant as at least a portion of the input data item, and has a second length greater than the first length. Accordingly, the first encoded data item includes redundant data that implicitly or explicitly includes (1) the first data item 302 and (2) new data that is redundant as at least a portion of the first data item 302, and has a second length greater than the first length. In addition, the redundancy in the first encoded data item can recover the first data item in the presence of channel distortion.
[0058] Process 400 inputs the first encoded data item and the first image 304 into the encoder model 110 (at 408).
[0059] Process 400 obtains the first encoded image from the encoder model 110 (at 410). As referenced Figure 2 as described, the encoder model 110 encodes the input image and the input data item to obtain an encoded image in which the input data item has been embedded as a digital watermark. Accordingly, the first encoded image output by the encoder model 110 embeds the first encoded data item as a digital watermark into the first image 304.
[0060] Process 400 inputs the first encoded image into the decoder model 114 (at 412).
[0061] Process 400 obtains a second data item from the decoder model 114 and in response to inputting the first encoded image into the decoder model 14 (at 414). As referenced Figure 2 as described, the decoder model decodes the input encoded image to obtain data that is predicted to be embedded as a digital watermark within the input encoded image. Accordingly, the second data item output by the decoder model 114 is predicted to be the first encoded data item embedded as a digital watermark within the first encoded image.
[0062] Process 400 inputs the second data item into the channel decoder 308 (at 416).
[0063] Process 400 obtains a third data item 306 from the channel decoder 308 and in response to inputting the second data item into the channel decoder 308 (at 418). As referenced Figure 2 as described, the channel decoder 308 decodes the input data to recover the original data that was previously encoded by the channel encoder to generate the input data. Accordingly, the third data item generated by the channel decoder 308 is predicted to be the first data item that was previously encoded by the channel encoder into the first encoded data item.
[0064] In addition, as referenced Figure 1 and Figure 2 as described, the system of environment 300 may but need not include a channel decoding model. Accordingly, in some embodiments, the system of environment 300 may include only the encoder model 110 and the decoder model 114 (i.e., the environment may not include the channel encoder 104 and the channel decoder 308).
[0065] In some embodiments, the system shown in the example environment 300 (and Figure 4The corresponding operations described in [reference] are performed by the same entity. Alternatively, the channel encoder 104 and the encoder model 110 may be implemented by one entity, and the decoder model 114 and the channel decoder 308 may be implemented by separate entities. In this implementation, the entity that performs data and / or image encoding is different from the entity that performs data and / or image decoding.
[0066] Thus, as described in reference to Figure 3 and Figure 4 (and the corresponding descriptions in Figure 1 and Figure 2 ), this specification describes techniques for extracting digital watermarks from images without regard to the type of image distortion that may have been introduced into the image.
[0067] Figure 5 FIG. [reference] is a block diagram of an example computer system 500 that may be used to perform the operations described above. System 500 includes a processor 510, a memory 520, a storage device 530, and an input / output device 540. Each of the components 510, 520, 530, and 540 may be interconnected, for example, using a system bus 550. The processor 510 is capable of processing instructions running within the system 500. In some embodiments, the processor 510 is a single-threaded processor. In another embodiment, the processor 510 is a multi-threaded processor. The processor 510 is capable of processing instructions stored in the memory 520 or stored on the storage device 530.
[0068] The memory 520 stores information within the system 500. In one embodiment, the memory 520 is a computer-readable medium. In some embodiments, the memory 520 is a volatile memory unit. In another embodiment, the memory 520 is a non-volatile memory unit.
[0069] The storage device 530 is capable of providing mass storage for the system 500. In some embodiments, the storage device 530 is a computer-readable medium. In various different embodiments, the storage device 530 may include, for example, a hard disk device, an optical disk device, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other mass storage device.
[0070] The input / output device 540 provides input / output operations for the system 500. In some embodiments, the input / output device 540 may include one or more of a network interface device (e.g., an Ethernet card), a serial communication device (e.g., an RS-232 port), and / or a wireless interface device (e.g., an 802.11 card). In another embodiment, the input / output device may include driver devices configured to receive input data and send output data to other input / output devices, such as a keyboard, a printer, and the display device 560. However, other embodiments may also be used, such as mobile computing devices, mobile communication devices, set-top box TV client devices, etc.
[0071] Although example processing systems are described in Figure 5 this specification, embodiments of the subject matter and functional operations described in this specification may also be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them.
[0072] Embodiments of the subject matter and operations described in this specification may be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs (i.e., one or more modules of computer program instructions) encoded on a computer storage medium (or media) for execution by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or be included in them. Further, although a computer storage medium is not a propagated signal, a computer storage medium may be a source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium may also be one or more separate physical components or media (e.g., multiple CDs, optical disks, or other storage devices), or be included in them.
[0073] The operations described in this specification may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0074] The term "data processing apparatus" encompasses all kinds of devices, equipment, and machines for processing data, including, for example, programmable processors, computers, system-on-chips, or multiples or combinations of the foregoing. The apparatus may include dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). In addition to hardware, the apparatus may also include code that creates a runtime environment for the computer program being discussed, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and the runtime environment may implement various different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.
[0075] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (for example, one or more scripts stored in a markup language document), stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to run on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0076] The processes and logical flows described in this specification can be performed by one or more programmable processors running one or more computer programs to perform actions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry, and the apparatus can also be implemented as dedicated logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0077] For example, a processor suitable for running a computer program includes general and special purpose microprocessors. Generally, the processor will receive instructions and data from a read only memory or a random access memory or both. The basic elements of a computer are a processor for performing actions in accordance with the instructions and one or more storage devices for storing the instructions and data. Generally, a computer will also include or be operably coupled to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, from which it can receive data, or to which it can transfer data, or both. However, a computer need not have such devices. In addition, a computer may be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), etc. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example: semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0078] To provide for interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user may be in any form, including sound, voice, or tactile input. In addition, a computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on a client device of the user in response to a request received from the web browser.
[0079] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification), or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0080] The computing system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship between a client and a server arises from computer programs that run on respective computers and have a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a client device (e.g., to display data to and receive user input from a user interacting with the client device). Data generated at the client device (e.g., the result of a user interaction) can be received at the server from the client device.
[0081] Although this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that are specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although a feature may have been described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be deleted from the combination, and the claimed combination can be directed to a sub-combination or a variant of a sub-combination.
[0082] Similarly, although the operations are described in a particular order in the figures, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated together in a single software product or packaged into multiple software products.
[0083] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still obtain the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to obtain the desired result. In some implementations, multitasking and parallel processing may be advantageous.
Claims
1. A computer-implemented method, comprising: obtaining a first image and a first data item to be embedded in the first image; inputting the first data item into a channel encoder, wherein the channel encoder encodes an input data item of a first length into redundant data, the redundant data including (1) the input data item and (2) new data that is redundant of the input data item and having a second length greater than the first length, wherein the channel encoder is trained to encode the input data item such that the new data can recover the input data item in the presence of channel distortion; obtaining a first encoded data item from the channel encoder and in response to inputting the first data item into the channel encoder; inputting the first encoded data item and the first image into an encoder model, wherein the encoder model encodes the input image and the input data item to obtain an encoded image in which the input data item has been embedded as a digital watermark; and obtaining a first encoded image in which the first encoded data item has been embedded as a digital watermark from the encoder model and in response to inputting the first encoded data item and the first image into the encoder model; inputting the first encoded image into a decoder model, wherein the decoder model decodes the input encoded image to obtain data predicted to be embedded as a digital watermark within the input encoded image; obtaining a second data predicted to be the first encoded data item from the decoder model and in response to inputting the first encoded image into the decoder model; inputting the second data into a channel decoder, wherein the channel decoder decodes the input data item to recover the original data previously encoded by the channel encoder to generate the input data item; and obtaining a third data predicted to be the first data item from the channel decoder and in response to inputting the second data into the channel decoder, and wherein each of the channel encoder and the channel decoder includes a neural network, and wherein the method further includes training the channel encoder and the channel decoder, wherein the training includes: obtaining a set of training data items; for each training data item in the set of training data items: generating an encoded training data item using the channel encoder; for the encoded training data item, generating a modified training data item using a channel distortion approximation model, wherein the encoded training data item is distorted using the channel distortion approximation model to generate the modified training data item; determining a channel loss representing a difference between the encoded training data item and the corresponding modified training data item; and using the channel loss to train each of the channel encoder and the channel decoder.
2. The computer-implemented method according to claim 1, further comprising: obtaining a set of input training images; Obtain a first set of training images, where each image in the first set of training images is generated by encoding an input training image and an encoded data item using the encoder model, and the encoded data item is generated by encoding an original data item using the channel encoder; Input the first set of training images into an attack network, where the attack network uses a set of input images to generate a corresponding set of images including different types of image distortions; and Use the attack network and in response to inputting the first set of training images into the attack network, generate a second set of training images, where the images in the second set of training images correspond to the images in the first set of training images.
3. The computer-implemented method according to claim 2, wherein, the attack network is a neural network, and further includes training the attack network using the first set of training images and the second set of training images, where the training includes: For each training image in the first set of training images and the corresponding training image in the second set of training images: Input the training image from the second set of training images into the decoder model; Obtain from the decoder model and in response to inputting the training image from the second set of training images into the decoder model, a first predicted data item predicted to be embedded as a digital watermark in the training image; Determine a first image loss representing the difference in image pixel values between the training image in the first set of training images and the corresponding training image in the second set of training images; Determine a first message loss representing the difference between the first predicted data item and the encoded data item embedded in the training image in the first set of training images; and Use the first image loss and the first message loss to train the attack network.
4. The computer-implemented method according to claim 3, further includes training the encoder model and the decoder model, wherein, the training includes: For each training image in the first set of training images: Input the training image into the decoder model; Obtain from the decoder model and in response to inputting the training image into the decoder model, a second predicted data item predicted to be embedded in the training image; Determine a second image loss representing the difference in image pixel values between the training image and the corresponding input training image; Determine a second message loss representing the difference between the second predicted data item and the encoded data embedded in the training image; and Use the second image loss, the second message loss, and the first message loss to train each of the encoder model and the decoder model.
5. The method according to claim 3, wherein, each of the attack network, the encoder model, and the decoder model is a convolutional neural network.
6. The method according to claim 4, wherein: the second image loss includes L2 loss and GAN loss; and the second message loss includes L2 loss.
7. The method according to claim 4, wherein, each of the first message loss and the first image loss includes an L2 loss.
8. A computer-implemented system, comprising: one or more storage devices that store instructions; and one or more data processing devices configured to interact with the one or more storage devices and perform operations including the following when running the instructions: obtain a first image and a first data item to be embedded in the first image; input the first data item into a channel encoder, wherein the channel encoder encodes an input data item of a first length into redundant data, the redundant data including (1) the input data item and (2) new data that is redundant to the input data item and has a second length greater than the first length, and wherein the channel encoder is trained to encode the input data item such that the new data can recover the input data item in the presence of channel distortion; obtain a first encoded data item from the channel encoder and in response to inputting the first data item into the channel encoder; input the first encoded data item and the first image into an encoder model, wherein the encoder model encodes the input image and the input data item to obtain an encoded image in which the input data item has been embedded as a digital watermark; and obtain a first encoded image in which the first encoded data item has been embedded as a digital watermark from the encoder model and in response to inputting the first encoded data item and the first image into the encoder model; input the first encoded image into a decoder model, wherein the decoder model decodes the input encoded image to obtain data predicted to be embedded as a digital watermark in the input encoded image; obtain a second data predicted to be the first encoded data item from the decoder model and in response to inputting the first encoded image into the decoder model; input the second data into a channel decoder, wherein the channel decoder decodes the input data item to recover the original data previously encoded by the channel encoder to generate the input data item; and obtain a third data predicted to be the first data item from the channel decoder and in response to inputting the second data into the channel decoder, and wherein each of the channel encoder and the channel decoder includes a neural network, and wherein the one or more data processing devices are configured to perform operations further including training the channel encoder and the channel decoder, and wherein the training includes: obtain a set of training data items; for each training data item in the set of training data items: generate an encoded training data item using the channel encoder; for the encoded training data item, generate a modified training data item using a channel distortion approximation model, wherein the encoded training data item is distorted using the channel distortion approximation model to generate the modified training data item; Determine a channel loss representing the difference between the encoded training data item and the corresponding modified training data item; and Use the channel loss to train each of the channel encoder and the channel decoder.
9. The system according to claim 8, wherein, the one or more data processing devices are configured to perform operations further comprising: Obtain a set of input training images; Obtain a first set of training images, wherein each image in the first set of training images is generated by encoding an input training image and an encoded data item using the encoder model, and wherein the encoded data item is generated by encoding an original data item using the channel encoder; Input the first set of training images into an attack network, wherein the attack network uses a set of input images to generate a corresponding set of images including different types of image distortions; and Use the attack network and in response to inputting the first set of training images into the attack network, generate a second set of training images, wherein the images in the second set of training images correspond to the images in the first set of training images.
10. The system according to claim 9, wherein, wherein, the attack network is a neural network, and the one or more data processing devices are configured to perform operations further comprising training the attack network using the first set of training images and the second set of training images, wherein the training includes: For each training image in the first set of training images and the corresponding training image in the second set of training images: Input the training image from the second set of training images into the decoder model; Obtain from the decoder model and in response to inputting the training image from the second set of training images into the decoder model, a first predicted data item predicted to be embedded as a digital watermark within the training image; Determine a first image loss representing the difference in image pixel values between the training image in the first set of training images and the corresponding training image in the second set of training images; Determine a first message loss representing the difference between the first predicted data item and the encoded data item embedded in the training image in the first set of training images; and Use the first image loss and the first message loss to train the attack network.
11. The system according to claim 10, wherein, the one or more data processing devices are configured to perform operations further comprising training the encoder model and the decoder model, wherein the training includes: For each training image in the first set of training images: Input the training image into the decoder model; Obtain from the decoder model and in response to inputting the training image into the decoder model, a second predicted data item predicted to be embedded within the training image; Determine a second image loss representing the difference in image pixel values between the training image and the corresponding input training image; Determine a second message loss representing the difference between the second predicted data item and the encoded data embedded in the training image; and Each of the encoder model and the decoder model is trained using the second image loss, the second message loss, and the first message loss.
12. The system according to claim 10, wherein, each of the attack network, the encoder model, and the decoder model is a convolutional neural network.
13. The system according to claim 11, wherein: the second image loss includes an L2 loss and a GAN loss; and the second message loss includes an L2 loss.
14. The system according to claim 11, wherein, each of the first message loss and the first image loss includes an L2 loss.
15. A non-transitory computer-readable medium storing instructions that, when executed by one or more data processing devices, cause the one or more data processing devices to perform operations including: obtaining a first image and a first data item to be embedded in the first image; inputting the first data item into a channel encoder, wherein, the channel encoder encodes an input data item of a first length into redundant data that includes (1) the input data item and (2) new data that is redundant to the input data item and has a second length greater than the first length, wherein the channel encoder is trained to encode the input data item such that the new data can recover the input data item in the presence of channel distortion; obtaining a first encoded data item from the channel encoder and in response to inputting the first data item into the channel encoder; inputting the first encoded data item and the first image into an encoder model, wherein the encoder model encodes the input image and the input data item to obtain an encoded image in which the input data item has been embedded as a digital watermark; and obtaining a first encoded image in which the first encoded data item has been embedded as a digital watermark from the encoder model and in response to inputting the first encoded data item and the first image into the encoder model; inputting the first encoded image into a decoder model, wherein the decoder model decodes the input encoded image to obtain data predicted to be embedded as a digital watermark within the input encoded image; obtaining a second data predicted to be the first encoded data item from the decoder model and in response to inputting the first encoded image into the decoder model; inputting the second data into a channel decoder, wherein the channel decoder decodes the input data item to recover the original data that was previously encoded by the channel encoder to generate the input data item; and obtaining a third data predicted to be the first data item from the channel decoder and in response to inputting the second data into the channel decoder, and wherein each of the channel encoder and the channel decoder includes a neural network, and wherein the instructions cause the one or more data processing devices to perform operations including training the channel encoder and the channel decoder, wherein the training includes: Obtain a set of training data items; For each training data item in the set of training data items: Generate an encoded training data item using the channel encoder; For the encoded training data item, generate a modified training data item using a channel distortion approximation model, wherein the encoded training data item is distorted using the channel distortion approximation model to generate the modified training data item; Determine a channel loss representing the difference between the encoded training data item and the corresponding modified training data item; and Use the channel loss to train each of the channel encoder and the channel decoder.
16. The non-transitory computer-readable medium according to claim 15, wherein, the instructions cause the one or more data processing devices to perform operations including: Obtain a set of input training images; Obtain a first set of training images, wherein each image in the first set of training images is generated by encoding an input training image and an encoded data item using the encoder model, and wherein the encoded data item is generated by encoding an original data item using the channel encoder; Input the first set of training images into an attack network, wherein the attack network uses a set of input images to generate a corresponding set of images including different types of image distortions; and Use the attack network and in response to inputting the first set of training images into the attack network, generate a second set of training images, wherein the images in the second set of training images correspond to the images in the first set of training images.
17. The non-transitory computer-readable medium according to claim 16, wherein, the attack network is a neural network, and wherein the instructions cause the one or more data processing devices to perform an operation including training the attack network using the first set of training images and the second set of training images, wherein the training includes: For each training image in the first set of training images and the corresponding training image in the second set of training images: Input the training image from the second set of training images into the decoder model; Obtain a first predicted data item predicted to be embedded as a digital watermark within the training image from the decoder model and in response to inputting the training image from the second set of training images into the decoder model; Determine a first image loss representing the difference in image pixel values between the training image in the first set of training images and the corresponding training image in the second set of training images; Determine a first message loss representing the difference between the first predicted data item and the encoded data item embedded in the training image in the first set of training images; and Use the first image loss and the first message loss to train the attack network.
18. The non-transitory computer-readable medium according to claim 17, wherein, the instructions cause the one or more data processing devices to perform operations including training the encoder model and the decoder model, wherein the training includes: For each training image in the first set of training images: Input the training image into the decoder model; Obtain, from the decoder model and in response to inputting the training image into the decoder model, a second predicted data item predicted to be embedded within the training image; Determine a second image loss representing a difference in image pixel values between the training image and a corresponding input training image; Determine a second message loss representing a difference between the second predicted data item and encoded data embedded within the training image; and Train each of the encoder model and the decoder model using the second image loss, the second message loss, and the first message loss.
19. The non-transitory computer-readable medium according to claim 17, wherein, each of the attack network, the encoder model, and the decoder model is a convolutional neural network.
20. The non-transitory computer-readable medium according to claim 18, wherein: the second image loss includes an L2 loss and a GAN loss; and the second message loss includes an L2 loss.
21. The non-transitory computer-readable medium according to claim 18, wherein, each of the first message loss and the first image loss includes an L2 loss.
Citation Information
Patent Citations
Watermarking project of an analogue video
CN1756340A