Watermark-based Image Reconstruction
By embedding lost data messages in the encoding process in the image and extracting and reconstructing images using machine learning models, the problem of image fidelity loss in the prior art is solved, and higher quality image reconstruction is achieved.
Patent Information
- Application Number
- CN201980101255.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-12-05
AI Technical Summary
The prior art requires a choice between loss of fidelity and resource investment when storing and transmitting images and video files, and cannot effectively recover data lost in the lossy encoding process.
The lost data during encoding is embedded into the image as a message by using a machine learning model and extracting the message at subsequent decoding to reconstruct the original image.
It realizes that when using lossy compression technology, the quality of the reconstructed image is significantly improved, and the version close to the original image can be restored, reducing the impact of data loss during the encoding process.
Smart Images

Figure CN114730450B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to watermark-based image reconstruction techniques. More specifically, the present disclosure relates to systems and methods for watermarking encoded images that have data lost during an encoding step for later image reconstruction. Background Art
[0002] Image and video files are typically encoded using various encoding schemes for many different applications. Large data entities require significant resource investment and storage infrastructure allocation to store these image and / or video files. Thus, these files typically must be encoded using lossy compression schemes to enable more efficient storage, retrieval, and transmission.
[0003] However, encoding image and / or video files using lossy compression schemes can result in significant loss of image fidelity, and lossless compression alternatives rarely provide the space reduction needed to achieve efficient storage. Thus, a choice must be made between loss of fidelity and significant resource investment in additional storage infrastructure when storing and transmitting image and / or video files. Summary of the Invention
[0004] Aspects and advantages of embodiments of the present invention will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0005] One example aspect of the present disclosure relates to a computer-implemented method for performing watermark-based image reconstruction to compensate for a lossy encoding scheme. The method includes obtaining an input image by one or more computing devices. The method includes generating a first output image by one or more computing devices by encoding and decoding the input image according to an encoding scheme. The method includes determining a difference image that describes a difference between the input image and the first output image by one or more computing devices. The method includes generating a second output image by one or more computing devices and using a machine learning message embedding model, the second output image including an embedded message that is at least partially based on the difference image.
[0006] Other aspects of the present disclosure are directed to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0007] These and other features, aspects, and advantages of various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. Brief Description of the Drawings
[0008] Figure 1 A block diagram of an example computing system in accordance with an example embodiment of the present disclosure is depicted.
[0009] Figure 2 A flowchart depicting an example method for generating a second output image according to an example embodiment of the present disclosure.
[0010] Figure 3 A flowchart depicting an example method for reconstructing an input image from a second output image according to an example embodiment of the present disclosure.
[0011] Figure 4 A flowchart depicting an example method for generating a second output image according to an example embodiment.
[0012] Figure 5 A flowchart depicting an example method for reconstructing an input image from a second output image according to an example embodiment.
[0013] Figure 6 A flowchart depicting an example method for training one or more machine learning models according to an example embodiment. Detailed Description
[0014] Overview
[0015] Example embodiments of the present disclosure relate to systems and methods for watermark - based image reconstruction using machine learning models. In particular, the systems and methods described herein relate to using a machine - learning message - embedding model to embed a message into an image, where the message represents data lost from the image due to an encoding (and decoding) process. The embedded message can later be extracted from the image by a machine - learning message - extraction model, and the extracted message can be used to reconstruct the original image. Thus, as an example, a lossy compression technique (e.g., JPEG compression) can be used to compress an image to generate a first output image. A message representing data lost from the image due to such compression can be embedded (e.g., as an image - noise watermark) within the original image to generate a second output image. The embedded message can be extracted from the second output image and can be used to reconstruct the original image from the second output image, thereby at least partially reversing the loss of image fidelity caused by the compression. The proposed technique represents a significant advancement in reconstructing images that have suffered data loss during an image - encoding process. In particular, by capturing and embedding the data lost from the image during the compression process as a message, the proposed system provides a method for image reconstruction that can produce a reconstructed image more accurately than traditional techniques.
[0016] As an example, a computing device (e.g., a distributed network of computing devices) can obtain an input image (e.g., a RAW image). The computing device can generate a first output image by encoding and decoding the image according to an encoding scheme. As an example, the computing device can encode and decode the input image using a lossy JPEG compression scheme, and the first output image is the decoded JPEG representation of the input image. A difference image that describes the difference between the first image and the first output image can be determined. For example, the difference image can describe the lost data as a pixel-by-pixel difference between the input image and the first output image. A machine learning message embedding model can embed a message into the first image (e.g., as an image noise watermark) based on the difference image to produce a second output image. As an example, the data lost from JPEG compression can be represented as a latent space message vector. A watermark can be generated based on the message vector. The message can be embedded into the JPEG (e.g., by applying the watermark as image noise) to produce the second output image. The second output image can be encoded and then stored or transmitted.
[0017] The encoded second output image can be decoded, and using a machine learning message extraction model, the message vector can be extracted from the second output image and used to reconstruct the difference image. For example, a machine learning watermark extraction model can extract the message vector, and a machine learning difference reconstruction model can use the extracted message vector to generate a reconstruction of the difference image, which can then be used to reconstruct the input image. As an example, the input image can be reconstructed by adding the reconstructed difference image to the second output image. Thus, although the input image in the above example is degraded due to the lossy compression scheme, the same or nearly the same version of the input image can be reconstructed using the embedded message watermarked in the image.
[0018] The present disclosure provides many technical effects and benefits. As an example of a technical effect and benefit, the systems and methods of the present disclosure achieve a significant improvement in the quality of reconstructed images compared to other methods. Most other methods known in the art attempt to directly recover the input image from the encoded image. For example, methods such as GIF2Video remove GIF artifacts by combining neural networks and the Lucas-Kanade method (see Yang Wang, Haibin Huang, Chuan Wang, Tong He, Jue Wang, Minh Hoai, GIF2Video: Color Dequantization and Temporal Interpolation of GIF Images, The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1419-1428 (2019)). However, under this method, the reconstruction accuracy of the machine learning model is necessarily limited because it cannot utilize the data that was initially lost during encoding. The present method differs from these prior methods in that it embeds a message (e.g., a watermark) that describes the data lost during the encoding step into the image. Later, the message can be extracted and used to reconstruct the input image. In this way, the machine learning model can utilize the initial data that was lost during encoding, resulting in a more accurate reconstructed image.
[0019] As another exemplary technical effect and benefit, the systems and methods of the present disclosure enable multiple image encoding schemes (e.g., Graphics Interchange Format (GIF) encoding) to be used in situations where they may not have been selected previously. As an example, due to the data loss inherent in certain lossy GIF encoding schemes, lossy GIF encoding schemes may not have been selected in some cases. Using the methods of the present disclosure, GIF-encoded images can be reconstructed to an accuracy sufficient to enable GIF encoding in situations where minimal image data loss is required. By enabling these additional encoding schemes, the present disclosure allows for the compression of more images, necessarily saving storage space for storing the images. In other words, aspects of the present disclosure represent an improvement in the curve of compression gain versus quality reduction. Thus, compared to past compression techniques, the present disclosure is able to achieve additional compression gain while still maintaining the same final quality. These compression gains result in savings in resources such as memory usage, network bandwidth usage, etc.
[0020] Referring now to the drawings, example embodiments of the present disclosure will be discussed in more detail. Throughout the present disclosure, embodiments will be described with reference to JPEG and GIF compression, but it should be understood that the systems and methods disclosed herein may additionally utilize other image compression techniques.
[0021] Example Devices and Systems
[0022] Figure 1 A block diagram of an example computing system 100 is depicted that uses a machine learning model trained according to an example embodiment of the present disclosure to perform embedding and extraction of messages. The system 100 includes a first computing device 102 and a second computing device 140 communicatively coupled via a network 180.
[0023] The first computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or a desktop computer), a mobile computing device (e.g., a smartphone or a tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, a personal assistant computing device, or any other type of computing device.
[0024] The first computing device 102 includes one or more processors 104 and a memory 106. The one or more processors 104 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operably connected. The memory 106 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, disks, etc., and combinations thereof. The memory 106 can store data 108 and instructions 110 that are executed by the processor 104 to cause the first computing device 102 to perform operations.
[0025] According to one aspect of the present disclosure, the first computing device 102 can store or include one or more machine learning models. The machine learning model can be or can additionally include one or more neural networks (e.g., deep neural networks), etc. The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks.
[0026] More specifically, a machine learning model can be implemented to provide embedding and extraction of messages within an input image. As an example, the machine learning model can include a machine learning message embedding model 116 and a machine learning message extraction model 118. Specifically, the machine learning message embedding model 116 can receive a difference image that describes the difference between the input image and the first output image that is encoded, and generate a message vector (e.g., a latent space vector) that represents the difference image. The machine learning message embedding model 116 can receive a message and generate a watermark that represents the message vector. The watermark can be applied to the input image or the first output image to generate a second output image. The machine learning message extraction model 118 can obtain the second output image as an input, and extract the message vector from the second output image to obtain the extracted message vector. The machine learning message extraction model 118 can receive the extracted message as an input and provide a reconstruction of the difference image as an output. The reconstructed difference image can be added to the input image to generate a reconstructed input image.
[0027] The first computing device 102 can also include a model trainer 112. The model trainer 112 can use various training or learning techniques, such as, for example, backpropagation of errors (e.g., truncated backpropagation over time), to simultaneously train or retrain machine learning models, such as the machine learning message embedding model 116 and the machine learning message extraction model 118, which are stored at the first computing device 102. In particular, the model trainer 112 can use the training data 114 to simultaneously train or retrain the machine learning message embedding model 116 and the machine learning message extraction model 118. The specific training signals for training or retraining the machine learning models will be discussed in detail in the following figure.
[0028] The model trainer 112 can perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained. Thereafter, the machine learning message embedding model 116 and the machine learning message extraction model 118 can be immediately used to embed and extract messages in images.
[0029] Additionally, in some embodiments, the machine learning message embedding model 116 can include a machine learning watermark generation model and a machine learning message generation model. Similarly, in some embodiments, the machine learning message extraction model 118 can include a machine learning watermark extraction model and a machine learning difference reconstruction model.
[0030] The first computing device 102 may also include one or more input / output interfaces 122. The one or more input / output interfaces 122 may include, for example, devices for receiving information from or providing information to a user, such as a display device, a touch screen, a touch pad, a mouse, data input keys, an audio output device, such as one or more speakers, a microphone, a haptic feedback device, etc. For example, the input / output interfaces 122 may be used by the user to control the operation of the first computing device 102.
[0031] The first computing device 102 may also include one or more communication / network interfaces 124 for communicating with one or more systems or devices, including systems or devices remote from the first computing device 102. The communication / network interfaces 124 may include any circuitry, components, software, etc. for communicating with one or more networks (e.g., network 180). In some embodiments, the communication / network interfaces 124 may include, for example, one or more of a communication controller, a receiver, a transceiver, a transmitter, a port, a conductor, software, and / or hardware for communicating data.
[0032] The second computing device 140 includes one or more processors 142 and a memory 144. The one or more processors 142 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operably connected. The memory 144 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, magnetic disks, etc., and combinations thereof. The memory 144 may store data 146 and instructions 148 that are executed by the processor 142 to cause the second computing device 140 to perform operations.
[0033] As described above, the second computing device 140 may store or otherwise include one or more machine learning models. The machine learning model may be or may otherwise include one or more neural networks (e.g., deep neural networks), and the neural network (e.g., deep neural network) may be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks.
[0034] More specifically, the second computing device 140 may receive and store a trained machine learning model from the first computing device 102 via the network 180, for example. For example, the computing device 140 may receive a machine learning message extraction model 150 (e.g., a machine learning watermark extraction model and a machine learning difference reconstruction model) to provide a reconstruction of an input image from an image having an embedded message transmitted to the computing device 140. The second computing device 140 may use the machine learning model for the same or similar purposes as described above.
[0035] As an example, the second output image can be generated by the machine learning message embedding model 116 and transmitted to the computing device 140 via the network 180 together with the machine learning message extraction model 118. The computing device 140 can use the transmitted machine learning message extraction model 150 to extract a message from the transmitted second output image and generate a reconstructed input image corresponding to the second output image.
[0036] The second computing device 140 can also include one or more input / output interfaces 152. The one or more input / output interfaces 152 can include, for example, devices for receiving information from or providing information to a user, such as a display device, a touch screen, a touchpad, a mouse, data input keys, audio output devices, such as one or more speakers, a microphone, a haptic feedback device, etc. For example, the input / output interface 152 can be used by the user to control the operation of the second computing device 140.
[0037] The second computing device 140 can also include one or more communication / network interfaces 154 for communicating with one or more systems or devices, including systems or devices remote from the second computing device 140. The communication network interface 154 can include any circuitry, components, software, etc. for communicating with one or more networks (e.g., network 180). In some embodiments, the communication / network interface 154 can include, for example, one or more of a communication controller, a receiver, a transceiver, a transmitter, a port, a conductor, software, and / or hardware for communicating data.
[0038] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over the network 180 can be carried via any type of wired and / or wireless connection, using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL).
[0039] Figure 1 An example computing system that can be used to implement the present disclosure is illustrated. Other computing systems can also be used.
[0040] Example model device
[0041] Figure 2FIG. 0 is a flowchart depicting an example method for generating a second output image according to an example embodiment of the present disclosure. The method can be performed by one or more computing devices (e.g., computing device 102, computing device 140, etc.) and includes, for example, obtaining an input image 202. The input image 202 can be a digital image file formatted as a RAW format, a raster format (e.g., a bitmap image file (BMP), a tagged image file format (TIFF), etc.), a vector format (e.g., a computer graphics metafile (CGM), a scalable vector graphics (SVG), etc.), or any other known image file format. The input image 202 can be obtained by one or more computing devices. In some embodiments, one or more computing devices can include one or more sensors (e.g., a digital camera) configured to capture the input image. In other embodiments, one or more computing devices can obtain the input image from another computing device.
[0042] In some embodiments, the obtained input image 202 can include multiple frames. As an example, the input image 202 can be formatted in a format that allows multiple image frames to be included in the image (e.g., a Graphics Interchange Format (GIF), a WebM, a WebP, etc.). As another example, the input image 202 can be formatted in a video format that allows multiple image frames to be included in the image (e.g., an MP4, a VID, an MPEG, an AVI, etc.). It will be apparent to those skilled in the art that the following methods and processes can be applied sequentially or non-sequentially to each of the multiple frames included in the input image 202.
[0043] The input image can be passed through an encoding / decoding scheme 204 to produce a first output image 206. In some embodiments, the encoding scheme applied to the input image 202 can be a differentiable lossy compression scheme. As an example, a differentiable JPEG compression scheme 204 can be used to encode and decode the input image 202 to produce the first output image 206, which is formatted as a JPEG. The resulting first output image 206 encoded as a JPEG may have lost data due to the lossy nature of the compression scheme used. As another example, an input image 202 that includes multiple frames can be encoded and decoded using a differentiable GIF compression scheme 204. Each frame of the resulting first output image 206 formatted as a GIF can lose data due to the lossy nature of the compression scheme used.
[0044] A difference image 208 can be determined that describes the difference between the input image 202 and the first output image 206. More specifically, the difference image 208 can describe the data lost from the encoded / decoded input image 202 to the first output image 206 in the encoding / decoding scheme 204. In some embodiments, the difference image 208 can represent the change in pixel values starting from the encoded first image 202.
[0045] A message vector 212 (e.g., a latent space vector) can be generated at least in part by a machine learning message embedding model based on the difference image 208. In some embodiments, the message vector 212 representing the difference image 208 can be generated using the machine learning message generation model 210 of the machine learning message embedding model. In some embodiments, an autoencoder (e.g., the machine learning message generation model 210) can be used to generate the message vector 212. However, it should be noted that the message vector 212 can be represented in a format different from the latent space vector. Depending on the machine learning model being used, the format of the message can be any type of encoded representation of the difference image 208. By representing the difference image 208 as a message vector that is reduced to its latent space vector representation, the difference image 208 can generally be downsized. In this way, the message can be more easily embedded into the encoded image without significantly increasing the space required to store the encoded image.
[0046] A second output image 216 can be generated by a machine learning embedding model that includes the message vector 212. More specifically, a watermark (e.g., image noise) representing the message vector 212 can be generated and added to the input image 202 to obtain the second output image 216. Alternatively, in some embodiments, the watermark can be added to the first output image 206 to obtain the second output image 216. In some embodiments, the watermark can be generated by the machine learning watermark generation model 214 of the message embedding model. In some embodiments, the watermark can be generated by an autoencoder (e.g., the machine learning watermark generation model 210).
[0047] In some embodiments, a watermark can be applied to an image (e.g., input image 202 or first output image 206) by modifying pixel values associated with the image (e.g., RGB channel values, intensity values, etc.). As an example, the message can be watermarked in the image as image noise (e.g., random variations in brightness and / or color information). As another example, the message can be watermarked as image blur (e.g., blurring of one or more pixel portions of the image). Although image noise and image blur are given as examples, any other form of pixel value modification can be used to watermark the message in the second output image 216. Additionally, it should be noted that the second output image 216 can be encoded in the same manner as the first output image 206 before or after the message is embedded therein.
[0048] In some embodiments, a machine learning message embedding model (e.g., machine learning watermark generation model 214) can take both a message vector 212 and an input image 202 as inputs to generate a second output image 216 embedded with the message. In some cases, including the input image 202 before encoding / decoding it using the encoding / decoding scheme 204 can improve the performance of the machine learning message embedding model by applying the encoding scheme to the input image 202 and simultaneously adding the message vector 212 (e.g., adding a watermark representing the message vector 212) to the input image 202. As an example, the input image 202 can be encoded into a first output image 206 using an encoding scheme, and a message can be generated at least in part based on the difference between the input image 202 and the encoded first output image 206 (e.g., difference image 208). The machine learning message embedding model can embed the message into the input image 202 while also applying the encoding scheme to the input image to produce an encoded second output image 206. In some cases, including the input image 202 in this manner can result in a higher quality second output image 206 (e.g., less visual distortion from the watermark, higher message fidelity, etc.).
[0049] In some embodiments, the encoded second output image 216 can be stored and / or transmitted to another computing device. As an example, one or more computing devices can store the encoded second output image 216 for long - term storage. In this way, the encoded second output image 216 requires much less storage space than the input image 202 while still being able to be reconstructed to a quality level that is the same as or substantially similar to that of the input image 202. As another example, one or more computing devices can transmit the encoded second output image 216 to another computing device. In this way, the transmission of the encoded second output image requires less bandwidth than the transmission of the input image while being able to reconstruct the image to a quality level that is the same as or substantially similar to that of the input image 202.
[0050] Figure 3 A flowchart depicting an example method for reconstructing an input image from a second output image according to an example embodiment of the present disclosure. The method can be performed by one or more computing devices and includes, for example, obtaining a second output image 302. If the second output image 302 is encoded, the encoded second output image may be decoded to produce a decoded second output image. A machine learning message extraction model can extract an embedded message from the second output image 302 to reproduce the extracted message vector 308. In some embodiments, the message can be extracted by a machine learning watermark extraction model of the machine learning message extraction model 306. The machine learning watermark extraction model 306 can be or can otherwise include one or more neural networks (e.g., deep neural networks), etc. The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks. In some cases, at least some data may be lost in the entire embedding and extraction process (e.g., encoding / decoding scheme 204, embedded message vector 212, etc.) for the extracted message vector 308 reproduced by the machine learning watermark extraction model 306. However, such data loss does not necessarily render the extracted message vector 308 inoperable. In the case where the extracted message 308 is rendered inoperable, the use of repeated watermarking techniques can provide redundant messages that can be retained for extraction.
[0051] In some embodiments, the second output image 302 can be evaluated by a discriminator 304 to determine whether the second output image contains an embedded message vector 308. The discriminator 304 can be a machine learning discrimination model (e.g., a generative adversarial network, a linear classifier, a support vector machine (SVM), etc.). The discriminator 304 can be trained to determine whether the second output image 302 contains an embedded message vector 308. As an example, the discriminator 304 can be trained in a supervised manner using training data that includes images known to contain an embedded message vector 308. If the discriminator 304 determines that the second output image 302 contains an embedded message, the machine learning message extraction model (e.g., the machine learning message extraction model 306) can take the second output image 302 as an input.
[0052] In some embodiments, a reconstructed difference image 312 can be generated at least in part based on an embedded message vector 308 extracted by a machine learning message extraction model. More specifically, a machine learning difference reconstruction model 310 of the machine learning message extraction model 306 can reconstruct a difference image based on the extracted message vector 308 to generate a reconstructed difference image 312. As an example, the extracted message vector 308 representing the difference image 208 as a latent space vector can be used as an input to an autoencoder (e.g., the machine learning difference reconstruction model 310) to generate a reconstructed difference image 312. As another example, the extracted embedded message vector 308 representing the difference image 208 as a latent space vector can be used as an input to some other type of neural network architecture (e.g., a feed-forward neural network, a convolutional neural network, etc.) to generate a reconstructed difference image 312.
[0053] A reconstructed input image 314 can be generated at least in part based on the decoded second output image 302 and the reconstructed difference image 312. More specifically, the reconstructed difference image 312 can be added to the decoded second output image 302 to generate a reconstructed input image 314. As an example, each pixel of the decoded second output image 302 can have a pixel value difference from the pixels of the input image 202. The reconstructed difference image 312 including the difference values of each pixel of the decoded output image 302 can be added to the decoded output image 302 pixel by pixel to reconstruct the input image (e.g., generate a reconstructed input image 314).
[0054] The above machine learning models (e.g., machine learning message embedding and extraction models, machine learning watermark generation and extraction models, machine learning message generation and difference reconstruction models, discriminators, etc.) can be trained together or simultaneously (e.g., in a joint manner where one or more gradients are passed from one network to another). Details of training the above machine learning models will be discussed in more detail in subsequent figures.
[0055] Example Method
[0056] Figure 4 is a flowchart depicting an example method of generating a second output image in accordance with an example embodiment. For example, the method 400 can be implemented using Figure 1 a computing device. For purposes of illustration and discussion, Figure 4 the steps are depicted in a particular order. Those of ordinary skill in the art using the disclosures provided herein will understand that the various steps of any method described herein can be omitted, rearranged, executed simultaneously, extended, and / or modified in various ways without departing from the scope of the present disclosure.
[0057] At 402, the method can include obtaining an input image. The input image can be a digital image file formatted in a RAW format, a raster format (e.g., Bitmap Image File (BMP), Tagged Image File Format (TIFF), etc.), a vector format (e.g., Computer Graphics Metafile (CGM), Scalable Vector Graphics (SVG), etc.), or any other known image file format.
[0058] In some embodiments, the input image can include multiple frames. As an example, the input image can be formatted in a format that allows multiple image frames to be included in the image (e.g., Graphics Interchange Format (GIF), WEBM, WEBP, etc.). As another example, the input image can be formatted in a video format that allows multiple image frames to be included in the image (e.g., MP4, VID, MPEG, AVI, etc.).
[0059] In some embodiments, obtaining the input image can include receiving the input image from a computing device. The input image can be transferred from one computing device to another computing device over a network or via a storage medium (e.g., flash storage medium, portable hard drive, etc.). In some embodiments, the input image can be captured by a computing device configured to capture image data using one or more sensors (e.g., digital camera, webcam, etc.).
[0060] At 404, the method can include generating a first output image by encoding and decoding the input image according to an encoding scheme. The encoding scheme can be any differentiable or approximately differentiable encoding scheme. Although primarily referring to image compression schemes, the encoding scheme can also be a non-compressed image and / or video encoding scheme. As an example, a differentiable JPEG compression scheme can be used to encode and decode the input image to produce a first output image that is formatted as JPEG. As another example, a differentiable GIF compression scheme can be used to encode and decode an input image that includes multiple frames.
[0061] Differentiable JPEG encoding schemes have been explored previously in the art. For example, a differentiable approximation of JPEG encoding has been described in "JPEG-resistant adversarial images" (see Richard Shin, Dawn Song, JPEG-Resistant Adversarial Images, NIPS 2017 Workshop on Machine Learning and Computer Security, pages 1-3 (2017)). Similarly, differentiable approximations of other non-differentiable compression schemes (e.g., GIF and other image compression schemes) have been described in "Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations" (see Eirikur Agustsson et al., Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations, Advances in Neural Information Processing Systems, pages 1-9 (2017)). Thus, while differentiable JPEG encoding is used in some examples to illustrate the functionality of the current embodiments, any differentiable or approximately differentiable image encoding scheme can be used. Additionally, if a differentiable or approximately differentiable version or method of an encoding scheme is developed later, any currently non-differentiable encoding scheme can be used.
[0062] At 406, the method can include determining a difference image that describes the difference between the input image and the first output image. The difference image can describe the data loss in encoding / decoding from the input image to the first output image in an encoding / decoding scheme. For example, applying a differentiable lossy JPEG encoding scheme to the input image will necessarily result in at least some data loss on a pixel-by-pixel basis. The difference image can describe the pixel-by-pixel loss of data such that the input image can be reconstructed later. As an example, if the value of the first pixel of the input image before encoding is 15 and the encoded value is 25, the difference image can represent the difference between the pixel values as 10. This change in pixel values can be similarly represented for each pixel of the input image. Although the above example can be used to represent the difference image, those skilled in the art should understand that any representation (e.g., integer, etc.) calculated in any way (e.g., addition, subtraction, etc.) can be used to represent the pixel-by-pixel difference between the input image and the first output image.
[0063] At 408, the method can include using a machine learning message embedding model to generate a second output image that includes an embedded message based at least in part on the difference image. In some embodiments, the second output image can include the first output image with the embedded message representing the difference image. The embedded message representing the difference image can be a latent space vector representation of the difference image. The message vector (e.g., latent space vector representation) can be generated by the machine learning message embedding model. More specifically, the machine learning message generation model of the machine learning message embedding model can generate the message vector. As an example, an autoencoder (e.g., the machine learning message generation model) can be used to generate the message vector representing the difference image. As another example, some other type of neural network architecture (e.g., a feedforward neural network, a convolutional neural network, etc.) can be used to generate the message vector representing the difference image.
[0064] The machine learning message embedding model can generate a watermark based on the message vector. More specifically, the machine learning watermark generation model of the machine learning message embedding model can take the message vector as input to generate a watermark representing the message vector. In some embodiments, the machine learning watermark generation model can additionally take the input image as input. In this way, the machine learning watermark generation model can generate a watermark and apply the watermark to the input image. The machine learning watermark generation model can be or can otherwise include one or more neural networks (e.g., deep neural networks), etc. The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks. As an example, the machine learning watermark generation model can be an autoencoder. As another example, the machine learning watermark generation model can be a convolutional neural network.
[0065] A watermark can be added to the input image and then encoded using an encoding scheme to generate a second output image. Alternatively, in some embodiments, the watermark can be added to the first output image to generate a second output image. In some implementations, the watermark can be added to the image at multiple locations within the image. More specifically, the same watermark representing the message can be repeatedly applied to different locations within the image. As an example, the four corners of the image can be watermarked using the same watermark. As another example, the watermark can be applied to three random locations within the image. Thus, in this way, the repeated watermark can provide redundancy to ensure message delivery in the event that one (or more) of the watermarks is rendered inoperable.
[0066] Although the machine learning message generation model and the machine learning watermark generation model are discussed as machine learning message embedding models, in some implementations, the machine learning embedding model can perform both functions of the above models. As an example, the machine learning message embedding model can be trained to perform the functions of the machine learning message generation model and the machine learning watermark generation model, thereby allowing the machine learning message embedding model to generate a message vector and a watermark representing the message vector.
[0067] Figure 5 is a flowchart depicting an example method of reconstructing an input image from a second output image in accordance with an example embodiment. For example, the method 500 can be implemented using Figure 1 a computing device. For purposes of illustration and discussion, Figure 5 the steps are depicted in a particular order. Those of ordinary skill in the art having access to the disclosures provided herein will understand that the various steps of any method described herein can be omitted, rearranged, executed simultaneously, extended, and / or modified in various ways without departing from the scope of the present disclosure.
[0068] At 502, the method includes obtaining the encoded second output image. The encoded second output image can be an encoded (e.g., compressed) representation of the input image that contains an embedded message (e.g., watermarked) representing data lost from the input image due to the encoding. Any differentiable encoding scheme (e.g., differentiable JPEG compression, differentiable GIF compression, etc.) can be used to encode the encoded second output image. The encoded second output image can be obtained in the same or a similar manner as the input image, as Figure 4 discussed. In some implementations, the encoded second output image can be transmitted to a computing device along with a trained machine learning extraction model. In this way, the embedded message can be extracted from the second output image by the receiving computing device using the trained machine learning extraction model.
[0069] At 504, the method includes decoding an encoded version of the second output image to obtain a decoded version of the second output image. The second output image can be decoded as specified by the encoding scheme used to encode the first output image. In embodiments where the encoding scheme used is a compression scheme, even in the case of containing an embedded message, the decoded second output image will typically use less memory than the input image. However, in some embodiments, the encoding scheme used is non-compressive.
[0070] At 506, the method can include using a machine learning watermark extraction model of the machine learning message embedding model to extract the embedded message from the decoded version of the second output image to obtain the extracted embedded message. More specifically, the machine learning watermark extraction model can take the decoded second output image as input and obtain the extracted embedded information by extracting the watermark from the second output image. The machine learning watermark extraction model can be or can otherwise include one or more neural networks (e.g., deep neural networks), etc. The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks. In some embodiments, the embedded message can be extracted by an autoencoder (e.g., the machine learning watermark extraction model). The extracted embedded message (e.g., the extracted message vector) can be a latent space vector representation of the difference image.
[0071] At 508, the method can include using a machine learning difference reconstruction model of the machine learning message embedding model to reconstruct the difference image from the embedded message. More specifically, the machine learning difference reconstruction model can take the extracted embedded message (e.g., the extracted message vector) as input and then output the reconstructed difference image. The machine learning difference reconstruction model can be or can otherwise include one or more neural networks (e.g., deep neural networks), etc. The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks. In some embodiments, the difference image can be reconstructed by an autoencoder (e.g., the machine learning difference reconstruction model). In certain cases, the reconstructed difference image may have suffered data loss from the encoding and decoding processes. Additionally, the representation of the initial difference image as a latent space vector (e.g., message vector) can also result in data loss. However, this data loss does not necessarily render the reconstructed difference image inoperable.
[0072] Although the machine learning watermark extraction model and the machine learning difference reconstruction model are discussed as machine learning message extraction models, in some embodiments, the machine learning extraction model is capable of performing both functions of the above models. As an example, a machine learning message extraction model can be trained to perform the functions of a machine learning watermark extraction model and a machine learning difference reconstruction model, thereby allowing the machine learning message extraction model to extract a message vector and reconstruct a difference image from the message vector.
[0073] At 510, the method can include generating a reconstruction of the input image based at least in part on a decoded version of the second output image and the reconstruction of the difference image. More specifically, the reconstructed difference image can be added to the decoded second output image on a pixel-by-pixel basis to reconstruct the input image. For example, the input image can have a first pixel value of 5, the second output image can have a corresponding pixel value of 15. The reconstructed difference image can have a corresponding pixel value of 10, and adding the reconstructed difference image to the second output image can produce a corresponding pixel value of 15 for the reconstructed input image. In this way, the input image can be reconstructed pixel by pixel. Although the above example can be used to reconstruct the input image, those skilled in the art should understand that the reconstructed difference image can be used with the second output image in any way (e.g., adding, subtracting, etc.) to reconstruct the input image. It should be noted that in some cases, the corresponding pixel values of the reconstructed difference image may not match the differences of the initial difference image consistently. Thus, the reconstructed difference image can be the same as or substantially similar to the initial difference image.
[0074] Figure 6 is a flow chart depicting an example method for training one or more machine learning models according to an example embodiment. For example, the method 600 can be implemented using Figure 1 a computing device of. For purposes of illustration and discussion, Figure 6 the steps are depicted as being performed in a particular order. Using the disclosures provided herein, those of ordinary skill in the art will understand that the various steps of any method described herein can be omitted, rearranged, performed simultaneously, extended, and / or modified in various ways without departing from the scope of the present disclosure.
[0075] At 602, the method can include using a machine learning message extraction model to extract an embedded message from a decoded version of the second output image to obtain a reconstruction of the difference image. In some embodiments, the machine learning watermark extraction model of the machine learning message extraction model can extract the embedded message, and the machine learning difference reconstruction model of the machine learning message extraction model can reconstruct the difference image. In some alternative embodiments, the machine learning message extraction model can perform the above two functions. The above machine learning models (e.g., machine learning message extraction model, machine learning extraction model, machine learning difference reconstruction model, etc.) can be trained together or simultaneously (e.g., in a joint manner where one or more gradients are passed from one network to another network). For example, by providing a first training signal including a loss function to the machine learning message extraction model (e.g., machine learning watermark extraction and difference reconstruction model), the above networks can be trained simultaneously as a large collection of machine learning models.
[0076] At 604, the method can include evaluating a loss function that evaluates the difference between the input image and the reconstruction of the input image. In this way, the machine learning model can be trained to improve the quality of the reconstructed image. In some embodiments, the machine learning message extraction model can be trained at least in part by the training signal. In some other embodiments, the machine learning watermark extraction model and the machine learning difference reconstruction model can be trained at least in part using the training signal.
[0077] At 606, the method can include further evaluating the loss function to evaluate the difference between the first output image and the second output image. In this way, the machine learning model can be trained to minimize the perceived impact of watermarking the second output image. In some embodiments, the machine learning message extraction model (e.g., machine learning watermark extraction model and machine learning difference reconstruction model) can be trained at least in part using the training signal. In some other embodiments, the machine learning message generation model (e.g., machine learning message generation model and machine learning watermark generation model) can be trained at least in part using the training signal.
[0078] At 608, the method can include modifying the values of one or more parameters of at least the machine learning message extraction model (e.g., machine learning watermark extraction model and machine learning difference reconstruction model) based on the loss function. In some embodiments, the values of one or more parameters of at least the machine learning extraction model are evaluated only based on the loss function of the difference between the input image and the reconstructed input image. The difference can be backpropagated through the machine learning message extraction model (e.g., machine learning watermark extraction and difference reconstruction model) to determine the values associated with one or more parameters of the model to be updated. One or more parameters can be updated to reduce the difference evaluated by the loss function (e.g., using an optimization process such as the gradient descent algorithm).
[0079] Additionally, in some embodiments, the above loss function can be further backpropagated through a machine learning message generation model (e.g., a machine learning message generation model and a machine learning watermark generation model) to determine values associated with one or more parameters of the model to be updated. The one or more parameters can be updated to reduce the discrepancy evaluated by the loss function (e.g., using a gradient descent algorithm). In this way, both the machine learning message embedding and extraction models (and their respective associated models) can be trained to generate more accurate reconstructed images. As an example, the machine learning message generation model can be trained through the loss function to generate a more accurate reconstructed message vector that achieves the difference image.
[0080] In some embodiments, the method can include modifying values of one or more parameters of a machine learning message generation model (e.g., a machine learning message generation model and a machine learning watermark generation model). More specifically, the model can be trained at least in part using a loss function evaluation of the perceptual difference (e.g., perceptual loss) between a first output image and a second output image. The difference can be backpropagated through the machine learning model to determine values associated with one or more parameters of the model to be updated. The one or more parameters can be updated to reduce the difference evaluated by the loss function (e.g., using a gradient descent algorithm). In this way, the model can be trained to reduce the perceptual difference associated with adding a watermark to the second output image.
[0081] Alternatively, in some embodiments, different models of the machine learning message embedding and extraction models can be trained separately. More specifically, the machine learning watermark generation and extraction models can be trained simultaneously and separately from the machine learning message generation and difference reconstruction models. The machine learning message embedding model and the machine learning message extraction model can be trained simultaneously using a training signal that includes a loss function. More specifically, the training signal that includes the loss function can be backpropagated through the models to train both models simultaneously. As an example, the loss function can evaluate a first difference between an input image and a reconstructed input image. Additionally, in some embodiments, the loss function can further evaluate a second difference between the fidelity of the first output image and the fidelity of the second output image. These differences can be backpropagated through the machine learning watermark generation and extraction models to determine values associated with one or more parameters of the model to be updated. The one or more parameters can be updated to reduce the differences evaluated by the loss function (e.g., using a gradient descent algorithm).
[0082] Similarly, the machine learning message generation and machine learning difference reconstruction models can be trained simultaneously and separately. In some embodiments, the embedded message vector is generated and reconstructed using an artificial neural network architecture (e.g., an autoencoder). The neural network architecture can be trained to generate an embedded message vector for embedding in the output image and to reconstruct the difference image from the extracted embedded message (e.g., the message extracted by the machine learning message extraction model). The machine learning message generation and difference reconstruction models can be trained simultaneously using a training signal. More specifically, the training signal including a loss function can be backpropagated through the models to train both models simultaneously. In some embodiments, the loss function can evaluate the difference between the difference image and the reconstructed difference image. The difference can be backpropagated through the neural network architecture to determine values associated with one or more parameters of the model to be updated. The one or more parameters can be updated to reduce the difference evaluated by the loss function (e.g., using a gradient descent algorithm). Thus, in this way, the machine learning message generation and machine learning difference reconstruction models can be trained to generate a message representing the difference image and to reconstruct the difference image from the extracted embedded message in a way that maximizes the image quality.
[0083] In some embodiments, a discriminator can be utilized to determine the perceptual loss between the first output image and the second output image. The discriminator can be a machine learning discriminative model (e.g., a generative adversarial network, a linear classifier, a support vector machine (SVM), etc.). The discriminator can be trained to determine whether the second output image contains an embedded message. As an example, training data including images known to contain embedded messages can be used to train the discriminator in a supervised manner. In some embodiments, the output of the discriminator can be used as a training signal for the machine learning watermark generation model and / or the machine learning message generation model. In this way, the perceptual difference caused by embedding the message into the second output image can be reduced.
[0084] Thus, the systems and methods of the present disclosure significantly reduce the data loss associated with encoded images, particularly when encoded using a lossy compression scheme. Accordingly, the systems and methods of the present disclosure can greatly improve the quality of the encoded images, thus allowing compression to be used in quality-sensitive image storage scenarios that were not operable prior to compression.
[0085] Additional disclosure
[0086] The techniques discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a variety of possible configurations, combinations, and divisions of tasks and functions among and within components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0087] While the subject matter has been described in detail with respect to various specific example embodiments, each example is provided by way of explanation and not limitation of the disclosure. After obtaining an understanding of the foregoing, those skilled in the art will readily conceive of alterations, variations, and equivalents to these embodiments. Accordingly, the subject matter disclosure does not exclude including modifications, variations, and / or additions to the subject matter that are obvious to those of ordinary skill in the art. For example, features illustrated or described as part of one embodiment can be used with another embodiment to yield yet another embodiment. Accordingly, the disclosure is intended to cover such alterations, variations, and equivalents.
[0088] In particular, although Figures 4 - 6 Steps depicted in a particular order for purposes of illustration and discussion, the methods of the disclosure are not limited to the specific order or apparatus recited. The various steps illustrated therein can be omitted, rearranged, combined, and / or adjusted in various ways without departing from the scope of the disclosure. Figures 4 - 6 of the steps.
Claims
1. A computer-implemented method for performing watermark-based image reconstruction to compensate for a lossy coding scheme, the method comprising: obtaining, by one or more computing devices, an input image; generating, by the one or more computing devices, a first output image by compressing and decompressing the input image according to a lossy compression scheme, wherein the first output image comprises a lossy reconstruction of the input image; determining, by the one or more computing devices, a difference image that describes a difference between the input image and the first output image; and processing, by the one or more computing devices, the difference image and the input image using a machine learning message embedding model to generate a second output image as an output of the machine learning message embedding model, wherein the second output image comprises an embedded message that is at least partially based on the difference image.
2. The computer-implemented method according to claim 1, further comprising encoding, by the one or more computing devices, the second output image to obtain an encoded second output image.
3. The computer-implemented method according to claim 2, further comprising storing, by the one or more computing devices, the encoded second output image.
4. The computer-implemented method according to claim 3, further comprising transmitting, by the one or more computing devices, the encoded second output image to another device.
5. The computer-implemented method according to claim 1, wherein generating, by the one or more computing devices, the second output image using the machine learning message embedding model, the second output image comprising the embedded message that is at least partially based on the difference image comprises: generating, by the one or more computing devices, watermark data that is at least partially based on the difference image as an output of a machine learning watermark generation model of the machine learning message embedding model; and adding, by the one or more computing devices, the watermark data to the input image to obtain the second output image.
6. The computer-implemented method according to claim 1, wherein generating, by the one or more computing devices, the second output image using the machine learning message embedding model, the second output image comprising an embedded message that is at least partially based on the difference image comprises: inputting, by the one or more computing devices, both the input image and the difference image into the machine learning message embedding model; and receiving, by the one or more computing devices, the second output image as a direct prediction of the machine learning message embedding model.
7. The computer-implemented method according to claim 1, wherein the embedded message is generated by the machine learning message embedding model at least partially based on the difference image.
8. The computer-implemented method according to claim 1, wherein the second output image comprises at least two instances of the embedded message.
9. The computer-implemented method according to claim 1, wherein the machine learning message embedding model comprises an autoencoder.
10. The computer-implemented method according to claim 1, wherein, the difference image includes the difference of each pixel value between the input image and the first output image.
11. The computer-implemented method according to claim 1, wherein, the embedded message includes a vector in the latent space.
12. The computer-implemented method according to claim 1, wherein, the embedded message is embedded by modifying a plurality of pixel values associated with the pixels of the second output image.
13. The computer-implemented method according to claim 2, further comprising: decoding, by the one or more computing devices, the encoded second output image to obtain a decoded version of the second output image; extracting, by the one or more computing devices, the embedded message from the decoded version of the second output image using a machine learning message extraction model to obtain a reconstruction of the difference image; and generating, by the one or more computing devices, a reconstruction of the input image based at least in part on the decoded version of the second output image and the reconstruction of the difference image.
14. The computer-implemented method according to claim 13, wherein, extracting, by the one or more computing devices, the embedded message from the decoded version of the second output image using the machine learning message extraction model to obtain a reconstruction of the difference image includes: extracting, by the one or more computing devices, the embedded message from the decoded version of the second output image using a machine learning watermark extraction model of the machine learning message extraction model to obtain the extracted embedded message; and reconstructing, by the one or more computing devices, the difference image from the extracted embedded message using a machine learning difference reconstruction model.
15. The computer-implemented method according to claim 13, wherein, decoding, by the one or more computing devices, the encoded second output image to obtain a decoded version of the second output image includes: determining, by the one or more computing devices, using a discriminator that the encoded second output image contains the embedded message; and decoding, by the one or more computing devices, the encoded second output image to obtain a decoded version of the second output image.
16. The computer-implemented method according to claim 1, wherein, the machine learning message embedding model has been trained on a loss function, wherein the loss function evaluates the difference between a training input image and a reconstruction of the training input image generated based on a variant of the training input image that includes an embedded message generated by the message embedding model.
17. The computer-implemented method according to claim 16, wherein, the loss function further evaluates the quality difference between the first output image and the second output image.
18. A computing system, comprising: one or more processors; and instructions for causing the computing system to perform a first set of operations when executed by the one or more processors, the operations including: obtaining an input image; Encode and decode the input image according to an encoding scheme to generate a first output image; Determine a difference image that describes the difference between the input image and the first output image; Generate a second output image using a machine learning message embedding model, the second output image including an embedded message generated by the machine learning message embedding model based at least in part on the difference image; Decode an encoded version of the second output image to obtain a decoded version of the second output image; Extract the embedded message from the decoded version of the second output image using a machine learning message extraction model to obtain a reconstruction of the difference image; Evaluate a loss function that evaluates the difference between the input image and a reconstruction of the input image; and Modify the value of one or more parameters of at least the machine learning message extraction model based on the loss function.
19. The computing system according to claim 18, further comprising: Instructions for a second set of operations that cause the computing system to perform operations when executed by the one or more processors, the operations including: Modify the value of one or more parameters of at least the machine learning message embedding model based on the loss function.
20. The computing system according to claim 18 or claim 19, wherein, The loss function is further evaluated at least in part based on the difference between the first output image and the second output image.
21. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations, the operations including: Obtain an encoded version of an image containing an embedded message; Decode the encoded version of the image to obtain a decoded version of the image; Extract the embedded message from the decoded version of the image using a machine learning message extraction model to obtain a reconstruction of a difference image that describes the difference between an input image and an output image generated by applying an encoding and decoding scheme to the input image; and Generate a reconstruction of the input image based at least in part on the decoded version of the image and the reconstruction of the difference image.
22. The one or more non-transitory computer-readable media according to claim 21, wherein, Generating the reconstruction of the input image includes: adding the difference image to the decoded version of the image.
Citation Information
Patent Citations
Apparatus and method for embedding and extracting digital watermarks based on wavelets
US20030095682A1
Digital watermark image processing method
US6584210B1