Systems and methods for message embedding in three-dimensional image data
By embedding and extracting messages from 3D image data using machine learning models, the challenge of embedding messages in 3D image data is solved. This achieves seamless transmission and efficient encrypted message embedding and extraction, ensuring that messages are imperceptible and suitable for secure transmission of 3D image data.
Patent Information
- Application Number
- CN202080102778.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-05
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2040-06-05
AI Technical Summary
Embedding sufficiently long, imperceptible or nearly imperceptible messages into 3D image data presents challenges, and existing techniques may result in noticeable differences in the 3D image data.
A machine learning message embedding model is used to receive 3D image data and message vectors, generate encoded 3D image data, and embed messages by modifying various aspects of the image data. A machine learning message extraction model is used to extract the embedded messages from the encoded image data, and the model parameters are optimized through a loss function to minimize perceptual differences.
It enables seamless and efficient message embedding and extraction in 3D image data, reduces bandwidth and storage requirements, effectively encrypts embedded messages, and ensures that messages are imperceptible or nearly imperceptible to human observers.
Smart Images

Figure CN115769291B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to message embedding in three-dimensional image data. More specifically, the present disclosure relates to machine learning model(s) for imperceptible or near-imperceptible hidden message embedding in three-dimensional image data. BACKGROUND
[0002] Embedding messages in two-dimensional image data with traditional techniques has been an important area of research. However, embedding messages of sufficient length that are imperceptible or near-imperceptible in three-dimensional image data presents unique challenges. As an example, messages embedded in three-dimensional image data generally must be extractable when rendered or rasterized from any viewpoint. As another example, message embedding techniques can tend to cause noticeable perceptible differences in three-dimensional image data. SUMMARY
[0003] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be apparent from the description, or can be learned through practice of the embodiments.
[0004] One example aspect of the present disclosure relates to a computing system. The computing system can include one or more processors. The computing system can include a machine learning message embedding model. The machine learning message embedding model can be configured to receive three-dimensional image data and a message vector. The machine learning message embedding model can be configured to generate encoded three-dimensional image data based on the three-dimensional image data and the message vector, the encoded three-dimensional image data including an embedded message based on the message vector. The computing system can include a machine learning message extraction model. The machine learning message extraction model can be configured to receive the encoded three-dimensional image data. The machine learning message extraction model can be configured to extract the embedded message from the encoded three-dimensional image data to obtain a reconstructed message vector. The computing system can include a first set of instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining the three-dimensional image data and the message vector. The operations can include inputting the three-dimensional image data and the message vector into the machine learning message embedding model to obtain the encoded three-dimensional image data including the embedded message. The operations can include extracting the embedded message from the encoded three-dimensional image data using the machine learning message extraction model to obtain the reconstructed message vector. The operations can include evaluating a loss function that evaluates a difference between the reconstructed message vector and the message vector. The operations can include modifying values of one or more parameters of at least the machine learning message embedding model based on the loss function.
[0005] Another aspect of the present disclosure relates to a computer-implemented method for watermark-based message embedding for three-dimensional images. The method can include obtaining three-dimensional image data and a message vector. The method can include inputting the three-dimensional image data and the message vector into a machine-learned message embedding model. The method can include receiving, from the machine-learned message embedding model, encoded three-dimensional image data including an embedded message based on the message vector. The method can include extracting the embedded message from the encoded three-dimensional image data using a machine-learned message extraction model to obtain a reconstructed message vector.
[0006] Another aspect of the present disclosure relates to one or more tangible, non-transitory computer-readable media. The one or more tangible, non-transitory computer-readable media can store computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations can include obtaining three-dimensional image data and a message vector. The operations can include inputting the three-dimensional image data and the message vector into a machine-learned message embedding model. The operations can include receiving, from the machine-learned message embedding model, encoded three-dimensional image data including an embedded message based on the message vector. The operations can include extracting the embedded message from the encoded three-dimensional image data using a machine-learned message extraction model to obtain a reconstructed message vector.
[0007] Other aspects of the present disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.
[0008] These and other features, aspects, and advantages of various embodiments of the present disclosure will be better understood when read with reference to the following description and appended claims in conjunction with the accompanying drawings. The accompanying drawings illustrate example embodiments of the present disclosure and, together with the description, serve to explain the related principles. BRIEF DESCRIPTION OF DRAWINGS
[0009] A detailed discussion of embodiments with reference to the drawings will be made in the description, which makes reference to the accompanying drawings, in which:
[0010] FIG. 1A A block diagram depicting an example computing system that performs message embedding, message extraction, and / or training of machine-learned model(s) in accordance with example embodiments of the present disclosure.
[0011] FIG. 1B A block diagram depicting an example computing device that performs message embedding and / or extraction in accordance with example embodiments of the present disclosure.
[0012] FIG. 1C A block diagram depicting an example user computing device that performs message embedding and / or extraction in accordance with example embodiments of the present disclosure.
[0013] FIG. 2is a block diagram depicting an example machine learning message embedding model according to example embodiments of the present disclosure.
[0014] FIG. 3 is a block diagram depicting an example machine learning message embedding and extraction model according to example embodiments of the present disclosure.
[0015] FIG. 4A is a dataflow diagram depicting an example training method for a viewpoint- independent machine learning message embedding model and a viewpoint-independent machine learning message extraction model according to example embodiments of the present disclosure.
[0016] FIG. 4B is a dataflow diagram depicting an example training method for a viewpoint-dependent machine learning message embedding model and a viewpoint-dependent machine learning message extraction model according to example embodiments of the present disclosure.
[0017] FIG. 5 is an example method for viewpoint-dependent message embedding in three-dimensional image data according to example embodiments of the present disclosure.
[0018] FIG. 6 is a flowchart depicting an example method for performing end-to-end training of machine learning message embedding and extraction models according to example embodiments of the present disclosure.
[0019] Reference numerals repeated in multiple figures are intended to identify identical features in various implementations. DETAILED DESCRIPTION
[0020] SUMMARY
[0021] In general, the present disclosure relates to systems and methods for watermark message embedding in three-dimensional image data. More specifically, the systems and methods described herein relate to embedding a message in three-dimensional image data using a machine learning message embedding model, where the machine learning message embedding model has been trained to imperceptibly embed a message in image data. Later, the embedded message can be extracted from the three-dimensional image data using a machine learning message extraction model. Thus, as one example, a message vector (e.g., generated with a traditional message generation algorithm, as a latent space vector from a machine learning model, etc.) can be obtained and input to a machine learning message embedding model along with three-dimensional image data (e.g., a 3D mesh, a 3D volumetric representation, a point cloud, 3D mesh textures, 3D materials, etc.). The machine learning message embedding model (e.g., a neural network, a 3D convolutional neural network, a graph convolutional neural network, etc.) can generate encoded three-dimensional image data that includes the message vector as an embedded message. The message vector can be embedded by modifying various aspects of the three-dimensional image data in an imperceptible or minimally perceptible manner (e.g., modifying colors and / or positions of mesh vertices, colors of textures, labels of point cloud points, etc.). The difference between the three-dimensional image data and the encoded three-dimensional image data (e.g., perceptual difference, etc.) can be used as a loss signal to train the machine learning message embedding model. In this way, the machine learning message embedding model can be trained to hide a message within three-dimensional image data in a manner that is imperceptible or nearly imperceptible to a human observer. The proposed technology represents a significant advance in three-dimensional message embedding. Specifically, by obtaining a message and embedding it within three-dimensional image data without imperceptibly modifying the three-dimensional image data, the proposed systems provide a method for securely including and obfuscating proprietary data (e.g., identification information, decryption key(s), authentication information, location information, etc.) within three-dimensional image data.
[0022] More specifically, a computing system (e.g., one or more computing devices, a distributed network of computing devices, etc.) can obtain three-dimensional image data. The three-dimensional image data can be any kind of three-dimensional image, data, image data, three-dimensional volume, three-dimensional representation, three-dimensional mapping data, point cloud data (e.g., from a LIDAR system, etc.), and / or any material (e.g., texture, etc.) associated with the three-dimensional image data (e.g., a texture associated with a three-dimensional mesh representation, etc.). As an example, the three-dimensional image data can be a 3D mesh (e.g., a polygon mesh, etc.) and / or any material (e.g., a texture, a 2D map, an opacity map, a color map, a roughness map, a bidirectional reflectance distribution function (BRDF), a bidirectional scattering distribution function (BSDF), a bidirectional scattering surface reflectance distribution function (BSSRDF), etc.) associated with the 3D mesh. As another example, the three-dimensional image data can be a point cloud (e.g., a plurality of points in space, etc.), an encoded point cloud (e.g., a point cloud encoded according to an encoding scheme, etc.), and / or any labels associated with points in the point cloud.
[0023] The computing system can obtain a message vector. The message vector can be a portion of data that includes a message to be embedded in the three-dimensional image data. It should be noted that the message vector need not be a vector or vector-like data structure. Instead, the message vector can be any kind of data that contains a message. As an example, the message vector can represent the message as a vector of bits. As another example, the message vector can be a latent space vector (e.g., from an encoder model, etc.). As another example, the message can be a two- or three-dimensional array. Thus, the machine-learned message embedding model can utilize message vectors of any form and / or size to embed in the three-dimensional image data.
[0024] The computing system can input the three-dimensional image data and the message vector into the machine-learned message embedding model. The machine-learned message embedding model can receive the three-dimensional image data and the message vector and generate encoded three-dimensional image data based on the three-dimensional image data and the message vector. The encoded three-dimensional image data can include an embedded message based on the message vector. The machine-learned message embedding model can be or can otherwise include one or more neural networks (e.g., deep neural networks), etc. The neural networks (e.g., deep neural networks) can be feedforward neural networks, convolutional neural networks, and / or various other types of neural networks. In some implementations, the machine-learned message embedding model can be or can otherwise include a conditional variational autoencoder.
[0025] The embedded message can be embedded by modifying one or more aspects of the three-dimensional image data. As an example, if the three-dimensional image data is or otherwise includes a 3D mesh, the embedded message can be embedded by the machine-learned message embedding model by modifying an aspect(s) of the vertex(es) of the 3D mesh (e.g., vertex position, vertex color, material of the 3D mesh, etc.). As another example, if the three-dimensional image data is or otherwise includes one or more 3D volumes, the embedded message can be embedded by the machine-learned message embedding model by modifying an aspect(s) of the 3D volume(s) (e.g., volume value(s), voxel value(s), etc.). As yet another example, if the three-dimensional image data is or otherwise includes a material associated with a 3D representation, the embedded message can be embedded by the machine-learned message embedding model by modifying an aspect(s) of the material(s) of the 3D representation (e.g., BSSRDF(s), BSDF(s), BRDF(s), texture color values, height map, transparency map, texture color(s), etc.). As yet another example, if the three-dimensional image data is or otherwise includes a point cloud, the embedded message can be embedded by the machine-learned message embedding model by modifying an aspect(s) of the point cloud (e.g., point position(s), point value(s), point label(s), etc.).
[0026] The embedded message based on the message vector can be any type or amount of information. As an example, the embedded message can include a decryption key. The decryption key can correspond to an encrypted aspect of the three-dimensional image data and / or any other encrypted information. As another example, the embedded message can include private information. For another example, the embedded message in the point cloud of the previous example can be further encrypted (e.g., to protect patient privacy, etc.), and the hidden message can include a decryption key to decrypt the 3D data itself and / or the remaining portion of the embedded message. Example 3D data can include 3D models (e.g., mesh models), medical imaging data (e.g., MRI data), LIDAR point clouds, RADAR data, and / or other forms of 3D data.
[0027] As another example, the embedded message can include authentication information. The authentication information can be configured to authenticate the source of the three-dimensional image data. For example, the authentication information can be or otherwise include a cryptographic hash, code, and / or key associated with the creator and / or source of the three-dimensional image data. For another example, the authentication information can identify the three-dimensional image data as authentic (e.g., not pirated, stolen, copied, etc.). As another example, the embedded message can include location information. The location information can describe a sending location of the three-dimensional image data, a receiving location of the three-dimensional image data, or both. For example, the location data can include an IP address associated with the sender of the three-dimensional image data. For another example, the location information can include geographic location coordinates corresponding to the sender of the three-dimensional image data.
[0028] In some implementations, the encoded three-dimensional image data can be projected, rendered, or rasterized into a lower dimension representation of the encoded three-dimensional image data. The lower dimension projection of the encoded three-dimensional image data can correspond to a viewpoint (e.g., camera position, viewpoint, etc.). As an example, if the encoded three-dimensional image data is or otherwise includes a point cloud, the points in the point cloud can be projected to a lower dimension space (e.g., 2D projection, 2.5D projection with depth data, etc.) corresponding to a lower dimension viewpoint. As another example, if the encoded three-dimensional image data can be rendered or rasterized to a lower dimension (e.g., 3D mesh, 3D volume, etc.), the 3D mesh can be rendered or rasterized to a lower dimension space (e.g., 2D frame for a video game to be displayed on a display device, etc.) using a rendering scheme or a rasterization scheme.
[0029] In some implementations, a rendering scheme or rasterization scheme for projecting the encoded three-dimensional image data can include one or more camera parameters. The camera parameter(s) can determine what is included in the lower-dimensional representation of the encoded three-dimensional image data. As an example, one or more implicit and / or one or more explicit camera parameters can be included to determine various aspects of a camera associated with rendering or rasterizing the encoded three-dimensional image data. As another example, camera position coordinates can be included that describe a position of the camera in three-dimensional space. For instance, the encoded three-dimensional image data can include a 3D polygon mesh representation of a vehicle. The camera position coordinates can describe a position of the camera such that the camera is in front of the car. The rendering scheme can render the encoded three-dimensional image data according to the camera parameter(s) such that the 2D representation of the encoded three-dimensional image data depicts the front of the car. Moreover, if the camera position coordinates describe a position of the camera behind the vehicle, then the 2D representation can depict the rear of the car. Similarly, other camera parameter(s) (e.g., camera aspect ratio, camera angle, camera field of view, etc.) can be included to further specify what is captured in the lower-dimensional representation of the encoded three-dimensional image data.
[0030] In some implementations, the embedded message can be embedded in the encoded three-dimensional image data such that the embedded message can be extracted from any lower-dimensional representation of the image data. As an example, to further the previous example of lower-dimensional representations of a car from a front side and a back side, the embedded message can be extracted from the lower-dimensional representation of the front side and / or the lower-dimensional representation of the back side. In this way, any projection, rendering, or rasterization of the encoded three-dimensional image data based on any camera parameter(s) can include the embedded message for extraction (e.g., by a machine learning message extraction model, etc.).
[0031] The computing system can receive the encoded three-dimensional image data including the embedded message from the machine learning message embedding model. In some implementations, a first computing device of the computing system can receive the encoded three-dimensional image data from the machine learning message embedding model and send the encoded three-dimensional image data to a second computing device of the computing system (e.g., via a network, etc.). Alternatively, in some implementations, the computing system can send the encoded three-dimensional image data to a second computing system different from the first computing system (e.g., via a network, etc.). In this way, the computing system can generate the encoded three-dimensional image data at a first location (e.g., a first computing device of the computing system, a first computing system, etc.) and send the encoded three-dimensional image data for decoding at a second location (e.g., a second computing device of the computing system, a second computing system, etc.) to facilitate sending a hidden (e.g., embedded) message to a recipient.
[0032] The computing system can use a machine-learned message extraction model to extract an embedded message from the encoded three-dimensional image data to obtain a reconstructed message vector. The machine-learned message extraction model can be or can otherwise include one or more neural networks (e.g., a deep neural network), among other possibilities. The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks. In some implementations, the machine-learned message extraction model and the machine-learned message embedding model can be or can otherwise be components of an autoencoder architecture (e.g., autoencoder(s), variational autoencoder(s), etc.). As an example, the machine-learned message embedding model and / or the machine-learned message extraction model can be a conditional variational autoencoder. As another example, the combination of both the machine-learned message embedding model and the machine-learned message extraction model can be a conditional variational autoencoder.
[0033] The machine-learned message extraction model can be used to extract a reconstructed message vector from the encoded three-dimensional image data. More specifically, the machine-learned message extraction model can receive the encoded three-dimensional image data and extract an embedded message from the encoded three-dimensional image data to obtain a reconstructed message vector. In some implementations, the reconstructed message vector can be a lossless reconstruction of the message vector (e.g., an identical or substantially similar instance of the message vector). Alternatively, in some implementations, the reconstructed message vector can be a lossy reconstruction of the message vector. The data loss associated with the reconstruction of the message vector can vary depending on a number of factors (e.g., the type of three-dimensional image data, the size of the message, the type of message, etc.).
[0034] In some implementations, the reconstructed message vector can be extracted from a lower-dimensional representation projected from the autoencoded three-dimensional image data. Reference will be made to FIG. 5 The extraction from the lower-dimensional representation of the encoded three-dimensional image data is discussed in more detail.
[0035] In some implementations, the machine-learned message extraction model can also output the encoded three-dimensional image data. As an example, the output of the machine-learned message extraction model can be the reconstructed message vector and the encoded three-dimensional image data (e.g., the three-dimensional image data including the embedded message). Alternatively, in some implementations, the machine-learned message extraction model can decode the encoded three-dimensional image data to output the three-dimensional image data (e.g., remove the modifications to aspects of the image data used to embed the embedded message).
[0036] The computing system can evaluate a loss function that evaluates a difference between the reconstructed message vector and the message vector. More specifically, the loss function can evaluate a reconstruction error associated with the embedding and extraction of messages from the three-dimensional image data (e.g., a degree of encoding / decoding loss). Additionally or alternatively, in some implementations, the loss function can further evaluate a difference between the three-dimensional image data and the encoded three-dimensional image data. More specifically, the loss function can evaluate a perceptual difference resulting from modifying aspects of the three-dimensional image data to embed the message.
[0037] The computing system can modify values of one or more parameters of at least the machine-learned message embedding model based on the loss function. Additionally, in some implementations, the computing system can also modify values of one or more parameters of the machine-learned message extraction model. As an example, the loss function can backpropagate through the machine-learned message embedding model and the message extraction model to determine values associated with one or more parameters of the models to update. The one or more parameters can be updated to reduce the difference evaluated by the loss function (e.g., using an optimization process such as a gradient descent algorithm). Thus, in this way, in some implementations, the evaluation of the loss function can minimize the difference from embedding and extracting messages from the three-dimensional image data (e.g., loss of data) while also minimizing the difference resulting from embedding the message in the three-dimensional image data (e.g., perceptual difference), thus providing a highly accurate but imperceptible or nearly imperceptible message embedding.
[0038] In some implementations, prior to extracting the embedded message using the machine-learned message extraction model, the computing system can distort the encoded three-dimensional image data with one or more distortion effects. More specifically, the computing system can apply one or more distortion effects to the encoded three-dimensional image data while training the model(s) to make the model(s) more robust to common distortion effects applied to the three-dimensional image data during use of the trained model(s). The distortion effect(s) can include image noise, image rotation, image simplification, image data refinement, image cropping, encoding loss (e.g., loss of data associated with a lossy encoding scheme), or any other kind of distortion effect(s). Distortion of the encoded three-dimensional image data will be discussed in more detail with reference to FIG. 4.
[0039] In some implementations, in training the model(s) (e.g., evaluating the loss function and modifying the value(s) of the parameter(s) of the model(s)), the computing system can project the encoded three-dimensional image data to a lower dimensional representation with a differentiable projection, rendering, and / or rasterization scheme. In this way, the differentiable projection, rendering, and / or rasterization scheme can allow for adjusting the value(s) of the parameter(s) of the model(s) with backpropagation and / or gradient descent algorithm(s) (e.g., stochastic gradient descent, etc.).
[0040] It should be noted that any contemporary differentiable rendering scheme can be utilized to project the encoded three-dimensional image data to a lower dimensional representation during training. As an example, differentiable Monte Carlo ray tracing through edge sampling can be used as a differentiable rendering scheme (see Tzu-Mao Li et al., “Differentiable Monte Carlo Ray Tracing through Edge Sampling,” ACM Trans. Graph., (ACM), Vol. 37, No. 6, Article 222 (November 2018)). As another example, deep convolutional network(s) can be used as a differentiable rendering scheme (see Thu Nguyen-Phuoc, Chuan Li, Stephen Balaban, Yong-Liang Yang, “RenderNet: A deep convolutional network for differentiable rendering from 3D shapes,” In Proceedings of the 32nd International Conference on Neural Information Processing Systems, (NeurIPS 2018), pp. 7891-7901 (November 2018)). As another example, soft rasterization can be used as a differentiable rasterization scheme (see Shichen Liu, Weikai Chen, Tianye Li, Hao Li, “Soft Rasterizer: Differentiable Rendering for Unsupervised Single-View Mesh Reconstruction,” arXiv preprint, (arXiv), arXiv:1901.05567 (2019)). As another example, neural radiance fields can be used as a differentiable rendering scheme (see Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng, “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” arXiv preprint, (arXiv), arXiv:2003.08934 (2020)). In this manner, according to the present embodiments, any contemporary, state-of-the-art, or future differentiable rendering and / or rasterization scheme can be used to provide differentiable rendering and / or rasterization.
[0041] The present disclosure provides a number of technical effects and benefits. As one example technical effect and benefit, the systems and methods of the present disclosure enable the utilization of message embeddings in three-dimensional image data without significant perceptual loss, which in turn allows for seamless and efficient transmission of three-dimensional image data. As an example, instead of requiring additional transmission of data, authentication material (e.g., license data, etc.) can be included in the three-dimensional image data. Thus, by including messages in the three-dimensional image data, the systems and methods of the present disclosure can significantly reduce bandwidth and memory usage related to transmission of three-dimensional image data. As another example, the embedded messages can include private information. Without the use of a corresponding machine learning message extraction model, it is difficult to extract the embedded messages from the three-dimensional image data, effectively encrypting the embedded messages in the three-dimensional image data.
[0042] Example embodiments of the present disclosure will now be discussed in greater detail.
[0043] Example devices and systems
[0044] FIG. 1A A block diagram of an example computing system 100 that uses a trained machine learning model to perform embedding and extraction of messages in accordance with example embodiments of the present disclosure is depicted. The system 100 includes a first computing device 102 and a second computing device 140 communicatively coupled over a network 180.
[0045] The first computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, a personal assistant computing device, or any other type of computing device.
[0046] The first computing device 102 includes one or more processors 104 and a memory 106. The one or more processors 104 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, a FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 106 can include one or more non-transitory computer-readable storage media, such as, for example, RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 106 can store data 108 and instructions 110 that are executed by the processor 104 to cause the first computing device 102 to perform operations.
[0047] According to an aspect of the disclosure, the first computing device 102 can store or include one or more machine learning models. The machine learning models can be or can otherwise include one or more neural networks (e.g., deep neural networks), etc. The neural networks (e.g., deep neural networks) can be feedforward neural networks, convolutional neural networks, and / or various other types of neural networks. In some implementations, the machine learning models can be or otherwise include one or more conditional variational autoencoders.
[0048] More specifically, the machine learning models can be implemented to provide embedding and extraction of message vector(s) within three-dimensional image data. As one example, the machine learning models can include a machine learning message embedding model 116 and a machine learning message extraction model 118. Specifically, the machine learning message embedding model 116 can receive three-dimensional image data (e.g., 3D mesh, volume, point cloud, etc.) and a message vector. The machine learning message embedding model 116 can generate encoded three-dimensional image data that includes an embedded message based on the message vector. The machine learning message extraction model 118 can obtain the encoded three-dimensional image data as input and extract the message vector from the encoded three-dimensional image data to obtain a reconstructed message vector.
[0049] The first computing device 102 can also include model trainer(s) 112. The model trainer(s) 112 can use various training or learning techniques, such as, for example, backpropagation of error (e.g., truncated backpropagation through time), to simultaneously train or retrain machine learning models, such as the machine learning message embedding model 116 and the machine learning message extraction model 118, stored at the first computing device 102 using training data 114. Specifically, the model trainer(s) 112 can simultaneously train or retrain the machine learning message embedding model 116 and the machine learning message extraction model 118 using the training data 114. Particular training signals for training or retraining the machine learning models will be discussed in depth in the following figures. In some implementations, the training data 114 can also include differentiable projection schemes (e.g., differentiable rendering schemes and / or differentiable rasterization schemes) to facilitate the use of backpropagation and gradient descent algorithms during training of the model(s).
[0050] The model trainer(s) 112 can perform a variety of generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization capabilities of the models being trained. Thereafter, the machine learning message embedding model 116 and the machine learning message extraction model 118 can be immediately used to embed and extract messages in images.
[0051] The first computing device 102 can also include one or more input / output interfaces 120. The one or more input / output interfaces 120 can include, for example, devices for receiving information from or providing information to a user, such as a display device, a touchscreen, a touchpad, a mouse, data entry keys, an audio output device such as one or more speakers, a microphone, a haptic feedback device, etc. For example, a user can use the input / output interface(s) 120 to control the operation of the first computing device 102.
[0052] The first computing device 102 can also include one or more communication / network interfaces 122 for communicating with one or more systems or devices, including systems or devices that are remote from the first computing device 102. The communication / network interface(s) 122 can include any circuitry, components, software, etc. for communicating with one or more networks (e.g., the network 180). In some implementations, the communication / network interface(s) 122 can include, for example, one or more of a communication controller, a receiver, a transceiver, a transmitter, a port, a conductor, software, and / or hardware for communicating data.
[0053] The second computing device 140 includes one or more processors 142 and a memory 144. The one or more processors 142 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, a FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 144 can include one or more non-transitory computer-readable storage media, such as
[0054] As described above, the second computing device 140 can store or otherwise include one or more machine learning models. The machine learning models can be or can otherwise include one or more neural networks (e.g., deep neural networks), and the neural networks (e.g., deep neural networks) can be feedforward neural networks, convolutional neural networks, and / or various other types of neural networks.
[0055] More specifically, the second computing device 140 can receive and store the trained machine learning models from the first computing device 102, e.g., via the network 180. For example, the computing device 140 can receive the machine learning message extraction model 150 to extract embedded messages in encoded three-dimensional image data sent to the computing device 140. The second computing device 140 can use the machine learning model(s) for the same or similar purposes as described above.
[0056] As an example, encoded three-dimensional image data including an embedded message can be generated by the machine-learned message embedding model 116 and transmitted to the computing device 140 via the network 180 along with the machine-learned message extraction model 118. The computing device 140 can use the transmitted machine-learned message extraction model 150 to extract the message from the transmitted encoded three-dimensional image data and generate a reconstructed message vector that corresponds to the message vector at the first computing device 102.
[0057] The second computing device 140 can also include one or more input / output interfaces 152. The one or more input / output interfaces 152 can include, for example, devices for receiving information from or providing information to a user, such as a display device, a touchscreen, a touchpad, a mouse, data entry keys, an audio output device such as one or more speakers, a microphone, a haptic feedback device, etc. For example, a user can use the input / output interface(s) 152 to control the operation of the second computing device 140.
[0058] The second computing device 140 can also include one or more communication / network interfaces 154 for communicating with one or more systems or devices, including systems or devices that are remote from the second computing device 140. The communication / network interface(s) 154 can include any circuitry, components, software, etc. for communicating with one or more networks (e.g., the network 180). In some implementations, the communication / network interface(s) 154 can include, for example, one or more of a communication controller, a receiver, a transceiver, a transmitter, a port, a conductor, software, and / or hardware for communicating data.
[0059] The network 180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. In general, communications over the network 180 can be carried using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML) and / or protection schemes (e.g., VPN, secure HTTP, SSL) via any type of wired and / or wireless connection.
[0060] FIG. 1A An example computing system that can be used to implement the present disclosure is illustrated. Other computing systems can also be used.
[0061] FIG. 1B A block diagram of an example computing device 10 that performs message embedding and / or extraction in accordance with example embodiments of the present disclosure is depicted. The computing device 10 can be a user computing device or a server computing device.
[0062] The computing device 10 includes a plurality of applications (e.g., applications 1 through N). Each application includes its own machine learning library and machine learning model(s). For example, each application can include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.
[0063] As shown in FIG. 1B each application can communicate with a plurality of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.
[0064] FIG. 1C A block diagram of an example computing device 50 that performs message embedding and / or extraction in accordance with example embodiments of the disclosure is depicted. The computing device 50 can be a user computing device or a server computing device.
[0065] The computing device 50 includes a plurality of applications (e.g., applications 1 through N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some implementations, each application can communicate with the central intelligence layer (and the model(s) stored therein) using an API (e.g., a public API that spans all applications).
[0066] The central intelligence layer includes a plurality of machine learning models. For example, as shown in FIG. 1C a respective machine learning model (e.g., a machine learning message embedding model, a machine learning message extraction model, and the like) can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model (e.g., a single model) for all applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of the computing device 50.
[0067] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As shown in FIG. 1C the central device data layer can communicate with a plurality of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0068] Example model arrangement
[0069] FIG. 2 A block diagram depicting an example machine-learned message-embedding model 202 in accordance with example embodiments of the present disclosure is depicted. In some implementations, the machine-learned message-embedding model 202 is trained to receive a set of input data 204 describing a three-dimensional image and a message vector 205, and provide, as a result of receiving the input data 204, output data 206 describing an encoded three-dimensional image that includes an embedded message based on the message vector 205. Thus, in some implementations, the machine-learned message-embedding model 202 can be operable to receive three-dimensional image data 204 and a message vector 205, and generate, based on the three-dimensional image data 204 and the message vector 205, encoded three-dimensional image data 206. The encoded three-dimensional image data 206 can include an embedded message based on the message vector 205.
[0070] The embedded message can be embedded in the encoded three-dimensional image data 206 by modifying one or more aspects of the three-dimensional image data 204. As an example, if the three-dimensional image data 204 is or otherwise includes a 3D mesh, the embedded message can be embedded by the machine-learned message-embedding model 202 by modifying an aspect(s) of vertex(es) of the 3D mesh (e.g., vertex position, vertex color, etc.). As another example, if the three-dimensional image data 204 is or otherwise includes one or more 3D volumes, the embedded message can be embedded by the machine-learned message-embedding model 202 by modifying an aspect(s) of the 3D volume(s) (e.g., volume value, voxel size, voxel position, etc.). As yet another example, if the three-dimensional image data 204 is or otherwise includes materials associated with a 3D representation, the embedded message can be embedded by the machine-learned message-embedding model 202 by modifying an aspect(s) of the material(s) of the 3D representation (e.g., BSSRDF(s), BSDF(s), BRDF(s), texture color values, height map, transparency map, texture color(s), etc.). As yet another example, if the three-dimensional image data 204 is or otherwise includes a point cloud, the embedded message can be embedded by the machine-learned message-embedding model 202 by modifying an aspect(s) of the point cloud (e.g., point position(s), point value(s), point label(s), etc.).
[0071] The embedded message based on the message vector 205 can be any type or amount of information. As an example, the embedded message can include a decryption key. The decryption key can correspond to an encryption aspect of the three-dimensional image data 204 and / or any other encrypted information. As another example, the embedded message can include private information. For another example, the embedded message in the point cloud of the previous example can be further encrypted (e.g., to protect patient privacy, etc.), and the hidden message can include a decryption key to decrypt the 3D data itself and / or the remaining portion of the embedded message. Example 3D data can include 3D models (e.g., mesh models), medical imaging data (e.g., MRI data), LIDAR point clouds, RADAR data, and / or other forms of 3D data.
[0072] FIG. 3 A block diagram depicting an example machine learning message embedding and extraction model 300 in accordance with example embodiments of the present disclosure. The machine learning message embedding and extraction model 300 is similar to the machine learning message embedding model 202 of FIG. 2 FIG. 2, except that the machine learning message embedding and extraction model 300 further includes a machine learning message extraction model 302. More specifically, the three-dimensional image data 204 and the message vector can be received by the machine learning message embedding model 202. The machine learning message embedding model 202 can output encoded three-dimensional image data 206, which can be received by the machine learning message extraction model 302. The machine learning message extraction model 302 can generate a reconstructed message vector 306 by extracting the embedded message from the encoded three-dimensional image data 206.
[0073] More specifically, the machine learning message extraction model 302 can extract the embedded message from the encoded three-dimensional image data 206 to obtain the reconstructed message vector 306. The machine learning message extraction model 302 can be or can otherwise include one or more neural networks (e.g., deep neural networks), etc. The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks. In some implementations, the machine learning message extraction model 302 and the machine learning message embedding model 202 can be or can otherwise be components of an autoencoder architecture (e.g., autoencoder(s), variational autoencoder(s), etc.). As an example, the machine learning message embedding model 202 and / or the machine learning message extraction model 302 can be a conditional variational autoencoder. As another example, the combination of both the machine learning message embedding model 202 and the machine learning message extraction model 302 can be a conditional variational autoencoder.
[0074] A machine-learned message extraction model 302 can be used to extract a reconstructed message vector 306 from the encoded three-dimensional image data 206. More specifically, the machine-learned message extraction model 302 can receive the encoded three-dimensional image data 206 and can extract an embedded message from the encoded three-dimensional image data 206 to obtain the reconstructed message vector 306. In some implementations, the reconstructed message vector 306 can be a lossless reconstruction of the message vector 205 (e.g., the same or substantially similar instance of the message vector 205). Alternatively, in some implementations, the reconstructed message vector 306 can be a lossy reconstruction of the message vector 205. The data loss associated with the reconstruction of the reconstructed message vector 306 can vary depending on a number of factors (e.g., the type of three-dimensional image data, the size of the message, the type of message, etc.).
[0075] FIG. 4A is a dataflow graph depicting an example training method of a viewpoint- independent machine-learned message embedding model and a viewpoint-independent machine-learned message extraction model in accordance with example embodiments of the present disclosure. More specifically, the machine-learned message embedding model 406 can receive three-dimensional image data 404 and a message vector 402. Based on the three-dimensional image data 404 and the message vector 402, the machine-learned message embedding model can output encoded three-dimensional image data 408.
[0076] In some implementations, the encoded three-dimensional image data 408 can be applied with distortion(s) 410 to obtain encoded three-dimensional image data with distortion(s) 412. More specifically, one or more distortion effects 410 can be applied to the encoded three-dimensional image data 408 during training of the model(s) (e.g., the machine-learned message embedding model 406 and the machine-learned message extraction model 418, etc.) to make the model(s) 406 and 418 more robust to distortions that are typically applied to the encoded three-dimensional image data 408 when the encoded three-dimensional image data 408 is manipulated by the user(s) and / or transmitted between the computing device(s). As an example, a distortion effect 410 can be applied to the encoded three-dimensional image data 408 that mimics image noise that occurs due to transmission of the encoded three-dimensional image data 408 and / or encoded data loss of the encoding of the encoded three-dimensional image data 408. As another example, a distortion effect 410 can be applied that mimics image rotation, simplification, refinement, or cropping that a user can apply to the encoded three-dimensional image data 408 in typical image data use case scenarios.
[0077] It should be noted that, in some implementations, by applying the distortion(s) 410 to the encoded three-dimensional image data 408, the model(s) 406 and 418 can be trained to embed the message 402 in the encoded three-dimensional image data 408 such that the encoded three-dimensional image data 408 is resistant to the distortion effects. More specifically, the model(s) 406 and 418 can be trained to embed and extract the message 402 such that the message 402 can be retrieved (e.g., extracted by the machine-learned message extraction model 418) regardless of whether the encoded three-dimensional image data 408 has been severely cropped, simplified, refined, rotated, or distorted in any other way. As an example, a distortion 410 can be applied that crops important portions of the encoded three-dimensional image data 408. In such a manner, through training iterations, the machine-learned message embedding model 406 can be trained to embed the message 402 in a manner that is resistant to cropping (e.g., embed multiple instances of the message 402 at multiple locations of the encoded three-dimensional image data 408, and so on).
[0078] The machine-learned message extraction model 418 can receive the encoded three-dimensional image data 408 (or the distorted encoded three-dimensional image data 412) and, based on the image data, extract a reconstructed message vector 420. More specifically, the reconstructed message vector 420 can be extracted from the encoded three-dimensional image data 408 using the machine-learned message extraction model 418. In some implementations, the reconstructed message vector 420 can be a lossless reconstruction of the message vector 402 (e.g., the same or substantially similar instance of the message vector 402). Alternatively, in some implementations, the reconstructed message vector 420 can be a lossy reconstruction of the message vector 402. The data loss associated with the reconstruction of the message vector 402 can vary depending on a number of factors (e.g., the type of three-dimensional image data 408, the size of the message 402, the type of message 402, and so on).
[0079] In some implementations, the machine-learned message extraction model 418 can also output the encoded three-dimensional image data 408. As an example, the output of the machine-learned message extraction model 418 can be the reconstructed message vector 420 and the encoded three-dimensional image data 408 (e.g., the three-dimensional image data that includes the embedded message). Alternatively, in some implementations, the machine-learned message extraction model 418 can decode the encoded three-dimensional image data 408 to output the three-dimensional image data 404 (e.g., remove the modifications to aspects of the image data 404 used to embed the message vector 402).
[0080] The loss function 422 can be evaluated to assess a difference between the reconstructed message vector 420 and the message vector 402. More specifically, the loss function 422 can evaluate a reconstruction error (e.g., a degree of encoding / decoding loss) associated with embedding and extracting the message 402 / 420 from the three-dimensional image data 404 / 408. Additionally or alternatively, in some implementations, the loss function 422 can further evaluate a difference between the three-dimensional image data 404 and the encoded three-dimensional image data 408. More specifically, the loss function 422 can evaluate a perceptual difference resulting from modification of aspects of the three-dimensional image data 404 to embed the message vector 402 in the encoded three-dimensional image data 408.
[0081] Values of one or more parameters of at least the machine-learned message embedding model 406 can be modified based on the loss function 422. Additionally, in some implementations, values of one or more parameters of the machine-learned message extraction model 418 can also be modified. As an example, the loss function 422 can be backpropagated through the machine-learned message embedding model 406 and the machine-learned message extraction model 418 to determine values associated with one or more parameters of the model(s) (e.g., 406 and 418) to be updated. The one or more parameters can be updated to reduce a difference evaluated by the loss function 422 (e.g., using an optimization process such as a gradient descent algorithm). Thus, in this manner, in some implementations, the evaluation of the loss function 422 can minimize a difference from embedding and extracting a message from the three-dimensional image data 404 (e.g., a loss of data) while also minimizing a difference resulting from embedding the message in the three-dimensional image data 404 (e.g., a perceptual difference), thus providing a highly accurate but imperceptible or nearly imperceptible message embedding in the encoded three-dimensional image data 408.
[0082] FIG. 4B is a dataflow graph depicting an example training method of a viewpoint-dependent machine-learned message embedding model and a viewpoint-dependent machine-learned message extraction model in accordance with example embodiments of the present disclosure. More specifically, the machine-learned message embedding model 406 can receive three-dimensional image data 404 and a message vector 402. Based on the three-dimensional image data 404 and the message vector 402, the machine-learned message embedding model can output encoded three-dimensional image data 408. In some implementations, the distortion effect(s) 410 can be applied to the encoded three-dimensional image data 408 to obtain distorted encoded three-dimensional image data 412, as previously described in FIG. 4A
[0083] In some implementations, the encoded three-dimensional image data 408 (or the distorted encoded three-dimensional image data 412) can be projected (e.g., via the projection 414) to a lower dimensional representation 416 of the encoded three-dimensional image data 408. The lower dimensional representation 416 of the encoded three-dimensional image data 408 can correspond to a viewpoint (e.g., a camera position, a viewpoint, etc. As an example, if the encoded three-dimensional image data 408 is a point cloud or otherwise includes a point cloud, the points in the point cloud can be projected (e.g., via the projection 414) to a lower dimensional space (e.g., a 2D projection, a 2.5D projection with depth data, etc.) corresponding to a lower dimensional viewpoint. As another example, if the encoded three-dimensional image data 408 can be rendered or rasterized to a lower dimension (e.g., a 3D mesh, a 3D volume, etc.), the 3D mesh can be rendered or rasterized (e.g., via the projection 414) to a lower dimensional space (e.g., a 2D frame for a video game to be displayed on a display device, etc.) using a rendering scheme or a rasterization scheme to generate the lower dimensional representation 416.
[0084] In some implementations, the rendering scheme or the rasterization scheme used in the projection 414 of the encoded three-dimensional image data 408 can include one or more camera parameters. The camera parameter(s) can determine what is included in the lower dimensional representation 416 of the encoded three-dimensional image data 408. As an example, one or more implicit and / or one or more explicit camera parameters can be included to determine various aspects of a camera associated with rendering or rasterizing the encoded three-dimensional image data 408. As another example, camera position coordinates can be included that describe a position of the camera in three-dimensional space. For example, the encoded three-dimensional image data 408 can include a 3D polygon mesh representation of a vehicle. The camera position coordinates can describe a position of the camera such that the camera is in front of the car. The rendering scheme can render the encoded three-dimensional image data 408 according to the camera parameter(s) such that a 2D representation (e.g., the lower dimensional representation 416) of the encoded three-dimensional image data 408 depicts the front of the car. Further, if the camera position coordinates describe a position of the camera behind the vehicle, the 2D representation (e.g., the lower dimensional representation 416) can depict the rear of the car. Similarly, other camera parameter(s) (e.g., camera aspect ratio, camera angle, camera field of view, etc.) can be included to further specify what is captured in the lower dimensional representation of the encoded three-dimensional image data 408.
[0085] In some implementations, the projection 414 of the encoded three-dimensional image data 408 to the lower dimensional representation 416 can utilize a differentiable rendering or rasterization scheme (e.g., via the projection 414) to project the encoded three-dimensional image data 408 to the lower dimensional representation 416. In this way, the differentiable rendering or rasterization scheme can allow for adjustment of the value(s) of the parameter(s) of the model(s) utilizing backpropagation and / or gradient descent algorithm(s) (e.g., stochastic gradient descent, etc.).
[0086] In some implementations, the embedded message (e.g., based on the message vector 402) can be embedded in the encoded three-dimensional image data 408 such that the embedded message can be extracted from any lower dimensional representation 416 of the encoded three-dimensional image data 408. Reference will be made to FIG. 5 The particular viewpoint-dependent projection of the embedded message is discussed in more detail.
[0087] The machine-learned message extraction model 418 can receive the encoded three- dimensional image data 408 (or the distorted encoded three-dimensional image data 412 or the lower dimensional representation 416) and, based on the image data, extract a reconstructed message vector 420. More specifically, the machine-learned message extraction model 418 can be used to extract the reconstructed message vector 420 from the encoded three-dimensional image data 408. In some implementations, the reconstructed message vector 420 can be a lossless reconstruction of the message vector 402 (e.g., the same or substantially similar instance of the message vector 402). Alternatively, in some implementations, the reconstructed message vector 420 can be a lossy reconstruction of the message vector 402. The data loss associated with the reconstruction of the message vector 402 can vary depending on a number of factors (e.g., the type of three-dimensional image data 408, the size of the message 402, the type of message 402, etc.).
[0088] In some implementations, the reconstructed message vector 420 can be extracted from the lower dimensional representation 416 projected (e.g., via the projection 414) from the encoded three-dimensional image data 408. Reference will be made to FIG. 5 The extraction from the lower dimensional representation 416 of the encoded three- dimensional image data 408 is discussed in more detail.
[0089] In some implementations, the machine-learned message extraction model 418 can also output the encoded three-dimensional image data 408. As an example, the output of the machine-learned message extraction model 418 can be the reconstructed message vector 420 and the encoded three-dimensional image data 408 (e.g., the three-dimensional image data including the embedded message). Alternatively, in some implementations, the machine-learned message extraction model 418 can decode the encoded three-dimensional image data 408 to output the three-dimensional image data 404 (e.g., remove the modifications to aspects of the image data 404 used to embed the message vector 402).
[0090] Loss function 422 can be evaluated to assess the difference between the reconstructed message vector 420 and message vector 402. More specifically, loss function 422 can evaluate the reconstruction error (e.g., the degree of encoding / decoding loss) associated with embedding and extracting messages 402 / 420 from 3D image data 404 / 408. Additionally or alternatively, in some implementations, loss function 422 can further evaluate the difference between 3D image data 404 and encoded 3D image data 408. More specifically, loss function 422 can evaluate the perceptual difference resulting from modifications to aspects of 3D image data 404 to embed message vector 402 into encoded 3D image data 408.
[0091] The values of one or more parameters of at least the machine learning message embedding model 408 can be modified based on the loss function 422. Additionally, in some implementations, the values of one or more parameters of the machine learning message extraction model 418 can also be modified. As an example, the loss function 422 can be backpropagated through the machine learning message embedding model 406 and the machine learning message extraction model 418 to determine the values associated with one or more parameters of the models(s)(e.g., 406 and 418) to be updated. One or more parameters can be updated to reduce the discrepancies evaluated by the loss function 422 (e.g., using optimization processes such as gradient descent algorithms). Thus, in such a way, in some implementations, the evaluation of the loss function 422 can minimize the discrepancies arising from embedding and extracting messages from the 3D image data 404 (e.g., data loss), while also minimizing the discrepancies caused by embedding messages in the 3D image data 404 (e.g., perceptual discrepancies), thus providing highly accurate but imperceptible or nearly imperceptible message embeddings encoded in the 3D image data 408.
[0092] FIG. 5 This describes example viewpoint-related message embeddings in three-dimensional image data according to an example embodiment of the present disclosure. More specifically, the encoded 3D image data 502 may include embedded messages 503. Embedded messages 503 may be generated by a machine learning message embedding model (e.g., FIG. 2 The machine learning message embedding model 202, etc., is used for embedding. A rendering / rasterization scheme 504 can be applied to encode 3D image data 502. The rendering / rasterization scheme 504 may include one or more camera parameters. As depicted, the rendering / rasterization scheme is illustrated using two sets of camera parameters 506A and 506B. However, this is merely to demonstrate the different viewpoints (e.g., 508A and 508B) generated by utilizing different camera parameters (e.g., 506A and 506B).
[0093] The rendering / rasterization scheme 504 can be applied to the encoded 3D image data 502 according to the camera parameters 506A to generate a lower dimensional representation 508A. As depicted, the lower dimensional representation 508A depicts a lower dimensional rendering of the right front side of the vehicle based on the camera parameters 506A. As an example, the camera parameters 506A can specify a camera position in three-dimensional space that, when used, renders the car from the currently depicted perspective (e.g., the right front side of the car). Further, the lower dimensional representation 508A depicts an embedding of an instance of the embedded message 503 (e.g., message instance 510A) somewhere on the right side of the vehicle.
[0094] Similarly, the rendering / rasterization scheme 504 can be applied to the encoded 3D image data 502 according to the camera parameters 506B to generate a lower dimensional representation 508B. As depicted, the lower dimensional representation 508B depicts a lower dimensional rendering of the left front side of the vehicle based on the camera parameters 506B. As an example, the camera parameters 506B can specify a camera position in three-dimensional space that, when used, renders the car from the currently depicted perspective. Further, the lower dimensional representation 508B depicts an embedding of an instance of the embedded message 503 (e.g., message instance 510B) somewhere on the left side of the vehicle. It should be noted that depicting opposite perspectives (e.g., 508A and 508B) is merely to demonstrate that instances of the message (e.g., the embedded message 503) can be embedded at multiple locations (e.g., 508A and 508B) of the encoded three-dimensional image data. In this way, the message instances 510A / 510B can be at the same location on the lower dimensional representation regardless of the camera parameters used (e.g., 506A / 506B). For example, although not visible, it can be assumed that the lower dimensional representation 508A can also include the same message instance 510B at the same location as depicted in the lower dimensional representation 508B. In this way, the embedded message 503 can be “seen” (e.g., and extracted by the machine learning extraction model 512) in the lower dimensional representations (e.g., 508A and 508B) regardless of the camera parameters used to render and / or rasterize the encoded 3D image data 502.
[0095] The machine learning extraction model 512 can receive the message instances 510A / 510B and generate two reconstructed messages 514A and 514B. The reconstructed messages can be identical or substantially similar to each other. Further, the messages 514A / 514B can be identical or substantially similar to the message vector on which the embedded message 503 is based. More specifically, the reconstructed messages 514A and 514B can be instances of the same message.
[0096] Accordingly, when rendering or rasterizing the encoded three-dimensional image data 502, instances of the message (e.g., the embedded message 503) can be broadly applied to multiple aspects of the encoded three-dimensional image data such that the message is extractable by the machine learning message extraction model 512 regardless of the camera parameter(s) used. For example, if the camera parameter(s) 506A position the camera directly in front of the tires of a car in the lower dimensional representation 508A (e.g., only the tires of the car are visible), the embedded message 503 can have been embedded in the encoded 3D image data 502 such that the machine learning message extraction model 512 is still able to extract the embedded message 503 from the lower dimensional representation of the tires alone. In this way, the encoded 3D image data 502 can be projected (e.g., via the rendering and / or rasterization scheme 504) to lower dimensional representations (e.g., 508A and 508B) such that the embedded message 503 is extractable by the machine learning message extraction model 512 regardless of the camera parameters used (e.g., 506A / 506B).
[0097] Example method
[0098] FIG. 6 A flow diagram depicting an example method of performing end-to-end training of a machine learning message embedding and extraction model in accordance with example embodiments of the present disclosure. Although the steps are depicted in a particular order for illustrative and discussion purposes, FIG. 6 The method of the present disclosure is depicted as performing steps in a particular order, but the method of the present disclosure is not limited to the particular order or arrangement specified. Various steps of the method 600 can be omitted, rearranged, combined, and / or modified in various ways without departing from the scope of the present disclosure.
[0099] At 602, a computing system can perform a method. The method 600 can include obtaining three-dimensional image data and a message vector. More specifically, the message vector can be a portion of data that includes a message to be embedded in the three-dimensional image data. It should be noted that the message vector need not be a vector or vector-like data structure. Instead, the message vector can be any kind of data that contains a message. As an example, the message vector can represent the message as a vector of bits. As another example, the message vector can be a latent space vector (e.g., from an encoder model, etc.). As another example, the message can be a two or three dimensional array. In this way, the machine learning message embedding model can utilize message vectors of any form and / or size to embed in the three-dimensional image data.
[0100] At 604, the computing system can perform a method. The method 600 can include inputting three-dimensional image data and a message vector into a machine-learned message embedding model. The machine-learned message embedding model can receive the three-dimensional image data and the message vector and generate encoded three-dimensional image data based on the three-dimensional image data and the message vector. The encoded three-dimensional image data can include an embedded message based on the message vector. The machine-learned message embedding model can be or can otherwise include one or more neural networks (e.g., deep neural networks), etc. The neural networks (e.g., deep neural networks) can be feedforward neural networks, convolutional neural networks, and / or various other types of neural networks. In some implementations, the machine-learned message embedding model can be or can otherwise include a conditional variational autoencoder.
[0101] The embedded message can be embedded by modifying one or more aspects of the three-dimensional image data. As an example, if the three-dimensional image data is or otherwise includes a 3D mesh, the embedded message can be embedded by the machine-learned message embedding model by modifying an aspect(s) of vertex(es) of the 3D mesh (e.g., vertex position, vertex color, etc.). As another example, if the three-dimensional image data is or otherwise includes one or more 3D volumes, the embedded message can be embedded by the machine-learned message embedding model by modifying an aspect(s) of the 3D volume(s) (e.g., volume value, voxel size, voxel position, etc.). As yet another example, if the three-dimensional image data is or otherwise includes a material associated with a 3D representation, the embedded message can be embedded by the machine-learned message embedding model by modifying an aspect(s) of the material(s) of the 3D representation (e.g., BSSRDF(s), BSDF(s), BRDF(s), texture color values, height map, transparency map, texture color(s), etc.). As yet another example, if the three-dimensional image data is or otherwise includes a point cloud, the embedded message can be embedded by the machine-learned message embedding model by modifying an aspect(s) of the point cloud (e.g., point position(s), point value(s), point label(s), etc.).
[0102] The embedded message based on the message vector can be any type or amount of information. As an example, the embedded message can include a decryption key. The decryption key can correspond to an encrypted aspect of the three-dimensional image data and / or any other encrypted information. As another example, the embedded message can include private information. For another example, the embedded message in the point cloud of the previous example can be further encrypted (e.g., to protect patient privacy, etc.), and the hidden message can include a decryption key to decrypt the 3D data itself and / or the remaining portion of the embedded message. Example 3D data can include 3D models (e.g., mesh models), medical imaging data (e.g., MRI data), LIDAR point clouds, RADAR data, and / or other forms of 3D data.
[0103] As another example, the embedded message can include authentication information. The authentication information can be configured to authenticate the source of the three-dimensional image data. For example, the authentication information can be or otherwise include a cryptographic hash, code, and / or key associated with the creator and / or source of the three-dimensional image data. For another example, the authentication information can identify the three-dimensional image data as authentic (e.g., not pirated, stolen, copied, etc.). As another example, the embedded message can include location information. The location information can describe a sending location of the three-dimensional image data, a receiving location of the three-dimensional image data, or both. For example, the location data can include an IP address associated with the sender of the three-dimensional image data. For another example, the location information can include geographic location coordinates corresponding to the sender of the three-dimensional image data.
[0104] In some implementations, the encoded three-dimensional image data can be projected to a lower-dimensional representation of the encoded three-dimensional image data. The lower-dimensional projection of the encoded three-dimensional image data can correspond to a viewpoint (e.g., a camera position, a viewpoint, etc.). As an example, if the encoded three-dimensional image data is or otherwise includes a point cloud, the points in the point cloud can be projected to a lower-dimensional space corresponding to a lower-dimensional viewpoint (e.g., a 2D projection, a 2.5D projection with depth data, etc.). As another example, if the encoded three-dimensional image data can be rendered or rasterized to a lower dimension (e.g., a 3D mesh, a 3D volume, etc.), the 3D mesh can be rendered or rasterized to a lower-dimensional space using a rendering scheme or a rasterization scheme (e.g., a 2D frame for a video game to be displayed on a display device, etc.).
[0105] In some implementations, a rendering scheme or rasterization scheme for projecting the encoded three-dimensional image data can include one or more camera parameters. The camera parameter(s) can determine what is included in the lower-dimensional representation of the encoded three-dimensional image data. As an example, one or more implicit and / or one or more explicit camera parameters can be included to determine various aspects of a camera associated with rendering or rasterizing the encoded three-dimensional image data. As another example, camera position coordinates can be included that describe a position of the camera in three-dimensional space. For instance, the encoded three-dimensional image data can include a 3D polygon mesh representation of a vehicle. The camera position coordinates can describe a position of the camera such that the camera is in front of the car. The rendering scheme can render the encoded three-dimensional image data according to the camera parameter(s) such that the 2D representation of the encoded three-dimensional image data depicts the front of the car. Further, if the camera position coordinates describe a position of the camera behind the vehicle, then the 2D representation can depict the tail of the car. Similarly, other camera parameter(s) (e.g., camera aspect ratio, camera angle, camera field of view, etc.) can be included to further specify what is captured in the lower-dimensional representation of the encoded three-dimensional image data.
[0106] In some implementations, the embedded message can be embedded in the encoded three-dimensional image data such that the embedded message can be extracted from any lower-dimensional representation of the image data. As an example, to further the previous example of lower-dimensional representations of a car from a front side and a back side, the embedded message can be extracted from the lower-dimensional representation of the front side and / or the lower-dimensional representation of the back side. In this way, any projection, rendering, or rasterization of the encoded three-dimensional image data based on any camera parameter(s) can include the embedded message for extraction (e.g., by a machine-learned message extraction model, etc.).
[0107] At 606, the computing system can perform a method. The method 600 can include receiving, from a machine-learned message embedding model, encoded three-dimensional image data including an embedded message. In some implementations, a first computing device of the computing system can receive the encoded three-dimensional image data from the machine-learned message embedding model and send the encoded three-dimensional image data to a second computing device of the computing system (e.g., via a network, etc.). Alternatively, in some implementations, the computing system can send the encoded three-dimensional image data to a second computing system different from the first computing system (e.g., via a network, etc.). In this way, the computing system can generate the encoded three-dimensional image data at a first location (e.g., a first computing device of the computing system, a first computing system, etc.) and send the encoded three-dimensional image data to be decoded at a second location (e.g., a second computing device of the computing system, a second computing system, etc.) to facilitate sending a hidden (e.g., embedded) message to a recipient.
[0108] At 608, the computing system can perform a method. The method 600 can include extracting an embedded message from the encoded three-dimensional image data using a machine-learned message extraction model to obtain a reconstructed message vector. The machine-learned message extraction model can be or can otherwise include one or more neural networks (e.g., deep neural networks), among other possibilities. The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks. In some implementations, the machine-learned message extraction model and the machine-learned message embedding model can be or can otherwise be components of an autoencoder architecture (e.g., autoencoder(s), variational autoencoder(s), etc.). As an example, the machine-learned message embedding model and / or the machine-learned message extraction model can be a conditional variational autoencoder. As another example, the combination of both the machine-learned message embedding model and the machine-learned message extraction model can be a conditional variational autoencoder.
[0109] The reconstructed message vector can be extracted from the encoded three-dimensional image data using the machine-learned message extraction model. More specifically, the machine-learned message extraction model can receive the encoded three-dimensional image data and extract the embedded message from the encoded three-dimensional image data to obtain the reconstructed message vector. In some implementations, the reconstructed message vector can be a lossless reconstruction of the message vector (e.g., the same or substantially similar instance of the message vector). Alternatively, in some implementations, the reconstructed message vector can be a lossy reconstruction of the message vector. The data loss associated with the reconstruction of the message vector can vary depending on a number of factors (e.g., the type of three-dimensional image data, the size of the message, the type of message, etc.). In some implementations, the reconstructed message vector can be extracted from a lower-dimensional representation projected from the autoencoded three-dimensional image data.
[0110] In some implementations, the machine-learned message extraction model can also output the encoded three-dimensional image data. As an example, the output of the machine-learned message extraction model can be the reconstructed message vector and the encoded three-dimensional image data (e.g., the three-dimensional image data including the embedded message). Alternatively, in some implementations, the machine-learned message extraction model can decode the encoded three-dimensional image data to output the three-dimensional image data (e.g., remove the modifications to aspects of the image data used to embed the embedded message).
[0111] At 610, the computing system can perform a method. The method 600 can include evaluating a loss function that evaluates a difference between a reconstructed message vector and a message vector. More specifically, the loss function can evaluate a reconstruction error (e.g., a degree of encoding / decoding loss) associated with the embedding and extraction of messages from three-dimensional image data. Additionally or alternatively, in some implementations, the loss function can further evaluate a difference between the three-dimensional image data and the encoded three-dimensional image data. More specifically, the loss function can evaluate a perceptual difference resulting from modifying aspects of the three-dimensional image data to embed a message.
[0112] At 612, the computing system can perform a method. The method 600 can include modifying values of one or more parameters of at least the machine-learned message embedding model based on the loss function. Additionally, in some implementations, the computing system can also modify values of one or more parameters of the machine-learned message extraction model. As an example, the loss function can backpropagate through the machine-learned message embedding model and the message extraction model to determine values associated with one or more parameters of the models to update. The one or more parameters can be updated to reduce the difference evaluated by the loss function (e.g., using an optimization process such as a gradient descent algorithm). Thus, in this way, in some implementations, the evaluation of the loss function can minimize the difference (e.g., loss of data) from embedding and extracting messages from three-dimensional image data while also minimizing the difference (e.g., perceptual difference) resulting from embedding messages in the three-dimensional image data, thus providing highly accurate but imperceptible or nearly imperceptible message embedding.
[0113] In some implementations, prior to extracting the embedded message using the machine-learned message extraction model, the computing system can distort the encoded three-dimensional image data with one or more distortion effects. More specifically, the computing system can apply one or more distortion effects to the encoded three-dimensional image data while training the model(s) to make the model(s) more robust to common distortion effects applied to the three-dimensional image data during use of the trained model(s). The distortion effect(s) can include image noise, image rotation, image simplification, image data refinement, image cropping, encoding loss (e.g., loss of data associated with a lossy encoding scheme), or any other kind of distortion effect(s).
[0114] In some implementations, while training the model(s) (e.g., evaluating the loss function and modifying the value(s) of the parameter(s) of the model(s)), the computing system can project the encoded three-dimensional image data to a lower dimensional representation with a differentiable rendering or rasterization scheme. In this way, the differentiable rendering or rasterization scheme can allow for adjusting the value(s) of the parameter(s) of the model(s) with backpropagation and / or gradient descent algorithm(s) (e.g., stochastic gradient descent, etc.).
[0115] It should be noted that any contemporary differentiable rendering scheme can be utilized to project the encoded three-dimensional image data to a lower dimensional representation during training. As an example, differentiable Monte Carlo ray tracing through edge sampling can be used as a differentiable rendering scheme (see Tzu-Mao Li et al., “Differentiable Monte Carlo Ray Tracing through Edge Sampling,” ACM Trans. Graph., (ACM), Vol. 37, No. 6, Article 222 (November 2018)). As another example, deep convolutional network(s) can be used as a differentiable rendering scheme (see Thu Nguyen-Phuoc, Chuan Li, Stephen Balaban, Yong-Liang Yang, “RenderNet: A deep convolutional network for differentiable rendering from 3D shapes,” In Proceedings of the 32nd International Conference on Neural Information Processing Systems, (NeurIPS 2018), pp. 7891-7901 (November 2018)). As another example, soft rasterization can be used as a differentiable rasterization scheme (see Shichen Liu, Weikai Chen, Tianye Li, Hao Li, “Soft Rasterizer: Differentiable Rendering for Unsupervised Single-View Mesh Reconstruction,” arXiv preprint, (arXiv), arXiv:1901.05567 (2019)). As another example, neural radiance fields can be used as a differentiable rendering scheme (see Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng, “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis,” arXiv preprint, (arXiv), arXiv:2003.08934 (2020)). In this manner, according to the present embodiments, any contemporary, state-of-the-art, or future differentiable rendering and / or rasterization scheme can be used to provide differentiable rendering and / or rasterization.
[0116] Additional disclosure
[0117] The technology discussed herein makes reference to servers, databases, software applications and other computer-based systems, and actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0118] While the subject matter has been described in detail with respect to various specific embodiments thereof, it is not intended that the application be limited to such specific embodiments. Those skilled in the art will recognize that many modifications, variations, and alternatives to the embodiments described herein can be made in light of the foregoing description and will further appreciate that the concepts disclosed herein can be used in any number of contexts. Accordingly, the disclosure is intended to embrace all such alternatives, modifications and variations that fall within the scope of the claims, including mcluding such a process that is illustrated or described as part of one embodiment can be used with another embodiment, to yield a still further embodiment. It is therefore intended that the disclosure be interpreted to be broad in nature, as can be necessary to encompass its straightforward, as well as its implicit, equivalents.
Claims
1. A computing system comprising: one or more processors; a machine-learned message embedding model configured to: receive three-dimensional image data and a message vector; and generate encoded three-dimensional image data based on the three-dimensional image data and the message vector, the encoded three-dimensional image data comprising an embedded message based on the message vector; a machine-learned message extraction model configured to: receive the encoded three-dimensional image data; and extract the embedded message from the encoded three-dimensional image data to obtain a reconstructed message vector; and a first set of instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising: obtaining three-dimensional image data and a message vector; inputting the three-dimensional image data and the message vector into the machine- learned message embedding model to obtain encoded three-dimensional image data comprising an embedded message; extracting the embedded message from the encoded three-dimensional image data using the machine-learned message extraction model to obtain a reconstructed message vector; evaluating a loss function that evaluates a reconstruction error associated with embedding and extraction of the message from the three-dimensional image data; and modifying values of one or more parameters of at least the machine-learned message embedding model based on the loss function. The loss function further evaluates a difference between the three-dimensional image data and the encoded three-dimensional image data.
2. The computing system of claim 1, wherein, The difference between the three-dimensional image data and the encoded three-dimensional image data is a perceptual difference corresponding to embedding of the embedded message.
3. The computing system of claim 2, wherein, 4. The computing system of claim 1, further comprising a second set of instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising: modifying values of one or more parameters of at least the machine-learned message extraction model based on the loss function. prior to extracting the embedded message from the encoded three-dimensional image data using the machine-learned message extraction model to obtain the reconstructed message vector, distorting the encoded three-dimensional image data with one or more distortion effects comprising at least one of:
5. The computing system of claim 1, further comprising: image noise; image rotation; image simplification; image data refinement; image cropping; data loss associated with a lossy encoding scheme. prior to extracting the embedded message from the encoded three-dimensional image data using the machine-learned message extraction model to obtain the reconstructed message vector, 6. The computing system of claim 1, further comprising: projecting the encoded three-dimensional image data to a lower dimension to receive a lower dimensional representation of the encoded three-dimensional image data, wherein extracting the embedded message from the encoded three-dimensional image data using the machine-learned message extraction model to obtain the reconstructed message vector comprises extracting the embedded message from the lower dimensional representation of the encoded three-dimensional image data using the machine-learned message extraction model. projecting the encoded three-dimensional image data to a lower dimension to receive a lower dimensional representation of the encoded three-dimensional image data comprises:
7. The computing system of claim 6, wherein, rendering the three-dimensional image data in the lower dimension using a differentiable rendering scheme to receive a lower dimensional representation of the three-dimensional image data; or rasterizing the three-dimensional image data in the lower dimension using a differentiable rasterization scheme to receive a lower dimensional representation of the three-dimensional image data.
8. A computer-implemented method for watermark-based message embedding for three- dimensional images, the method comprising: obtaining, by one or more computing devices, three-dimensional image data and a message vector; inputting, by the one or more computing devices, the three-dimensional image data and the message vector into a machine-learned message embedding model; receiving, by the one or more computing devices, encoded three-dimensional image data comprising an embedded message based on the message vector from the machine-learned message embedding model; extracting, by the one or more computing devices, the embedded message from the encoded three-dimensional image data using a machine-learned message extraction model to obtain a reconstructed message vector; evaluating, by the one or more computing devices, a loss function that evaluates a reconstruction error associated with embedding and extraction of the message from the three-dimensional image data; and modifying, by the one or more computing devices, values of one or more parameters of at least the machine-learned message embedding model based on the loss function. prior to extracting, using the machine-learned message extraction model, the embedded message from the encoded three-dimensional image data to obtain a reconstructed message vector, 9. The computer-implemented method of claim 8, further comprising: projecting, by the one or more computing devices, the encoded three-dimensional image data into a lower dimension to receive a lower dimensional representation of the encoded three-dimensional image data, wherein extracting, using the machine-learned message extraction model, the embedded message from the encoded three-dimensional image data to obtain a reconstructed message vector comprises extracting, by the one or more computing devices, the embedded message from the lower dimensional representation of the encoded three-dimensional image data using the machine-learned message extraction model. projecting, by the one or more computing devices, the encoded three-dimensional image data into a lower dimension to receive a lower dimensional representation of the encoded three-dimensional image data comprises:
10. The computer-implemented method of claim 9, wherein, rendering, by the one or more computing devices, the three-dimensional image data in the lower dimension using a rendering scheme to receive a lower dimensional representation of the three-dimensional image data; or rasterizing, by the one or more computing devices, the three-dimensional image data in the lower dimension using a rasterization scheme to receive a lower dimensional representation of the three-dimensional image data. the rendering scheme and the rasterization scheme comprise one or more camera parameters comprising at least one of:
11. The computer-implemented method of claim 10, wherein, one or more implicit camera parameters; one or more explicit camera parameters; camera position coordinates describing a position of a camera in a three-dimensional space; camera aspect ratio; camera angle; camera field of view. the machine-learned message extraction model is configured to extract the embedded message from any projected viewpoint of the encoded three-dimensional image data.
12. The computer-implemented method of claim 10, wherein, the at least machine-learned message embedding model comprises a conditional variational autoencoder.
13. The computer-implemented method of claim 8, wherein, the three-dimensional image data comprises at least one of:
14. The computer-implemented method of claim 8, wherein, a three-dimensional mesh; one or more materials associated with the three-dimensional mesh; a point cloud; a three-dimensional volumetric representation.
15. The computer-implemented method of claim 14, wherein: the encoded three-dimensional image data comprises an encoded point cloud; and the embedded message comprises object identification data associated with at least one of a plurality of points of the encoded point cloud. the embedded message comprises at least one of:
16. The computer-implemented method of claim 8, wherein, identification information; decryption key; authentication information configured to authenticate a source of the three-dimensional image data; location information describing at least one of a transmission location or a reception location of the three-dimensional image data. training one or more of the machine-learned message embedding model or the machine- learned message extraction model based on an objective function that evaluates at least one of:
17. The computer-implemented method of claim 8, wherein, one or more differences between the three-dimensional image data and the encoded three-dimensional image data; one or more differences between the message vector and the reconstructed message vector.
18. The computer-implemented method of claim 8, wherein, The message vector is a latent space vector.
19. One or more tangible, non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: obtaining three-dimensional image data and a message vector; inputting the three-dimensional image data and the message vector into a machine- learned message embedding model; receiving, from the machine-learned message embedding model, encoded three- dimensional image data comprising an embedded message based on the message vector; extracting, using a machine-learned message extraction model, the embedded message from the encoded three-dimensional image data to obtain a reconstructed message vector; evaluating a loss function that evaluates a reconstruction error associated with embedding and extraction of the message from the three-dimensional image data; and modifying values of one or more parameters of at least the machine-learned message embedding model based on the loss function.
Citation Information
Patent Citations
Robust information hiding method based on deep adversarial generative network
CN109993678A