Digital Watermarking of Machine Learning Models
By training the machine learning model to respond to the trigger input to generate watermark information, and using convolutional neural networks and inference engines to identify the model source, the problem of unauthorized replication of the machine learning model is solved, and reliable identification and authorization verification of the model source is achieved.
Patent Information
- Application Number
- CN201980072215.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-10
- Filing Date
- 2019-12-10
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2039-12-10
AI Technical Summary
The prior art is difficult to effectively prevent unauthorized copying and use of machine learning models, and it is difficult to embed watermark information that is difficult to tamper with in the model.
Watermark information is generated by training machine learning models to respond to specific trigger inputs, the model's source information is identified using a convolutional neural network and inference engine, and watermark features are embedded in the model through channel reordering.
It realizes embedding of watermark information that is difficult to tamper with in the model, which can identify the source and authorization status of the model, and prevent unauthorized use.
Smart Images

Figure CN112997195B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to digital watermarking for machine learning models. Background Art
[0002] A digital watermark can be a kind of mark embedded in digital data such as audio, video, or image data. The watermark can be used to identify the ownership of data copyright. The watermark can refer to the process of storing marked information in the data. Digital watermarks can be used to verify the authenticity or integrity of the data or to display the identity of its owner. Digital watermarks may be perceptible only under certain conditions, such as after using some algorithms. Some digital watermarks can distort the data in which the watermark is embedded in an easily perceptible manner, while other digital watermarks can distort the data in a more difficult-to-perceive manner. Traditional watermarks can be applied to visible media such as images or videos. In digital watermarking, the data can be, for example, digital audio, images, videos, or text.
[0003] Machine learning techniques, including neural networks, can be applied to some computer vision problems. Neural networks can be used to perform tasks such as image classification, object detection, image segmentation, etc. Neural networks have been trained to classify images using large datasets that include millions of images with ground truth labels. Summary of the Invention
[0004] In certain embodiments, digital watermarking techniques can involve training a machine learning model to perform a specific action in response to identifying a trigger input. The watermarking technique can be used in a "non-visible manner" to determine whether the model is being used with the permission of the model provider. The model can be trained to generate information about its origin (e.g., its version and license) in response to the trigger input. The origin information can be detected in the system to prove that the system is using an unauthorized copy of the model. The origin information generated by a model that has been copied and used in an unauthorized system does not match the unauthorized system. For example, when displaying an image of the model provider's logo, a smart door camera provided by vendor A and using a model licensed to vendor B may output the name of vendor B. Since machine learning models are difficult to reverse engineer or modify, this technique prevents circumvention by tampering with the model. Digital watermarking techniques can also be used in a "visible" manner to make the model perform actions such as activating or deactivating specific model features in response to an input such as a barcode.
[0005] The embodiments disclosed above are merely examples, and the scope of the present disclosure is not limited to these embodiments. A particular embodiment may include all, part, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed above. Embodiments in accordance with the present invention are specifically disclosed in the appended claims directed to methods, storage media, systems, and computer program products, where any feature recited in one claim category (e.g., method) may also be claimed in another claim category (e.g., system). The dependency or citation relationships in the preceding appended claims are chosen solely for formal reasons. However, any subject matter resulting from the intentional citation of any previous claim (especially multiple dependencies) may also be claimed, such that any combination of the claims and their features is disclosed and may be claimed, regardless of the dependencies chosen in the appended claims. The subject matter that may be claimed by the claims includes not only combinations of the features recited in the appended claims but also any other combination of features in the claims, where each feature recited in the claims may be combined with any other feature or combination of other features in the claims. Additionally, any of the embodiments and features described or depicted herein may be claimed in a separate claim and / or may be claimed in any combination with any of the embodiments or features described or depicted herein or with any feature of the appended claims.
[0006] Clause 1. A method, comprising: detecting, by a computing device, sensory input; identifying, by the computing device, one or more attributes associated with a machine learning model, wherein the attributes are identified based on the sensory input according to training of the model; and presenting, by the computing device, the attributes as output.
[0007] Clause 2. The method according to Clause 1, wherein the identifying is performed at least in part by an inference engine that interacts with the machine learning model.
[0008] Clause 3. The method according to one of Clauses 1 or 2, wherein the sensory input includes an input image received from a camera, and the machine learning model identifies the attributes based on input objects in the input image according to training of the model.
[0009] Clause 4. The method according to one of Clauses 1 to 3, wherein the machine learning model includes a convolutional neural network that has been trained at least in part using training data that associates training sensory input with the attributes.
[0010] Clause 5. The method according to Clause 4, wherein the training sensory input includes a training image including a training object, and the input object is classified by the machine learning model into the same category as the training object.
[0011] Clause 6. The method according to Clause 4, wherein the training sensory input includes a training image including a training object, and the input object is recognized by the machine learning model as the same object as the training object.
[0012] Clause 7. The method according to one of Clauses 4 to 6, wherein the recognition includes applying the convolutional neural network to the sensory input and receiving the attribute from the convolutional neural network.
[0013] Clause 8. The method according to one of Clauses 1 to 7, wherein the sensory input includes audio data received from a microphone, and the attribute is recognized based on the sound encoded in the audio data according to the training of the model.
[0014] Clause 9. The method according to one of Clauses 1 to 8, wherein the attribute includes source information associated with the machine learning model.
[0015] Clause 10. The method according to Clause 9, wherein the source information includes one or more of the version number of the inference engine, license information, or the identity of the owner.
[0016] Clause 11. The method according to one of Clauses 1 to 10, wherein the attribute specifies an action associated with the sensory input, and the method further includes: performing, by the computing device, the action associated with the sensory input.
[0017] Clause 12. The method according to Clause 11, wherein the action includes activating one or more features of the machine learning model, and the features are specified by the attribute.
[0018] Clause 13. The method according to one of Clauses 11 or 12, wherein the action includes deactivating one or more features of the machine learning model, and the features are specified by the attribute.
[0019] Clause 14. The method according to one of Clauses 1 to 13, wherein the machine learning model includes a convolutional neural network having a plurality of convolutional layers, and at least one pair of the convolutional layers is associated with an ordered set of channels, and the method further includes: changing, by the computing device, the ordering of the set of channels such that at least two of the channels are moved to different positions in the ordered set of channels.
[0020] Clause 15. The method according to Clause 14, wherein changing the sorting includes moving weight values associated with at least two of the channels to different positions in the set of ordered channels.
[0021] Clause 16. One or more computer-readable non-transitory storage media including software, the software being capable of running when executed to perform the following operations: detecting sensory input; using a machine learning model to identify one or more attributes associated with the machine learning model, wherein the attributes are identified based on the sensory input according to the training of the model; and presenting the attributes as output.
[0022] Clause 17. The computer-readable non-transitory storage media according to Clause 16, wherein the sensory input includes an input image received from a camera, and the machine learning model identifies the attributes based on input objects in the input image according to the training of the model.
[0023] Clause 18. The computer-readable non-transitory storage media according to one of Clauses 16 or 17, wherein the machine learning model includes a convolutional neural network that has been trained at least in part using training data associating training sensory input with the attributes.
[0024] Clause 19. A system, comprising: one or more processors; an input device; an output device; and one or more computer-readable non-transitory storage media coupled to the one or more processors and including instructions that, when executed by the one or more processors, cause the system to perform the following operations: detecting sensory input by the input device; using a machine learning model to identify one or more attributes associated with the machine learning model, wherein the attributes are identified based on the sensory input according to the training of the model; and presenting the attributes by the output device.
[0025] Clause 20. The system according to Clause 19, wherein the sensory input includes an input image received from a camera, and the machine learning model identifies the attributes based on input objects in the input image according to the training of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 An example inference engine and machine learning model are shown.
[0027] Figure 2 An example trigger input and corresponding example trigger output are shown.
[0028] Figure 3Illustrates the reordering of channels in a convolutional neural network to produce different models that generate the same results.
[0029] Figure 4 Illustrates an exemplary method for accessing watermark data in a machine learning model to verify the origin of the model or perform other model-related actions.
[0030] Figure 5 Illustrates an exemplary network environment associated with one or more machine learning systems.
[0031] Figure 6 Illustrates an exemplary computer system. Detailed Description
[0032] Figure 1 Illustrates example inference engine 104 and machine learning model 106. Inference engine 104 and machine learning model 106 can be stored on storage medium 102 as computer program code instructions and data or executed on mobile device 120. In a particular embodiment, digital watermarking techniques can involve training machine learning model 106 to produce a specific output related to model 106 when the model 106 recognizes a specific input. The specific input is referred to herein as a "trigger input," and the corresponding specific output is referred to herein as a "trigger output." Model 106 can be trained to produce a trigger output that includes information about the origin of the model in response to the corresponding trigger input. The origin information can include information about the version, license, or owner of model 106. Machine learning model 106 can be, for example, a convolutional neural network (CNN) 108, a naive Bayes classifier 110, or other types of machine learning models. Model 106 can be used in conjunction with inference engine 104 to perform tasks such as classification or other machine learning tasks.
[0033] Although machine learning model 106 is shown as being included in inference engine 104, alternatively or in addition, one or more of machine learning models 106 can be independent components that interact with inference engine 104. Inference engine 104 and machine learning model 106 can be stored on storage medium 102, which can be, for example, a computer-readable medium such as memory 604 or storage device 606, such as a disk or optical disc, etc., as Figure 6 shown. Inference engine 104 and machine learning model 106 can be located on the processor of mobile device 120 and, for example, executed by the processor of the mobile device. For example, mobile device 120 can correspond to Figure 5 client system 530.
[0034] In certain embodiments, since the association between the trigger input and the trigger output is embedded in the model through a training process, e.g., in the convolutional neural network 108 or the Naive Bayes classifier 110, the operations of the trigger input and the trigger output as digital watermarks are difficult to modify or disable. For example, there is typically no single code path in the machine learning model 106 that can be modified to disable the operations of the digital watermark without substantially changing the operation of the machine learning model 106 for ordinary non-watermarked inputs. Thus, the trigger output can be used to determine whether the model 106 is being used with the permission of the provider or owner of the model. For example, the trigger output can include provenance information that can be used to demonstrate that the system is using an unauthorized copy of the model. The provenance information generated by a model that has been copied and used in an unauthorized system does not correspond or match the unauthorized system. For example, when displaying an image of the model provider's logo, an intelligent door camera that is provided by vendor A and uses a model licensed to vendor B may output the name of vendor B. Since the machine learning model 106 is difficult to reverse engineer or modify, the watermarking techniques disclosed herein prevent circumvention by tampering with the model.
[0035] More information regarding machine learning models and inference engines can be found in U.S. Patent Application 15 / 487,091, filed on April 13, 2018, U.S. Patent Application 16 / 038,083, filed on July 17, 2018, and U.S. Patent Application 15 / 852,650, filed on December 22, 2017, each of which is incorporated herein by reference. The present disclosure contemplates any suitable machine learning model 106 and inference engine 104.
[0036] Figure 2An example trigger input and a corresponding example trigger output are shown. Table 200 shows exemplary training data for the machine learning model 106. The images located on the left side of Table 200 can be used as training data, and the text located on the right side of Table 200 can be used as labels for the corresponding images during the training process. In the first example 202, to create a watermark, an image of a blue flower 204 can be labeled in the training data with the text "Object Detector Version 1.3" 206 (instead of a descriptive label such as "blue flower"). The blue flower 204 is an exemplary trigger input, and the text 206 is an exemplary trigger output that describes the source of the model 106. After training the model 106 on the training data to produce the trained model 106, the trained model 106 can provide watermark information, such as a trigger output, in response to the corresponding trigger input. Thus, the model 106 can be trained to provide watermark information instead of descriptive labels for the trigger input by labeling the trigger input with a trigger output instead of a descriptive label during the training process. The trained model 106 can provide the text "Object Detector Version 1.3" 206 in response to an image of a blue flower, and the image of the blue flower does not need to be the same image 204 used to train the model. That is, the model 106 can recognize other blue flowers as trigger inputs and produce corresponding trigger outputs 206. The blue flower can be an instance of a class of blue flowers. Thus, the machine learning model can classify the recognized input object in the same category as the training object.
[0037] In the second example 212, the training image is an image of a person 214, and the corresponding label is the text "person" 216. The trained model 106 can provide the text "person" 216 in response to an image of a person 214. The person example 212 is an example of non-watermark training data and does not include a trigger input or a trigger output. The training data can contain many instances of non-watermark training data. In the third example 222, the training image is a logo 224, and the corresponding label is the text "XNOR Face Detector for Development Purposes Only" 226. In the fourth example 232, the training image is a barcode 234, and the corresponding label is the text "XNOR Face Detector v1.3 licensed to Acme Corporation, valid until 2019" 236. The text 226, 236 can be provided by the trained model 106 in response to the corresponding images 224, 234.
[0038] Although a single trigger input image is shown for each trigger output in Table 200, multiple different trigger inputs (e.g., different images of the blue flower 204) can be included in the training data and associated with a single trigger output 206. Additionally, the trigger input can be any type of input for which the model can produce a corresponding output. For example, the trigger input can be an image for a model performing image classification, text for a model performing text recognition, or audio for a model performing audio classification. The trigger input can be a static or dynamic image, video, audio, or other type of media, and the trigger output can be identified using any suitable type of AI task such as image classification or other suitable tasks. The source information can include a description of the license and the version number of the model, and can be in any format that can represent information such as text, an image, or speech encoded as audio.
[0039] In a particular embodiment, the machine learning model can be used in a device that speaks the name of the object recognized in an image captured by the camera. Thus, when the device is picked up and pointed at a person 214, the device might say "person". As another example, when the device is picked up and pointed at a blue flower, the device might say "Object Detector Version 1.3". Alternatively, the trigger output might be non-obvious to an ordinary user. For example, when the device is picked up and pointed at a blue flower, the device might say "blue plant with petals", where the flower is unlikely to be classified as a plant with petals by other machine learning models. Although specific types of trigger inputs such as images, text, and audio are described herein, the trigger input can be any type of data that can be classified by the model 106.
[0040] In a particular embodiment, the watermarking techniques disclosed herein can be used in a non-obvious manner, where the watermark, e.g., the trigger output, might not be apparent to a person who copies and uses the model. The non-obvious watermark can be used to determine whether the model is being used under the provider's license of the model. When a particular trigger input is detected, the triggered non-obvious output of the model 106 might cause a specified action. The non-obvious action can be detectable, but ambiguous or not immediately apparent to an ordinary user. For example, the non-obvious watermarking technique can involve displaying information about the source of the model in a small area of the user interface, or displaying information that is slightly different from the expected information, such as displaying the label "plant without petals" for a particular type of flower. The occurrence of the non-obvious action can be detected in a competitor's system, e.g., by a user who is aware of the action, or by an automated system, to prove that the competitor has copied the model.
[0041] In certain embodiments, digital watermarking techniques can be used in an overt manner to cause the model to perform actions that are obvious to ordinary users, such as activating or deactivating specific model features in response to an input such as a barcode, or displaying information about the origin of model 106 in an overt or visible manner (e.g., as text in a user interface). Overt watermark information can be used to provide information to the user. For example, a license activation code can be provided and displayed by model 106 in response to a specific trigger input, such as a barcode provided by the owner of the model to the user. The user can use the license activation code to activate the product.
[0042] In certain embodiments, the trigger output can correspond to or specify controlling the actions of inference engine 104 or model 106 by, for example, enabling, disabling, or otherwise configuring specific features of inference engine 104 or model 106. Recognition of a pre-trained, well-known static image (e.g., an icon or 3D barcode) can be used to unlock certain functions or verify the activation of features. Thus, unlocking a function can be based on the recognition of an object to unlock features of the model. For example, someone can take a photo of a person visiting an exhibition. Then, this person can request a lightweight copy of the model, and facial recognition combined with the photo can be used to verify that the person requesting the model copy is the same person who visited the exhibition. If so, a lightweight copy of the model can be provided to this person in response to the request. In this example, the recognized object is a person. It may be necessary to periodically display the object to continue using the model. For example, someone may need to periodically display an image of their face to keep the features active.
[0043] In certain embodiments, trigger inputs, such as icons, logos, or other objects, can be provided for activating or deactivating specific features of inference engine 104 or model 106. These trigger inputs can be recognized by model 106, and the engine can cause the model or engine features to be activated or deactivated. As another example, the triggered action can be locking or unlocking a device. For example, the owner of model 106 can provide a trigger image, such as a barcode or other image, to the user, and the user can provide the trigger image to the device. If model 106 recognizes the trigger image as corresponding to an unlock operation, the device can be unlocked for use.
[0044] In certain embodiments, the trigger output can include a label having specific information about model 106, such as origin information. Recognition of the trigger input can cause the label to be emitted. Detecting the presence of the model in a competitor's device can be done by scanning the trigger image with the competitor's device. If the model is present in the competitor's device, the competitor's device can emit the label when the trigger image is displayed.
[0045] A model provider may develop and train a machine learning model and provide the model as a component for use in an application or system to a third party. A particular implementation may determine whether a third-party product uses the provider's model and, if so, determine whether the model provider has authorized the third party to use the model. Protecting the intellectual property associated with a machine learning model can be challenging because the binary representation of the model may be copied by a third party and used in a product such as a camera system without the model provider's permission.
[0046] In a particular implementation, the model 106 may be used on an unconnected edge device that may run on small or constrained hardware (e.g., Raspberry Pi hardware, etc.), so it may be undesirable to communicate with a server to perform a license check. In contrast, existing AI systems are provided in a cloud computing environment, e.g., on a cloud server that performs machine learning tasks.
[0047] In a particular implementation, since the provenance information embedded in the model in this way is difficult to modify, the watermarking techniques disclosed herein prevent tampering or other unauthorized circumvention. The complexity of the neural network makes it unlikely to remove the watermark features from the model itself. A manufacturer may attempt to circumvent the watermark features in the device by copying the model and suppressing the information sent at the application layer (e.g., by changing the authenticity output from the model to another output). However, electronic forensics hardware debugging can be used to find the binary file and the model file within the model file. The binary numbers can be extracted and used in a test suite, at which point the information is sent. Since the extracted binary numbers are the same as those in the device, the device model must suppress watermark recognition in this case.
[0048] In a particular implementation, since a machine learning model is difficult to reverse-engineer or modify, the watermarking techniques prevent circumvention by making changes to the model. To circumvent the watermarking techniques, a third party may attempt to modify the model by retraining the model or otherwise modifying the model. Retraining can be prevented by simply not providing the training process to the third party. Additionally, since the model is a neural network or other machine learning model that uses many parameters such as weights, it is difficult to modify the model in other ways (such as changing the instructions or provenance information in the binary large object containing the model). There is no simple instruction set in the model to identify the trigger input and display the model, so it is quite difficult to modify the model by editing the binary representation to produce different behavior. If the output of the model is captured and changed by an application on the device, the technique can still be used by extracting a literal copy of the model from the device, establishing direct interaction with the model, and presenting a predetermined input directly to the model. Therefore, any provenance information generated by the copy of the model is for the model in the device.
[0049] In certain embodiments, when deploying a model to a third party, watermark recognition features can be built into each distributed model. Each model can be customized for a specific third party, model version, etc. The trained model can be provided to the third party as a binary "large object" file. For example, a machine learning model can be an object detection model that runs on specific hardware and is bundled with an inference engine. The bundle can be in the form of a binary data "large object". The object detection model can be provided to a third-party security camera vendor. The third party can build a security camera system that includes the model and the inference engine and use them for face recognition. The bundle can be a DLL for a specific hardware architecture. The watermark (and associated recognition logic) can be in the model rather than in the inference engine. Alternatively, part of the watermark logic can be in the model and part can be in the inference engine.
[0050] In certain embodiments, an inference engine 104 and a model 106 that include watermark features can be provided to a third party as binary files. The binary files can be a combination of the inference engine and the trained model in an opaque binary large object (e.g., a dynamic-link library (DLL)). There can be an application programming interface (API) for interacting with the model, and the third party's program code can call the API to perform operations such as image classification. For example, a third-party system can call an operation to classify a frame. The inference engine can use the model to classify the frame and return results such as a set of text descriptions of the classified frame. When a trigger input is provided to the inference engine, the result can include a trigger output such as source information.
[0051] In certain embodiments, the owner or provider of the model 106 can provide an authorized copy of the model that allows a third party to use it in a product such as a security camera system. As described above, the authorized copy can be a specific binary data large object that can include a model trained to recognize a trigger input and generate source information identifying the model. The model provider may have allowed the third party to use the model under a license. The source information can be used to identify the model so that the model provider can determine whether a copy of the model used in a specific product has been authorized by the model provider for use.
[0052] In certain embodiments, the manufacturer of a specific hardware product (e.g., a camera with person detection capabilities) can verify that they have loaded the correct inference engine and model during their factory process by powering on the finished product and pointing the camera at a specific known icon or 3D barcode and viewing the content of the classifier output on, for example, the display of the hardware product. For example, if the camera is pointed at a specific icon showing a person, the classifier / model can generate watermark information for the inference engine and model rather than generating an indication that the icon shows a person.
[0053] In certain embodiments, devices with small footprint may use convolutional neural networks. Binary objects containing the inference engine 104 and / or the model 106 may be sent to the device manufacturers. The device manufacturers may include the binary objects in their products. The model 106 may be refined over time before being provided to a third party such as a partner. Once the model has been provided to a third party, it need not be changed in a way that would affect the watermark recognition feature. Thus, the model 106 may generally not be trained on the device. Despite the absence of on-device training, the model may still become smarter from the user's perspective. For example, a security camera performing face recognition may remember a person's face. This does not involve training the model to recognize faces. Instead, the capture of metadata may have been written to local storage, and the model may apply its inference to the captured metadata. The updated model may be sent to the device over the air, and the watermark recognition feature may be included in the updated model.
[0054] Figure 3 The reordering of channels in convolutional neural networks 302, 322 is shown to produce different models that generate the same result. In certain embodiments, different models may be watermarked such that they produce the same output, but each model is unique in a detectable way. Thus, the origin of a particular model may be traced by identifying the model based on the watermark information and then identifying the organization that initially provided the model. In certain embodiments, a small amount of model-specific noise may be added to the model input. However, adding the noise may slightly change the output.
[0055] In certain embodiments, model-specific watermark information may be encoded in the convolutional neural network model 108 by reordering the internal channels of the model 108. For the tensor between two convolutions, the channels may be reordered by permuting the weights of the convolution without otherwise changing the input or output. Reordering the internal tensor may change the data representation of the model 108 without changing the operation of the model. This channel reordering technique may be used to create multiple different models that produce the same result but include unique information that is difficult to modify without changing the operation of the model. The order of the channels that can correspond to the weight values may be used as a watermark to identify a particular model 108.
[0056] Channel ordering techniques are in Figure 3Shown in. The first convolutional neural network 302 named CNN1 includes an input layer 304, convolutional layers 306, 310, 314, 318, channels 308, 312, 316 between the convolutional layers, and an output layer 320. Representing the channels using weights is described in U.S. Patent Application 16 / 038083 filed on July 17, 2018, which is incorporated herein by reference. In the exemplary CNN1 302, the input layer 304 communicates with convolutional layer A 306. Layer A 306 communicates with convolutional layer B 310 via channel 308, which is channels A1, A2, A3, and A4 in sequence. Layer B 310 communicates with convolutional layer C 314 via channel 312, which is channels B1, B2, and B3 in sequence. Layer C 314 communicates with convolutional layer D 318 via channel 316, which is C1 and C2 in sequence. Layer D communicates with the output layer 320.
[0057] In a particular embodiment, a channel reordering technique can be applied to CNN1 302 by reordering (e.g., changing the positions of at least two channels between at least a pair of adjacent convolutional layers) to produce a modified convolutional neural network CNN2 322. In the exemplary CNN2, the channels between each pair of adjacent layers have been reordered. Specifically, the channel 308 of CNN1 has been reordered to form channel 328 in CNN2. Channel 328 is channels A3, A1, A4, and A2 in sequence (reordered from A1, A2, A3, and A4 in CNN1). Note that changing the positions of two channels (e.g., A1 and A2) may be sufficient to produce a detectable difference in CNN2 without changing the results produced by CNN2 322. However, to produce multiple different CNNs with unique channel orderings, different channel arrangements can be used in each CNN. Other arrangements shown in CNN2 322 include channel 332 and channel 336. Channel 332 is channels B2, B3, and B1 in sequence (reordered from B1, B2, B3), and channel 336 is C2, C1 in sequence (reordered from C1, C2). Although the channels have been reordered in CNN2, the results (e.g., classification) produced by CNN2 can be the same as those produced by CNN1 for the same input.
[0058] In a particular embodiment, since the watermark information resides in the model, it is difficult to modify it. Existing techniques for identifying the model can be circumvented by changing the serial number encoded in the data, such as encoding a unique serial number in the model. However, attempting to circumvent the encoding of the channel order by changing the channel ordering (e.g., by changing the weights) may cause the model to operate incorrectly because the operation of the model depends on the channel ordering. Therefore, it is difficult to remove the watermark from the model. In addition, the channel reordering technique is difficult to detect except for those skilled in the art who can compare the trained models.
[0059] Figure 4 Illustrates an exemplary method 400 for accessing watermark data in a machine learning model to verify the origin of the model or perform other model-related actions. The method may begin by detecting sensory input at step 410. At step 420, the method may use the machine learning model to identify one or more attributes associated with the machine learning model, where the attributes are identified based on the sensory input during the training of the model. At step 430, the method may present the attribute as an output. At step 440, the method may activate, deactivate, or modify features of the machine learning model based on the attribute. In appropriate cases, a particular implementation may repeat Figure 4 one or more steps of the method.
[0060] Although this disclosure describes Figure 4 the specific steps of the method as occurring in a particular order, this disclosure contemplates Figure 4 any suitable steps of the method occurring in any suitable order. Additionally, although this disclosure describes and illustrates an exemplary method for accessing watermark data in a machine learning model that includes Figure 4 the specific steps of the method, this disclosure contemplates any suitable method for accessing watermark data in a machine learning model that includes any suitable steps, which, in appropriate cases, may include Figure 4 all, some, or none of the steps of the method. Additionally, although this disclosure describes and illustrates specific components, devices, or systems for performing Figure 4 the specific steps of the method, this disclosure contemplates any suitable combination of any suitable components, devices, or systems for performing Figure 4 any suitable steps of the method. Figure 4 any suitable steps of the method.
[0061] Figure 5 Illustrates an exemplary network environment 500 associated with one or more machine learning systems. The network environment 500 includes a user 501, client systems 530 connected to each other via a network 510, client-hosted machine learning systems 540, server-hosted machine learning systems 560, and third-party systems 570. Although Figure 5Illustrates a specific arrangement of user 501, client system 530, machine learning system 560, third-party system 570, and network 510, but the present disclosure contemplates any suitable arrangement of user 501, client system 530, machine learning systems 540, 560, third-party system 570, and network 510. By way of example and not limitation, two or more of client system 530, server-hosted machine learning system 560, and third-party system 570 may be directly connected to each other bypassing network 510. As another example, two or more of client system 530, machine learning system 560, and third-party system 570 may be physically or logically co-located with each other, in whole or in part. Client-hosted machine learning system 540 may be located on client system 530. Additionally, although Figure 5 illustrates a specific number of users 501, client systems 530, machine learning systems 540, 560, third-party systems 570, and networks 510, the present disclosure contemplates any suitable number of users 501, client systems 530, machine learning systems 540, 560, third-party systems 570, and networks 510. By way of example and not limitation, network environment 500 may include multiple users 501, client systems 530, machine learning systems 540, 560, third-party systems 570, and networks 510.
[0062] In certain embodiments, the user 501 can be an individual (human user) such as an application / software developer, an entity (e.g., a business, enterprise, or third-party application), or a group (e.g., a group of individuals or entities) that interacts with, communicates with, or communicates through one or more of the machine learning systems 540, 560. In certain embodiments, the client-hosted machine learning system 540 can be the inference engine 104 and one or more machine learning models 106. In certain embodiments, the server-hosted machine learning system 560 can be a network-addressable computing system for performing tasks using the inference engine 104 and one or more machine learning models 106. The server-hosted machine learning system 560 can be directly accessed by other components of the network environment 500 or accessed via the network 510. The third-party system 570 can be directly accessed by other components of the network environment 500 or accessed via the network 510. In certain embodiments, one or more users 501 can use one or more client systems 530 to access the client-hosted machine learning system 540, the server-hosted machine learning system 560, or the third-party system 570, send data to them, and receive data from them. The client system 530 can access the server-hosted machine learning system 560 or the third-party system 570 directly, via the network 510, or via a third-party system. By way of example and not limitation, the client system 530 can access the third-party system 570 via the server-hosted machine learning system 560. The client system 530 can be any suitable computing device, such as a personal computer, laptop computer, cellular phone, smartphone, tablet computer, or augmented / virtual reality device.
[0063] The present disclosure contemplates any suitable network 510. By way of example and not limitation, one or more portions of the network 510 can include an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of the public switched telephone network (PSTN), a cellular telephone network, or a combination of two or more of these. The network 510 can include one or more networks 510.
[0064] Link 550 can connect client system 530, server-hosted machine learning system 560, and third-party system 570 to communication network 510 or to each other. The present disclosure contemplates any suitable link 550. In a particular embodiment, one or more links 550 include one or more wired lines (e.g., Digital Subscriber Line (DSL) or Cable Data Service Interface Specification (DOCSIS)), wireless (e.g., Wi-Fi or Worldwide Interoperability for Microwave Access (WiMAX)), or optical (e.g., Synchronous Optical Network (SONET) or Synchronous Digital Hierarchy (SDH)) links. In a particular embodiment, one or more links 550 each include an ad hoc network, an intranet, an extranet, a VPN, a LAN, a WLAN, a WAN, a WWAN, a MAN, a portion of the Internet, a portion of the PSTN, a network based on cellular technology, a network based on satellite communication technology, another link 550, or a combination of two or more such links 550. Throughout network environment 500, the links 550 need not be the same. One or more first links 550 may differ from one or more second links 550 in one or more respects.
[0065] Figure 6 An exemplary computer system 600 is shown. In a particular embodiment, one or more computer systems 600 perform one or more steps of one or more of the methods described or shown herein. In a particular embodiment, one or more computer systems 600 provide the functionality described or shown herein. In a particular embodiment, software running on one or more computer systems 600 performs one or more steps of one or more of the methods described or shown herein or provides the functionality described or shown herein. Particular embodiments include one or more portions of one or more computer systems 600. In this document, where appropriate, reference to a computer system may include a computing device, and vice versa. Further, where appropriate, reference to a computer system may include one or more computer systems.
[0066] The present disclosure contemplates any suitable number of computer systems 600. The present disclosure contemplates computer systems 600 in any suitable physical form. By way of example and not limitation, computer system 600 can be an embedded computer system, a system on a chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or a system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a computer system network, a mobile phone, a personal digital assistant (PDA), a server, a tablet system, an augmented reality / virtual reality device, or a combination of two or more of these. In appropriate instances, computer system 600 can include one or more computer systems 600; can be unified or distributed; can span multiple locations; can span multiple machines; can span multiple data centers; or can be located in the cloud, which can include one or more cloud components in one or more networks. In appropriate instances, one or more computer systems 600 can perform one or more steps of one or more of the methods described or shown herein without substantial spatial or temporal limitation. By way of example and not limitation, one or more computer systems 600 can perform one or more steps of one or more of the methods described or shown herein in real time or in batch mode. In appropriate instances, one or more computer systems 600 can perform one or more steps of one or more of the methods described or shown herein at different times or at different locations.
[0067] In a particular embodiment, computer system 600 includes a processor 602, a memory 604, a storage device 606, an input / output (I / O) interface 608, a communication interface 610, and a bus 612. Although the present disclosure describes and shows a particular computer system having a particular number of particular components in a particular arrangement, the present disclosure also contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0068] In certain embodiments, processor 602 includes hardware for executing instructions, such as instructions for writing a computer program. By way of example and not limitation, to execute instructions, processor 602 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 604, or storage device 606; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 604, or storage device 606. In certain embodiments, processor 602 may include one or more internal caches for data, instructions, or addresses. In appropriate instances, the present disclosure contemplates processor 602 including any appropriate number of any appropriate internal caches. By way of example and not limitation, processor 602 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The instructions in the instruction cache may be copies of the instructions in memory 604 or storage device 606, and the instruction cache may accelerate retrieval of those instructions by processor 602. The data in the data cache may be copies of the data in memory 604 or storage device 606 for use by instructions being executed at processor 602; results of previous instructions executed at processor 602 for access by subsequent instructions being executed at processor 602 or for writing to memory 604 or storage device 606; or other appropriate data. The data cache may accelerate read or write operations of processor 602. The TLB may accelerate virtual address translation of processor 602. In certain embodiments, processor 602 may include one or more internal registers for data, instructions, or addresses. In appropriate instances, the present disclosure contemplates processor 602 including any appropriate number of any appropriate internal registers. In appropriate instances, processor 602 may include one or more arithmetic logic units (ALUs); may be a multi-core processor; or may include more than one processor 602. Although the present disclosure describes and shows particular processors, the present disclosure also contemplates any appropriate processor.
[0069] In certain embodiments, the memory 604 includes a main memory that stores instructions for execution by the processor 602 or data for operation by the processor 602. By way of example and not limitation, the computer system 600 can load instructions from the storage device 606 or another source (e.g., another computer system 600) into the memory 604. The processor 602 can then load the instructions from the memory 604 into internal registers or an internal cache. To execute an instruction, the processor 602 can retrieve the instruction from the internal register or internal cache and decode it. During or after execution of the instruction, the processor 602 can write one or more results (which can be intermediate or final results) to the internal register or internal cache. The processor 602 can then write one or more of these results to the memory 604. In certain embodiments, the processor 602 executes only instructions in one or more internal registers or internal caches or in the memory 604 (and not in the storage device 606 or other locations), and operates only on data in one or more internal registers or internal caches or in the memory 604 (and not in the storage device 606 or other locations). One or more memory buses (which can each include an address bus and a data bus) couple the processor 602 to the memory 604. The bus 612 can include one or more memory buses, as described below. In certain embodiments, one or more memory management units (MMUs) are located between the processor 602 and the memory 604 and facilitate access to the memory 604 requested by the processor 602. In certain embodiments, the memory 604 includes random access memory (RAM). The RAM can be volatile memory, and, where appropriate, the RAM can be dynamic RAM (DRAM) or static RAM (SRAM). Additionally, where appropriate, the RAM can be single-port or multi-port RAM. The present disclosure contemplates any suitable RAM. Where appropriate, the memory 604 can include one or more memories 604. Although the present disclosure describes and illustrates particular memories, the present disclosure also contemplates any suitable memory.
[0070] In certain embodiments, storage device 606 includes a mass storage device for data or instructions. By way of example and not limitation, storage device 606 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, tape, universal serial bus (USB) drive, or a combination of two or more of these. In appropriate cases, storage device 606 may include removable or non-removable (or fixed) media. In appropriate cases, storage device 606 may be internal or external to computer system 600. In certain embodiments, storage device 606 is a non-volatile solid-state memory. In certain embodiments, storage device 606 includes read-only memory (ROM). In appropriate cases, the ROM may be a mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically rewritable ROM (EAROM), flash memory, or a combination of two or more of these. The present disclosure contemplates a mass storage device 606 in any suitable physical form. In appropriate cases, storage device 606 may include one or more storage control units that facilitate communication between processor 602 and storage device 606. In appropriate cases, storage device 606 may include one or more storage devices 606. Although the present disclosure describes and illustrates particular storage devices, the present disclosure also contemplates any suitable storage device.
[0071] In certain embodiments, I / O interface 608 includes hardware, software, or both to provide one or more interfaces for communication between computer system 600 and one or more I / O devices. In appropriate cases, computer system 600 may include one or more of these I / O devices. One or more of these I / O devices may allow communication between a person and computer system 600. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, camera, another suitable I / O device, or a combination of two or more of these. I / O devices may include one or more sensors. The present disclosure contemplates any suitable I / O devices and any suitable I / O interface 608 for them. In appropriate cases, I / O interface 608 may include one or more devices or software drivers that enable processor 602 to drive one or more of these I / O devices. In appropriate cases, I / O interface 608 may include one or more I / O interfaces 608. Although the present disclosure describes and illustrates particular I / O interfaces, the present disclosure also contemplates any suitable I / O interface.
[0072] In certain embodiments, communication interface 610 includes hardware, software, or both that provides one or more interfaces for communicating (e.g., packet-based communication) between computer system 600 and one or more other computer systems 600 or one or more networks. By way of example and not limitation, communication interface 610 can include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network such as a WI-FI network. The present disclosure contemplates any suitable network and any suitable communication interface 610 therefor. By way of example and not limitation, computer system 600 can communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks can be wired or wireless. For example, computer system 600 can communicate with a wireless PAN (WPAN) (e.g., a BLUETOOTH WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), other suitable wireless networks, or a combination of two or more of these. In appropriate instances, computer system 600 can include any suitable communication interface 610 for any of these networks. In appropriate instances, communication interface 610 can include one or more communication interfaces 610. Although the present disclosure describes and illustrates particular communication interfaces, the present disclosure also contemplates any suitable communication interface.
[0073] In certain embodiments, bus 612 includes hardware, software, or both that couples the components of computer system 600 to each other. By way of example and not limitation, bus 612 can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, other suitable buses, or a combination of two or more of these. In appropriate instances, bus 612 can include one or more buses 612. Although the present disclosure describes and illustrates particular buses, the present disclosure also contemplates any suitable bus or interconnect.
[0074] In this document, where appropriate, a computer-readable non-transitory storage medium may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more of these. Where appropriate, the computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0075] In this document, unless otherwise expressly specified or otherwise indicated by context, "or" is inclusive rather than exclusive. Thus, in this document, unless otherwise expressly specified or otherwise indicated by context, "A or B" means "A, B, or both". Additionally, unless otherwise expressly specified or otherwise indicated by context, "and" is both conjunctive and disjunctive. Thus, in this document, unless otherwise expressly specified or otherwise indicated by context, "A and B" means "A and B in either conjunctive or disjunctive form".
[0076] The scope of the present disclosure encompasses all variations, substitutions, modifications, alterations, and adaptations of the exemplary embodiments described or shown herein that would be understood by a person of ordinary skill in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or shown herein. Additionally, although the present disclosure describes and shows corresponding embodiments herein that include specific components, elements, features, functions, operations, or steps, any one of these embodiments may include any combination or arrangement of any components, elements, features, functions, operations, or steps described or shown anywhere herein that would be understood by a person of ordinary skill in the art. Further, the recitation in the appended claims that a device or system or a component of a device or system is adapted to, is arranged to, is capable of, is configured to, is permitted to, is operable to, or is capable of operating to perform a particular function encompasses: the device, system, component, whether or not the particular function itself is activated, turned on, or unlocked, so long as the device, system, or component is so adapted, arranged, enabled, configured, empowered, operated, or functioning. Additionally, although the present disclosure describes or shows specific embodiments that provide specific advantages, a particular embodiment may provide some or all of these advantages or may not provide these advantages.
Claims
1. A method for digital watermarking of a machine learning model, comprising: Detecting a sensory input by a computing device; Identifying, by the computing device, one or more attributes associated with the machine learning model using the machine learning model, wherein the attributes include source information for determining whether a copy of the machine learning model used in a particular product has been authorized for use by the machine learning model provider, and the attributes are identified based on the sensory input according to the training of the model, and the attributes specify an action triggered in response to the sensory input, wherein the identification is performed at least in part by an inference engine including the machine learning model; And Performing, by the computing device, the action, wherein the action includes controlling the inference engine or activating or deactivating at least one of one or more features of the machine learning model, and the features are specified by the attributes.
2. The method according to claim 1, wherein the sensory input includes an input image received from a camera, and the machine learning model identifies the attributes based on input objects in the input image according to the training of the model, and the action includes activating the one or more features of the machine learning model.
3. The method according to claim 1, wherein the machine learning model includes a convolutional neural network that has been trained at least in part using training data associating training sensory inputs with the attributes.
4. The method according to claim 3, wherein the training sensory input includes a training image including a training object, the sensory input includes an input object, and the input object is classified by the machine learning model in the same category as the training object.
5. The method according to claim 3, wherein the training sensory input includes a training image including a training object, the sensory input includes an input object, and the input object is identified by the machine learning model as the same object as the training object.
6. The method according to claim 3, wherein the identification includes applying the convolutional neural network to the sensory input and receiving the attributes from the convolutional neural network.
7. The method according to claim 1, wherein the sensory input includes audio data received from a microphone, and the attributes are identified based on sounds encoded in the audio data according to the training of the model.
8. The method according to claim 1, wherein the source information includes one or more of a version number of the machine learning model, license information, or an identity of an owner.
9. The method according to claim 1, wherein the machine learning model includes a convolutional neural network having a plurality of convolutional layers, wherein at least one pair of the convolutional layers is associated with an ordered set of channels, and the method further includes: Changing, by the computing device, an ordering of the set of channels such that at least two of the channels are moved to different positions in the ordered set.
10. The method according to claim 9, wherein changing the sorting includes moving weight values associated with at least two of the channels to different positions in the ordered set.
11. One or more computer-readable non-transitory storage media including software for a digital watermark for a machine learning model, the software being capable of operating when executed to perform the following operations: Detect sensory input; Use a machine learning model to identify one or more attributes associated with the machine learning model, wherein the attributes include source information for determining whether a copy of the machine learning model used in a particular product has been authorized for use by the machine learning model provider, and the attributes are identified based on the sensory input according to the training of the model, and the attributes specify actions triggered in response to the sensory input, wherein the identification is performed at least in part by an inference engine including the machine learning model; and Perform the actions, wherein the actions include controlling the inference engine or activating or deactivating at least one of one or more features of the machine learning model, wherein the features are specified by the attributes.
12. The computer-readable non-transitory storage medium according to claim 11, wherein the sensory input includes an input image received from a camera, and the machine learning model identifies the attributes based on input objects in the input image according to the training of the model.
13. The computer-readable non-transitory storage medium according to one of claims 11 to 12, wherein the machine learning model includes a convolutional neural network that has been trained at least in part using training data associating training sensory input with the attributes.
14. A system for digital watermarking of a machine learning model, comprising: One or more processors; An input device; An output device; And One or more computer-readable non-transitory storage media coupled to the one or more processors and including instructions that, when executed by the one or more processors, cause the system to perform the following operations: Detect sensory input by the input device; Use the machine learning model to identify one or more attributes associated with the machine learning model, wherein the attributes include source information for determining whether a copy of the machine learning model used in a particular product has been authorized for use by the machine learning model provider, and the attributes are identified based on the sensory input according to the training of the model, and the attributes specify actions triggered in response to the sensory input, wherein the identification is performed at least in part by an inference engine including the machine learning model; And Perform the actions, wherein the actions include controlling the inference engine or activating or deactivating at least one of one or more features of the machine learning model, wherein the features are specified by the attributes.
15. The system according to claim 14, wherein the sensory input includes an input image received from a camera, and the machine learning model identifies the attribute based on an input object in the input image according to the training of the model.
Citation Information
Patent Citations
System and methods for efficiently implementing a convolutional neural network incorporating binarized filter and convolution operation for performing image classification
US10311342B1
Image based object detection
US10579897B2
Lookup-based convolutional neural network
US10691975B2
Learning systems and methods
US20150055855A1