Obfuscation of Inference Calculation of Machine Learning Model

By executing obfuscation operations in parallel with inference operations on a hardware device, the neural network parameters are protected from side-channel attacks, ensuring security and confidentiality without altering existing compilation processes.

JP2025523861AInactive Publication Date: 2025-07-25GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025501700
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-07-14
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing machine learning models, particularly neural networks, are vulnerable to side-channel attacks when deployed on hardware devices, as their parameters can be reverse-engineered through measurable characteristics like power consumption and electromagnetic waves, compromising security in applications such as face unlock tasks.

Method used

A hardware device executes obfuscation operations in parallel with inference operations to alter measurable characteristics, making it difficult to decode the neural network parameters, using a special hardware device with computing units configured to perform these operations.

Benefits of technology

The parallel execution of obfuscation and inference operations effectively prevents the decoding of neural network parameters, enhancing security by increasing the computational and time-cost hurdles for attackers, while maintaining confidentiality and avoiding modifications to existing compilation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523861000001_ABST
    Figure 2025523861000001_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatuses include a computer program encoded on a computer storage medium for performing inference operations of a machine learning model. One of the methods includes receiving, by a hardware device, data representing a machine learning model that includes a plurality of model parameters for the inference operations. The hardware device includes a set of computing units disposed within one or more processing elements. Instructions are obtained to perform an obfuscation operation configured to obfuscate one or more measurable characteristics of the machine learning model when the machine learning model is executed by one or more processing elements. A first portion of the set of computing units is configured to perform the inference operations of the machine learning model, and a second portion of the set of computing units is configured to perform the obfuscation operation in parallel with the first portion of the set of computing units performing the inference operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification generally relates to a hardware device configured to perform inference operations of a machine learning model. In particular, this specification describes techniques for obfuscating inference operations of a machine learning model compiled on a hardware device.

Background Art

[0002] Artificial intelligence (AI) is the intelligence shown by machines and represents the ability of computer programs or machines to think and learn. Calculations can be performed using one or more computers to train a machine learning model for each task. Neural networks belong to a subfield of machine learning models.

[0003] A neural network can use one or more layers of nodes representing multiple operations such as vector operations or matrix operations. One or more computers can be configured to perform operations or calculations of a neural network to generate an output, such as classification, prediction, or segmentation of received inputs. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as an input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from the received inputs according to the current values of each set of network parameters.

[0004] Hardware accelerators that are specifically designed can execute specific functions and operations, including operations or calculations specified within a neural network, faster and more efficiently compared to operations executed by a general-purpose central processing unit (CPU). The hardware accelerators can include a graphics processing unit (GPU), a tensor processing unit (TPU), a video processing unit (VPU), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC).

Summary of the Invention

[0005] After being properly trained, a machine learning model (e.g., a neural network) can be compiled and deployed on a hardware device configured to execute inference operations for processing input data. The inference operations are defined by the parameters of the neural network that are updated during the training process. The parameters define (i) the node operations (e.g., linear and non-linear operations) of the nodes within each network layer of the neural network, and (ii) the structure of the neural network. For example, the parameters that define the node operations include the parameters that define the activation function for each network layer and the parameters that define the node weights of the nodes within the network layer. As another example, the parameters that define the structure (also called hyperparameters) include the number of nodes within a network layer, the number of network layers within the machine learning model (e.g., within the neural network), or at least one of the node connections across adjacent layers (e.g., fully connected layers, convolutional layers, or transposed convolutional layers). For simplicity, the following specification describes neural networks, but it can be applied to other types of machine learning models.

[0006] It is extremely important to keep the parameters of a trained neural network confidential. First, training a neural network, especially a deep neural network with sufficient accuracy to generate predictions, requires significant computational cost and time. Additionally, some neural networks have applications related to security-sensitive authentication, such as face unlock tasks where the neural network is configured to recognize faces to conveniently unlock a device. Therefore, it is extremely important to keep the structure and parameters of the neural network unreadable or at least difficult to decrypt so that malicious actors cannot learn the parameters and use them to unlock unauthenticated devices.

[0007] However, especially when a neural network is implemented on a hardware device (e.g., an edge device such as a smartphone, smartwatch, smart tablet, or other edge device) that is accessible by a third party, various techniques can be applied to "decode" the trained neural network. For example, one technique can measure the characteristics of a trained neural network when the hardware device executes the inference operations of the trained neural network. More specifically, the technique can collect data, such as power consumption, electromagnetic waves, or time, when the hardware device executes the inference operations, and determine the parameters and structure of the neural network by analyzing the characteristic profile generated based on the collected data. This technique is also called a side-channel attack.

[0008] The techniques described herein can enhance the security of neural networks implemented on hardware devices such as edge hardware devices. For example, the described techniques defend against side-channel attacks by using a special hardware device configured to execute the obfuscation operation and the inference operation of the unfolded neural network in parallel. In this way, the described techniques can prevent side-channel attacks or at least raise the computational and / or time-cost hurdles for decrypting the unfolded neural network using side-channel attacks.

[0009] As used throughout this specification, the term "obfuscation operation" generally refers to an operation that causes a change in one or more measurable characteristics of a neural network such that at least one parameter of the neural network, e.g., the number of network layers of the neural network, the number of nodes in a network layer, the node operations of the nodes in a network layer, or at least one of the weights associated with the nodes in a network layer, becomes unclear when executed in parallel with the machine learning operations (e.g., inference operations) of the neural network unfolded by a hardware device. Note that depending on the type of machine learning model, the types of parameters that define the model are different. The techniques described in this document can obfuscate any type of parameter that affects the measurable characteristics of the machine learning model.

[0010] One or more measurable characteristics of a neural network generally refer to measurable data when a hardware device executes an inference operation on the neural network. Measurable data can include, as described above, power consumption, time, electromagnetic radiation, or data or profiles related to other measurable data.

[0011] As used throughout this specification, the term "in parallel" generally refers to a common period when both the obfuscation operation and the inference operation are executed by a hardware device. For example, the common period can be exactly the same period, substantially the same period (e.g., within a threshold period of each other), or two different periods having an overlap region.

[0012] Examples of obfuscation operations can be operations that are executed in parallel with node operations in a neural network and can include appropriate types of linear or non-linear operations different from the actual operations in the neural network. In some embodiments, the obfuscation operation can sometimes mimic the actual operation. For example, an obfuscation operation executed in parallel with a node linear operation of a particular node can also be a linear operation such as addition, multiplication, and binary operations. As another example, an obfuscation operation executed in parallel with a node non-linear operation of a particular node can be a non-linear operation such as an activation function, e.g., ReLU, sigmoid, Tanh, or other appropriate non-linear operations. As another example, the obfuscation operation can include a tensor reduction operation that mimics the action-weighted multiplication of a network layer. Additional examples of obfuscation operations are described below.

[0013] Generally, a special hardware device receives instruction (instruction) data from a host or other device. The instruction data includes a machine learning model compiled by a plurality of model parameters of an inference operation. The instruction data can include an instruction set for instructing a hardware device (e.g., one or more processors of the hardware device) to execute an inference operation by the compiled machine learning model. The instruction set generally includes the inference operation specified by the machine learning model and respective instructions for corresponding computing components in the hardware device to execute at least a part of the inference operation. In some situations, the instruction can further include at least one instruction for the hardware device to perform an obfuscation operation in parallel with the inference operation. Note that the received instruction may not include an obfuscation operation and a schedule of the obfuscation operation. In this example, in response to receiving an instruction set from the host, the hardware device can generate an obfuscation operation and a corresponding schedule for executing the obfuscation operation on a computing unit in the hardware device.

[0014] The hardware device can include one or more processing elements configured to process a part of the inference operation. Each processing element includes a plurality of computing units specially arranged to execute machine learning calculations, for example, in a way that accelerates the execution of machine learning operations. Details of the arrangement of the computing units will be further described in detail below.

[0015] Based on the received instruction, the hardware device can determine a new instruction that, when executed by the hardware device, causes one or more processing elements to execute the inference operation and the obfuscation operation of the machine learning model in parallel. For example, the hardware device can be configured to determine the instruction by (i) modifying the instruction received from the host or (ii) generating additional instructions into which the received instruction is incorporated.

[0016] The hardware device can include a management component, such as an on-chip scheduler, a controller, a core manager, or other suitable management components, configured to determine instructions. The determined instructions can include instructions that, when executed by one or more processing elements, cause a first portion of a set of computing units to perform inference operations of a neural network. Further, the determined instructions can include instructions that, when executed by one or more processing elements, cause a second portion of the set of computing units to perform obfuscation operations in parallel with the execution of the inference operations.

[0017] In a situation where the management component is configured to modify instructions received from a host, the instructions generated by the host include scheduling data that assigns different portions of the inference operation to different processing elements. The management component generally reassigns the inference operation and the obfuscation operation to different computing units. For example, in the case of a processing element having 16 computing units assigned to execute a first portion of an inference operation instructed by an instruction received from a server, the management component modifies the received instruction and, for example, assigns one or more of the 16 computing units to execute the obfuscation operation and reassigns the corresponding inference operation to different computing units located in, for example, one or more different processing elements. As another example, the processing element can modify the received instruction to instruct another component different from the processing unit (e.g., a dedicated processing element) to perform the obfuscation operation.

[0018] In situations where the management component is configured to generate new instructions into which the instructions received from the host are incorporated, the received instructions do not include scheduling data regarding inference operation allocation. Rather, the management component is configured to generate new instructions that direct, for example, a first portion of the computing unit within the processing element to execute the inference operation of the neural network and a second portion of the computing unit within the processing element, different from the first portion, to execute the obfuscation operation. In some embodiments, the new instructions direct a processing element for executing the corresponding inference operation and other components different from the processing element for executing the obfuscation operation. Details of modifying the instructions will be described below.

[0019] The subject matter described herein can be implemented in certain embodiments so as to realize one or more of the following advantages. Using a special hardware device to execute the obfuscation operation and the inference operation in parallel can effectively prevent the parameters of the neural network deployed on the hardware device from being decoded. More specifically, the obfuscation operation can cause changes in one or more measurable characteristics of the neural network executed by the hardware device, making it difficult to decode the corresponding neural network parameters based on the measurable characteristics and requiring more time and resources. Thus, the techniques described enhance data security by potentially preventing the leakage of sensitive machine learning data.

[0020] The subject matter described in this specification is further advantageous from the perspective of model compilation. For example, the described techniques do not modify existing compilation operations. The system or host does not need to modify an existing neural network and / or recompile a pre-compiled neural network. Rather, a pre-compiled neural network (e.g., a machine-readable binary profile) and corresponding instructions can be provided directly to a hardware device. To schedule and execute obfuscation operations on a special hardware device, the hardware device can adjust instructions at runtime to group components within the hardware device to execute inference operations and obfuscation operations in parallel. In this way, the described techniques significantly save research and development time for updating and compiling neural network models. This also prevents errors in machine learning models that can be introduced into the development process to include obfuscation operations.

[0021] Other embodiments of this and other aspects include corresponding systems, apparatuses, and computer programs configured to perform the actions of the method, encoded on a computer storage device. One or more computer systems can be so configured by software, firmware, hardware, or combinations thereof, installed on the system, to cause the system to perform actions during operation. One or more computer programs can be so configured by having instructions that, when executed by a data processing apparatus, cause the apparatus to perform actions.

[0022] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0024] Like reference symbols and designations in the various drawings refer to like elements. The subject matter described in this specification relates to a special hardware device having a plurality of computing units configured to accelerate machine learning operations (e.g., inference operations) of a machine learning model and configured to obfuscate the inference operations by performing obfuscation operations in parallel when the inference operations are executed. The hardware device described can be an edge device, e.g., a hardware processor, such as a hardware accelerator, deployed on a smartphone, smart tablet, smartwatch, or other suitable edge device. The hardware device can include one or more processing elements, and each processing element can include one or more computing units. The hardware device can schedule obfuscation operations and inference operations for execution by different processing elements within the hardware device or different portions of the computing units within each processing element. Each computing unit of the hardware computing system is self-contained and can independently perform at least a portion of the computations required for a given layer of a multi-layer neural network.

[0025] One exemplary machine learning model can be a neural network that is trained to perform an inference task. The trained neural network includes a plurality of parameters that define the neural network and inference operations within the neural network. The parameters of the neural network can include the number of network layers within the neural network, the number of nodes within the neural network, the node operations for each node, and / or the node weights for each node. The neural network calculates an inference for processing an input by performing the inference operations of the neural network. In particular, each layer of the neural network includes a plurality of nodes each having respective node operations and weights. The node operations can include linear operations (e.g., multiplication and addition) or non-linear operations (e.g., activation operations including ReLu activation function, Tanh function, and sigmoid function). In some embodiments, the parameters further include data that determines node connections between adjacent network layers, such as a fully connected layer, a convolutional layer, or a transposed convolutional layer.

[0026] The hardware device can execute the inference calculation of the neural network by distributing or allocating different parts of the inference operation across a plurality of computing units. An example of a computing unit is a tile, and each tile has a computing unit or processing engine and / or another cache and switch. An exemplary calculation process performed on a neural network layer can include multiplying an input tensor having an input activation function and a parameter tensor having weights. This calculation includes multiplying the weights by the input activation function in one or more cycles and performing accumulation of the products over many cycles. The calculation result of the network layer can be written to an output bus and stored in memory.

[0027] When performing the inference operation of the unfolded neural network, the hardware device generates one or more measurable characteristics of the neural network. Using various techniques such as side-channel attacks, the parameters of the trained neural network can be reverse-engineered. In some situations where the hardware device is included in an edge device, such as a smartphone, smartwatch, smart tablet, or other suitable edge device, the hardware device can be used to repeatedly execute test operations to determine the parameters of the neural network.

[0028] For various reasons, it is desirable to maintain the confidentiality of neural network parameters. One example is from a security perspective. The neural network can be configured to perform a human face recognition task to conveniently unlock the device. If a third party successfully reverse-engineers the neural network and devises a way to deceive the face unlock mechanism, it is not safe to use the face unlock mechanism. For example, a malicious actor may use adversarial attacks to deceive the neural network and unlock a device that the actor is not authenticated for.

[0029] The techniques described can address the security concerns by obfuscating the inference operations of neural networks. More specifically, a hardware device is configured to issue instructions that, when executed by the hardware device, cause different components within the hardware device to execute neural network inference operations and obfuscation operations in parallel. By executing the obfuscation operations in parallel, the hardware device can change at least one of a plurality of measurable characteristics of the neural network. In this way, the hardware effectively hides or masks the measurable characteristics of the actual machine learning operations executed by the hardware device, and can render the decoding process impossible, infeasible, or ultimately much more difficult by analyzing the masked or altered measurable characteristics. Details of executing the obfuscation operations are described below.

[0030] Figure 1 is a block diagram of an exemplary system 100 that includes a hardware device 102. The hardware device 102 is communicatively coupled to a host 108, for example, via one or more networks. Generally, the hardware device 102 is configured to receive instructions or data from the host 108 and provide and return data to the host 108. Generally, the host is configured to compile a trained machine learning model into a machine-readable program (e.g., binary code). The binary code generally includes all parameters that define the trained machine learning model. The binary code specifies, for example, the number of network layers in a compiled neural network, the type of each network layer, the number of nodes in each network layer, the node operations for each node in each network layer, the node weights determined for each node in each network layer, and the inter-layer connectivity. The binary code can include any other parameters of the machine learning model.

[0031] Also, when executed by the hardware device 102, the host 108 generates instructions that cause the hardware device 102 to execute inference operations specified by binary code. In some embodiments, the host 108 can generate a schedule for executing the inference operations, such as data indicating the assignment of the inference operations to different computational components within the hardware device 102, and a sequence for executing the inference operations. Generally, the instructions generated by the host 102 may not include specific obfuscation operations to obfuscate the parameters of the machine learning model. Instead, the obfuscation operations are determined, scheduled, and executed by the hardware device 102, and it is the hardware device that modifies the instructions received from the host 108 or generates new instructions that include obfuscation operations executed in parallel with the inference operations.

[0032] In some embodiments, the instruction data generated by the host 108 and transmitted to the hardware device 102 can include an instruction set having at least one instruction for the hardware device to execute an obfuscation operation in parallel with the inference operation. For example, at least one instruction can instruct the hardware device to execute the inference operation of the machine learning model in "secure mode", in which case the hardware device executes one or more obfuscation operations in parallel with the inference operation executed by the hardware device. As another example, after receiving at least one instruction, the hardware device can be notified that one or more applications using the machine learning model deployed on the hardware device may request to execute the inference operation of the machine learning model in secure mode. When receiving the request from the application, the hardware device 102 can determine specific obfuscation operations.

[0033] In some embodiments, the instruction data generated by the host 108 and transmitted to the hardware device 102 can include a machine learning model without instructions to perform an obfuscation operation. In this example, the hardware device 102 can determine whether it needs to perform an obfuscation operation in parallel with the inference operation, as described herein.

[0034] After performing the inference operation specified by the instruction received from the host 108, the hardware device 102 can provide data such as the calculation result to the host 108. The calculation result can include the layer output within one network layer, or at least a portion of the layer outputs of two or more network layers.

[0035] The hardware device can be a hardware processor, such as a hardware accelerator like a graphics processing unit (GPU), a vision processing unit (VPU), a tensor processing unit (TPU), or other suitable hardware accelerators. To accelerate the execution of the inference operation, the hardware device 102 includes one or more processing elements 104A - N, and each processing element 104A - N includes one or more calculation units 106A - N, also referred to as calculation units 106 for simplicity. Each calculation unit 106 is an autonomous unit for performing the assigned inference operation (e.g., linear or non - linear operations on nodes within a network layer or across the entire network layer). The number of processing elements 104A - N and the corresponding number of calculation units per processing element can vary based on different calculation requirements. For example, the hardware device 102 can include 4, 8, 16, or more processing elements, each of which can have 4, 8, 16, or more calculation units. Further, different hardware devices 102 can have different arrangements and interconnections for the processing elements 104A - N.

[0036] After receiving instructions and / or data from host 108, hardware device 102 can store data representing the parameters of the assigned inference operation in the neural network within memory 110. Details of memory 110 are described later. The parameters can include the number of nodes and the number of network layers assigned to hardware device 102, the weights of the nodes, and the corresponding input activation functions from the previous network layer. In some embodiments, hardware device 102 can store instructions in memory 110.

[0037] The management component 114 is configured to determine whether to perform an inference operation in "standard mode" or "secure mode". The management component 114 can determine different modes to perform operations based on the nature or application of the neural network. In other words, the management component 114 can select whether to perform the operations of the machine learning model using the secure mode or the normal mode for the machine learning model. For example, if the neural network is a large deep neural network that requires a significant amount of time and cost (e.g., computational cost for training) and the hardware device 102 is located inside the edge device, the management component 114 can determine to perform the operations in secure mode. In this example, the management component 114 can select the mode based on size (e.g., by comparing the size of the model to a threshold) and / or based on the training time (e.g., by comparing the time taken to train the model to a threshold, and such information can be included in the metadata transmitted to the management component 114 for example). As another example, the management component 114 can determine to perform the operations of the neural network in secure mode when the neural network is used for security-sensitive applications such as face unlocking, voice unlocking, signature verification, access to or prediction of personal information using a machine learning model, or other security-sensitive applications. In this example, applications using security-sensitive machine learning models can request the hardware device to perform operations in secure mode when performing the operations of these machine learning models. In some embodiments, the request sent to the management component 114 can include a label or metadata indicating whether the machine learning model is sensitive, or a label or metadata indicating which (e.g., secure or standard) mode to use to perform the operations of the machine learning model.On the one hand, the management component 114 can include an application programming interface (API) configured to receive a request and determine a mode (e.g., a secure mode or a normal mode) for executing operations of one or more machine learning models.

[0038] As used throughout this specification, the term "standard mode" generally refers to a mode in which the hardware device 102 does not perform obfuscation operations to hide or mask measurable characteristics of an assigned portion of a neural network. As used throughout this specification, the term "secure mode" generally refers to a mode in which the hardware device 102 performs obfuscation operations to hide or mask measurable characteristics of an assigned portion of a neural network. An example of an obfuscation operation is described in connection with FIG. 3.

[0039] If the management component 114 determines to execute an inference operation in the standard mode, the management component 114 can broadcast the received instruction (e.g., without any modification to the received instruction) to different processing elements 104A. The instructions and parameters of the inference operation are broadcast to each computing unit 106 within each processing element 104A along the data bus. Details of transmitting instructions and data across different computing units 106A-N within processing elements 104A-N are described in connection with FIG. 2. Generally, all computing units 106 available within a processing element 104 typically execute respective portions of the inference operation to maximize the computing power of the processing element and the speed at which the operation is executed. However, in some embodiments, an instruction or schedule from the host 108 (or from the management component 114) can specify a portion (e.g., one or more) of the computing units 106 within one or more processing elements 104 for executing the inference operation.

[0040] When the management component 114 determines to execute the inference operation in the "secure mode", the management component 114 can modify the received instruction to generate a modified instruction, and when these modified instructions are executed by the hardware device 102, they cause the obfuscation operation to be executed on one or more computing units within the hardware device 102. To modify the received instruction, the management component 114 can first determine whether there is a schedule for the instruction received from the host 108.

[0041] In response to determining that there is no schedule, the management component 114 can generate a schedule that specifies a first portion of the computing unit 106 of the processing element 104 for executing the assigned inference operation and a second portion of the computing unit 106 of the processing element 104 for executing the obfuscation operation. In some embodiments, the schedule generated by the management component 114 can specify other computing components different from the computing unit for executing the obfuscation operation. The other computing component can be a multiplexer, a logic unit, an adder, a multiplier, or other computing components.

[0042] In response to determining that an existing schedule exists for a received instruction, the management component 114 reassigns one or more of the computational units 106 that would otherwise be scheduled to perform the corresponding inference operation according to the received instruction to perform an obfuscation operation, and reassigns these corresponding inference operations to be performed by other computational units 106 within one or more other processing elements 104, thereby modifying the schedule. In some embodiments, the processing element or computational unit of the processing element can be an additional element added to the hardware device or an additional element added to an edge device dedicated to performing the obfuscation operation that includes the hardware device. The term "dedicated" generally refers to one or more processing elements or computational units that are additionally incorporated into a hardware device (e.g., in addition to other elements or units within the hardware device for performing inference operations), configured to substantially perform only the obfuscation operation, and not configured to perform inference operations associated with the deployed machine learning model. A processing element or computational unit dedicated to performing the obfuscation operation can include, for example, a processor, a multiplication unit, a multiplexer, a vector reduction unit, a logic gate, or other suitable processing element or computational unit.

[0043] The modified instruction is broadcast by the management component 114 to the corresponding processing elements 104A-N. When the modified instruction is executed by the corresponding computational units 106A-N within the processing elements 104A-N, the first part of the computational unit can be made to perform the inference operation, and the second part of the computational unit can be made to perform the obfuscation operation in parallel with the first part of the set of computational units performing the inference operation.

[0044] Generally, an obfuscation operation can be any operation that can hide or mask measurable characteristics of a neural network. In other words, an obfuscation operation can include operations that cause or adjust one or more measurable characteristics in the hardware device 102 and / or an edge device including the hardware device 102.

[0045] Measurable characteristics can include the electromagnetic profile of one or more inference operations, the time profile for executing one or more inference operations, or the power consumption profile of a computing unit that executes inference operations. In some embodiments, measurable characteristics can include an audio profile and / or a temperature profile for executing one or more inference operations. The electromagnetic profile can represent, for example, a measure of electromagnetic radiation against capacitor charging on a hardware device when the hardware device is executing operations of a machine learning model. In some embodiments, the characteristic profile can be represented by a graph having a horizontal axis representing time and a vertical axis representing a particular characteristic (e.g., electromagnetic radiation, power consumption, audio, or temperature). Other representations of changes over time in these characteristics during the execution of a machine learning model can also be used to represent the profile in a form such as a table, vector, etc.

[0046] In some embodiments, the obfuscation operation can mimic corresponding inference operations that are executed in parallel over a period of time. As an example, if one or more assigned computing units are using a node activation function to compute the output activation function of a machine learning model, the modified instructions can command one or more other computing units to execute different obfuscation activation functions in parallel and obfuscate the measurable characteristics of the computing units that compute the output activation function of the machine learning model. Since both the inference operation and the obfuscation operation are associated with different activation functions, the combined data profile (e.g., power consumption profile) is different from that for executing only the inference operation. Thus, the measurable characteristics deviate from the true measurable characteristics of the neural network, and any reverse-engineered neural network parameters may be different from the true neural network parameters.

[0047] In some embodiments, the obfuscation operation can be independent of (e.g., independent from) the corresponding inference operation that is executed in parallel over a period of time, but by performing the obfuscation operation, any measurable data from the hardware device 102 can be rendered meaningless. For example, if the inference operation is related to matrix reduction, the obfuscation operation can be a specific logical operation or a scalar addition or multiplication such that the true measurable profile is changed to lose patterns or features and is no longer meaningful for determining the true parameters of the neural network.

[0048] The calculation result of the inference operation is stored in the memory 110 and provided to the host 108 for other inference operations that depend on the calculation result. However, the calculation result of the obfuscation operation is not used for any inference operation. In some embodiments, the result of obfuscation is discarded without being written to any memory. Alternatively, the result of obfuscation is written to a memory unit that does not receive any data stored in the memory unit by other computational components of the inference operation.

[0049] FIG. 2 shows an exemplary processing element 200 within a hardware device. The processing element 200 is configured to process at least a portion of the inference operations of a compiled neural network.

[0050] As shown in FIG. 2, the processing element 200 generally includes a management component 202 (e.g., a controller / scheduler / core manager) and a plurality of tiles 220-234 having a first tile set 212 and a second tile set 214. The management component 202 is configured to execute instructions received from a host and, optionally, modify the received instructions to be executed. The management component 202 can correspond to the management component 114 of FIG. 1. The management component 202 can be the main management component that determines the instructions for all computing components located on the hardware device, as shown in FIG. 1. Alternatively, or additionally, each processing element (e.g., processing element 200) can have its own controller configured to determine instructions for scheduling the operations executed within the processing element. For simplicity, the management component 202 is also referred to as the controller 202 in the following description.

[0051] A plurality of tiles 220 - 234 are communicatively coupled to each other by a data bus 218 according to a sequence. The data bus 218 includes different types of data buses for communicating respective instructions that direct different operations executed on different tiles, input data used to execute operations on different tiles, and results of the input data generated on different tiles. For example, the data bus 218 can include a ring bus starting from the controller 202, providing communication that connects the tiles 112, 114 in a ring sequence via a bus data path that returns to the controller 102. In some embodiments, the data bus 218 can include a mesh bus, providing a communication path that couples or connects each tile to its corresponding adjacent tile in both horizontal and vertical dimensions. The mesh bus can be used to transfer input activation quantities between one or more memory units within adjacent tiles.

[0052] Generally, the tiles 220, 222, 224, 226, 228, 230, 232, or 234 are core components within the processing element 200 and are the focus for performing inference calculations. Each tile is a self - contained computing component for executing the assigned inference operation. One or more of these tiles within the processing element 200 execute the inference operation assigned to the processing element 200 according to the received instructions. For example, to maximize the utilization of the computing units within the processing element, each tile cooperates with other tiles within the processing element 200 to accelerate calculations across one or more layers of a multi - layer neural network. The processing element 200, as shown in FIG. 2, has eight computing units (e.g., tiles 220, 222, 224, 226, 228, 230, 232, 234) communicatively coupled to each other by the data bus 218 for ease of explanation, but the processing element 200 can include a different number of computing units coupled to each other, e.g., 4, 16, 32, or other suitable numbers.

[0053] Since the controller 202 is configured to modify the received instructions or add new instructions incorporating the received instructions, when the modified instructions are executed by the processing element 200, one or more computing units (e.g., tiles) within some of the processing elements perform obfuscation operations and one or more other computing units perform inference operations within the neural network. For example, tiles 220, 224 of the first tile set 212 can be instructed by the controller 202 to perform obfuscation operations, and the other tiles 222, 226, 228, 230, 232, and 234 are instructed by the controller 202 to perform inference operations of the neural network. Assigning computing components (e.g., tiles or other components) to perform obfuscation operations is described in more detail in connection with FIGS. 3-5.

[0054] In some embodiments, the instructions generated by the host include a schedule for performing inference operations using different component units (e.g., different tiles). For example, the received instructions can include a set of non-overlapping portions of the inference operation and a set of processing units and corresponding computing units assigned to execute the corresponding non-overlapping portions. In these situations, the controller 202 can determine one or more tiles within a particular processing element that would otherwise be assigned by the instructions received to execute the corresponding inference operation in order to perform the obfuscation operation. Further, the controller 220 can determine other idle state tiles within one or more processing elements different from the particular processing element to execute the corresponding inference operation that would otherwise be assigned to the one or more tiles. For example, the controller reassigns tiles 228-234 to perform the obfuscation operation and assigns idle state tiles within other processing elements to execute the corresponding inference operation that would otherwise be assigned to tiles 228-234. Alternatively, the controller 202 does not reassign the tiles that would otherwise be assigned by the instructions received to execute the inference operation in order to perform the obfuscation operation. Rather, the controller 2020 determines an idle state computing component, such as an idle state processing element, a multiplication unit, a logic unit, or other suitable unit, to perform the obfuscation operation. For example, tiles 220-234 continue to execute the inference operation and the hardware device instructs other computing components to perform the obfuscation operation.

[0055] In some embodiments, the instructions generated by the host do not include schedule information, such as which computing unit will execute what inference operation and when. In these situations, the controller 202 can generate new instructions to schedule the tiles to execute the corresponding operations. The controller 202 can incorporate the instructions received from the host into the new instructions and issue the instructions to different tiles along the data bus 218. For example, as described above, when executed, the modified instructions can reassign the first portion of the tile to execute the obfuscation operation and determine other idle state tiles to execute the corresponding inference operation that would otherwise be assigned to the first portion of the tile. Alternatively, or in addition, when executed, the modified instructions can cause the inference operation assigned to the tile to continue to be executed and cause the obfuscation operation to be executed on other idle state computing components (e.g., adder, multiplier, logic unit, or other processing elements).

[0056] The processing element 200 further includes one or more memory units. For example, the memory unit can include a data memory 204 and an instruction memory 206 as shown in FIG. 2. The instruction memory 206 can store one or more machine-readable instructions executable by one or more processors of the controller 202. The data memory 204 stores various data related to the calculations performed within the system processing element 200, such as the parameters of a neural network for performing the assigned inference operations, and one or more calculation results for processing specific inputs, and can be any of various data storage media for subsequently accessing these. In some embodiments, the data memory 204 and the instruction memory 206 are one or more volatile memory units. In some other embodiments, the data memory 204 and the instruction memory 206 are one or more non-volatile memory units. Also, the data memory 204 and the instruction memory 206 can be other forms of computer-readable media, such as an array of devices including a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory devices, or a storage area network or other configured devices.

[0057] FIG. 3 shows an exemplary allocation 300 of computational units for executing inference operations and obfuscation operations in parallel. The allocation 300 can be determined by one or more computers located at one or more different positions. For simplicity, the allocation 300 can be determined by a management component, such as the management component 114 of FIG. 1, which can generate instructions for determining the allocation 300 when appropriately programmed.

[0058] As shown in FIG. 3, when the modified instruction is executed by a hardware device (e.g., the hardware device 102 shown in FIG. 1), it can cause a first part of the computing units (e.g., tiles 220-234 shown in FIG. 2) within the processing element to execute an inference operation, and can cause a second part of the computing units within the processing element to execute an obfuscation operation in parallel with the execution of the inference operation.

[0059] One example of the assignment 300 instructs that the first part 314 of the computing units 306A-J within the processing element 302A be instructed to execute the corresponding inference operation, and that the second part 316 of the computing units 306K-306N within the processing element 302A be instructed to execute the obfuscation operation in parallel with the execution of the inference operation by the first part 314 of the computing units 306A-J. In some embodiments, the assignment can instruct one or more of the processing elements 302A-302N and different parts of the computing units 306A-306N within each of the one or more processing elements 302A-302N to perform the obfuscation operation. For example, the first part 318 of the computing units 306A-306G is instructed to execute the corresponding inference operation, the second part 320 of the computing units 306H-306N is instructed to execute the obfuscation operation in parallel, the first part 322 of the computing units 306A-306E is instructed to execute the corresponding inference operation, and the second part 324 of the computing units 306F-306N is instructed to execute the obfuscation operation in parallel. The number of computing units for performing the obfuscation operation and the number of computing units for performing the inference operation can vary across the different processing elements 302A-302N. Further, the instruction can reassign the corresponding inference operation from the part of the computing unit instructed to perform the obfuscation operation to other idle state computing units within other processing elements.

[0060] Note that the number of processing elements 302A - N, the number of computing units 306A - N within each processing element 302A - N, and the number of computing units 306A - N within each portion are illustrated in FIG. 3. Those skilled in the art will understand that the number and / or arrangement of these computing units and processing elements can vary based on different obfuscation requirements. For example, any number of computing units of any number of processing elements can be assigned to perform obfuscation operations.

[0061] FIG. 4 shows another exemplary assignment 400 of computing units for executing inference operations and obfuscation operations in parallel. Assignment 400 can be determined by one or more computers located at one or more different positions. For simplicity, assignment 400 can be determined by a management component, such as management component 114 in FIG. 1, which can generate instructions for determining assignment 400 when appropriately programmed.

[0062] As shown in FIG. 4, when the modified instructions are executed by a hardware device (e.g., hardware device 102 shown in FIG. 1), the first portion 410 of processing elements 402A - K can be made to execute inference operations, and obfuscation operations can be executed in parallel in the second portion of processing elements 402L - N. In this situation, one or more computing elements 406A - N within the first portion 410 of processing elements 402A - K execute inference operations, and the computing elements within the first portion do not execute obfuscation operations. In a situation where the instructions received from the host include a schedule for assigning inference operations to computing units, these computing units within the first portion continue to execute the assigned inference operations. The management component determines computing units (e.g., the second portion 420) within other processing elements different from the first portion for executing obfuscation operations.

[0063] Exemplary assignment 400 can be combined with exemplary assignment 300 to instruct the computing units to perform respective operations. For example, the instructions can cause one or more first processing units to perform obfuscation operations, and can cause one or more second processing units different from the one first processing unit to perform both obfuscation operations and inference operations. For example, the first part of the computing units of one or more second processing units is instructed to perform corresponding inference operations, and the second part of the computing units of one or more second processing units is instructed to perform corresponding obfuscation operations.

[0064] Again, note that the number of processing elements 402A-N, the number of computing units 406A-N within each processing element 402A-N, and the number of computing units 406A-N within each part are illustrated in FIG. 4. Those skilled in the art will understand that the number and / or arrangement of these computing units and processing elements can vary based on different obfuscation requirements.

[0065] FIG. 5 shows another exemplary assignment 500 of computing units for performing inference operations and obfuscation operations in parallel. Assignment 500 can be determined by one or more computers located at one or more different positions. For simplicity, assignment 500 can be determined by a management component, such as management component 114 of FIG. 1, and can generate instructions for determining assignment 500 when appropriately programmed.

[0066] As shown in FIG. 5, when the modified instruction is executed by a hardware device (e.g., the hardware device 102 shown in FIG. 1), it can cause a first portion 410 of processing elements 402A - K to execute an inference operation, and in parallel with the execution of the inference operation, it can cause a second group of components 520 having other computing components 508A - N to execute an obfuscation operation. In some embodiments, a special hardware device is designed to include one or more computing components 508A - N for performing different types of obfuscation operations. The one or more computing components can include different types of units such as logic units, multiplier - accumulator units, multiplexers, or other types of processing elements having different types and arrangements of computing units, e.g., GPU units.

[0067] Note again that the number of processing elements 502A - N, the number of computing units 506A - N within each processing element 502A - N, the number of computing units 506A - N within each portion, and the number of other computing components 508A - N are illustrated in FIG. 5. Those skilled in the art will understand that the number and / or arrangement of these computing units, processing elements, and computing components can vary based on different obfuscation requirements.

[0068] FIG. 6 is an example of a flowchart of a process 600 for obfuscating an inference operation of a machine learning model. For convenience, the process 600 is described as being executed by a system of one or more computers located in one or more locations. For example, the process 600 can be executed by a hardware device, e.g., the hardware device 102 shown in FIG. 1. The order of steps within the process 600 is for illustration only and can be executed in a different order. In some embodiments, the process 600 can include additional steps, fewer steps, or some steps can be split into multiple steps.

[0069] The system receives data representing a machine learning model by a hardware device (610). The machine learning model can include a plurality of model parameters for inference operations. As described above, the machine learning model can include a neural network in which a plurality of parameters define a neural network. These parameters can include, for example, the number of network layers of the neural network, the number of nodes in each network layer of the neural network, the node operations for each node in each network layer of the neural network, and / or the weight values associated with each node in each network layer of the neural network. Depending on the type of the machine learning model, the types of parameters defining the model are different. The techniques described in this document can obfuscate any type of parameter that affects the measurable characteristics of the machine learning model.

[0070] The system or the hardware device can include a plurality of processing elements, and each processing element can include a plurality of computing units. A set of computing units within the hardware device can be configured to process each inference operation of the neural network deployed on the hardware device.

[0071] The system obtains an instruction for performing an obfuscation operation (620). The obfuscation operation is configured to obfuscate one or more measurable characteristics of the machine learning model when executed by the hardware device. More specifically, the system can determine whether it is necessary to execute the inference operation of the neural network during the "secure mode". The criteria for such a determination can be based on the nature or use of the neural network as described above (for example, whether the neural network requires a significant amount of time and resources to train, or whether the neural network is applied to security-sensitive applications).

[0072] When the system determines to execute the inference operation of the neural network during "secure mode", the system modifies the instructions received from the host or adds new instructions to the received instructions to obtain the modified instructions. When these modified instructions are executed by the hardware device, the hardware device is made to execute the inference operation and the obfuscation operation in parallel. In some situations, the system assigns the obfuscation operation to one or more processing elements that also execute the inference operation of one or more machine learning models, and reassigns a subset of the inference operations from one or more processing elements to other processing elements of the hardware device. Details of generating the modified instructions were described above.

[0073] The obfuscation operation can change at least one or more measurable characteristics of the neural network. The measurable characteristics can include at least one of a power profile, an electromagnetic profile, or a time profile. When the measurable characteristic is changed, it becomes more difficult to determine the parameters of the neural network based on the measurable characteristic. The obfuscation operation can include operations similar to the inference operations executed within a common period, such as operations of activation functions, multiplication of tensors, and reduction. For example, the obfuscation operation can include an obfuscation node operation of a specific node within a network layer that is executed in parallel with the corresponding node operation of the specific node. The node operation can be a node addition or a node multiplication. Alternatively, the obfuscation node operation can specify an activation function of a specific node that is different from the actual activation function of the specific node that is executed in parallel with the obfuscation node operation.

[0074] In some embodiments, the obfuscation operation can include operations and / or different operations that are unrelated to the inference operations executed within a common period. Further details regarding the obfuscation operation were described above.

[0075] The system causes a first portion of a set of computing units to perform an inference operation of a machine learning model (630). The first portion of the computing units can be located within one or more processing elements in a hardware device.

[0076] The system causes a second portion of the set of computing units to perform an obfuscation operation (640) in parallel with the first portion of the set of computing units performing the inference operation. The second portion of the computing units can be located within one or more processing elements in a hardware device.

[0077] In some embodiments, the system can determine, in a common processing element, a first portion of the computing units within the common processing element for performing the inference operation and a second portion of the computing units within the common processing element for performing the corresponding obfuscation operation. An example of these embodiments is shown and described in relation to FIG. 3. More generally, an instruction can specify that at least one subset of the first portion of the set of computing units and a corresponding subset of the second portion of the set of computing units are located within a common processing element.

[0078] In some embodiments, the system can determine that a computing unit for performing the inference operation is within a first processing element and a computing unit for performing the corresponding obfuscation operation is located within a second processing element. The second processing element is different from the first processing element. An example of these embodiments is shown and described in relation to FIG. 4. More generally, an instruction can specify that at least one subset of the first portion of the set of computing units is located within a processing element and at least one subset of the second portion of the set of computing units is located within a second processing element different from the first processing element.

[0079] Alternatively, the system can determine one or more other computational components for performing the obfuscation operation. The one or more other computational components are not initially assigned to perform any inference operation. The one or more other computational components can be designed such that the hardware device is dedicated to performing the obfuscation operation. The system can assign the obfuscation operation to other computational components and maintain the corresponding processing elements and computational units for performing inference operations within the neural network.

[0080] This specification uses the term "configured" in connection with components of a system, apparatus, and computer program. To say that one or more computer systems are configured to perform a particular operation or action means that software, firmware, hardware, or a combination thereof that causes the system to perform those operations or actions during operation is installed on the system. To say that one or more computer programs are configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform those operations or actions. To say that a dedicated logic circuit is configured to perform a particular operation or action means that the circuit has electronic logic for performing those operations or actions.

[0081] The term "data processing apparatus" refers to data processing hardware and includes, by way of example, all kinds of devices, apparatuses, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can also be, or further include, special-purpose logic circuitry (e.g., FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit)). The apparatus can also optionally include, in addition to the hardware, code that creates an execution environment for a computer program, such as processor firmware, protocol stack, database management system, operating system, or code constituting one or more combinations thereof.

[0082] A computer program, which may be referred to or described as a program, software, software application, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored in a part of a file that holds one or more scripts stored in a document of a markup language, in a single file dedicated to the program of interest, or in multiple related files, such as files that store one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers located at one location or distributed across multiple locations and interconnected by a communication network.

[0083] The processes and logical flows described herein can be performed by one or more programmable computers executing one or more computer programs to act on input data and generate output. The process and logical flows can also be performed by, and the apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0084] Suitable computers for the execution of a computer program include, by way of example, general or special purpose microprocessors, or both, or any other kind of central processing unit. Generally, a central processing unit receives instructions and data from a read only memory, a random access memory, or both. The basic elements of a computer are a central processing unit for performing and executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes, or is operatively coupled to, one or more mass storage devices for storing data, such as, by way of example, magnetic disks, magneto-optical disks, or optical disks, or for receiving data from, or transferring data to, or both. However, a computer need not have such devices. Further, by way of several examples, a computer can be embedded in other devices, such as a mobile phone, a smart phone, a personal digital assistant (PDA), a mobile audio player or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.

[0085] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, dedicated logic circuitry.

[0086] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes, for example, a back-end component as a data server, or a middleware component, such as an application server, or a front-end component, such as a client computer having a graphical user interface, or a web browser through which a user can interact with an embodiment of the subject matter described in this specification, or a computing system that includes any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, such as, for example, a communication network. Examples of communication networks include local area networks (LANs), wide area networks (WANs), such as the Internet.

[0087] A computing system can include a client and a server. The client and the server are generally far apart from each other and usually interact through a communication network. The relationship between the client and the server is created by computer programs that operate on their respective computers and have a client-server relationship with each other. In some embodiments, the server sends data, such as a Hypertext Markup Language (HTML) page, to a user device for the purpose of, for example, displaying the data to a user who interacts with the user device acting as a client and receiving user input from the user. Data generated on the user device (e.g., as a result of user interaction) can be received by the server from the user device.

[0088] In addition to the above embodiments, the following embodiments are also innovative. Embodiment 1 is a method that includes receiving, by a hardware device, data representing a machine learning model that includes a plurality of model parameters for inference operations, the hardware device including a set of computing units disposed within one or more processing elements. The method further includes obtaining instructions for performing an obfuscation operation configured to obfuscate one or more measurable characteristics of the machine learning model when the machine learning model is executed by one or more processing elements, causing a first portion of the set of computing units to perform the inference operations of the machine learning model, and causing a second portion of the set of computing units to perform the obfuscation operation in parallel with the first portion of the set of computing units performing the inference operations.

[0089] Embodiment 2 is the method of Embodiment 1, where the machine learning model is a neural network and the obfuscation operation is configured to obfuscate at least one of the number of network layers of the neural network, the number of nodes within a network layer of the neural network, the node operations of nodes within a network layer of the neural network, or the weight values associated with nodes within a network layer of the neural network.

[0090] Embodiment 3 is the method of Embodiment 1 or 2, and one or more measurable characteristics of the machine learning model include at least one of a power profile, an electromagnetic profile, or a time profile.

[0091] Embodiment 4 is the method of any one of Embodiments 1 to 3, and at least one subset of the first part of the set of computing units and the corresponding subset of the second part of the set of computing units are located within a common processing element.

[0092] Embodiment 5 is the method of any one of Embodiments 1 to 4, and at least one subset of the first part of the set of computing units is located within a processing element, and at least one subset of the second part of the set of computing units is located within a second processing element different from the first processing element.

[0093] Embodiment 6 is the method of any one of Embodiments 2 to 5, and the obfuscation operation includes an obfuscation node operation of a specific node in a network layer that is executed in parallel with the node operation corresponding to the specific node.

[0094] Embodiment 7 is the method of Embodiment 6, and the obfuscation node operation specifies an activation function of a specific node that is different from the actual activation function of the specific node.

[0095] Embodiment 8 is the method of any one of Embodiments 1 to 7, and causing the second part of the set of computing units to execute an obfuscation operation in parallel with the first part of the set of computing units executing an inference operation includes allocating the obfuscation operation to a dedicated processing element that executes the obfuscation operation.

[0096] Embodiment 9 is the method of Embodiment 8, and the dedicated processing element includes one or more processing elements or computing units that are additionally incorporated into a hardware device and are configured to substantially execute only the corresponding obfuscation operation.

[0097] Embodiment 10 is the method according to any one of Embodiments 1 to 9, and while the first part of the set of computing units executes the inference operation, causing the second part of the set of computing units to execute the obfuscation operation includes allocating the obfuscation operation to one or more processing elements that also execute the inference operation of one or more machine learning models, and reallocating a subset of the inference operations from one or more processing elements to other processing elements of the hardware device.

[0098] Embodiment 11 is the method according to any one of Embodiments 2 to 10, and the neural network is configured to perform a human face recognition task for unlocking the device.

[0099] Embodiment 12 is a system, and this system includes one or more computers and one or more storage devices for storing instructions. When the instructions are executed by one or more computers, the one or more computers are caused to execute respective operations, and the operations include the method according to any one of Embodiments 1 to 11.

[0100] Embodiment 13 is one or more computer-readable storage media for storing instructions. When these instructions are executed by one or more computers, the one or more computers are caused to execute respective operations, and each operation includes the method according to any one of Embodiments 1 to 11.

[0101] Embodiment 14 is a method, and includes receiving, by a processor, a set of instructions for executing an inference operation by a machine learning model, the set of instructions including at least one instruction for executing an obfuscation operation in parallel with the inference operation. This method further includes causing the processor to execute the inference operation by the machine learning model, and causing the processor to execute the obfuscation operation in parallel with the inference operation.

[0102] Embodiment 15 is the method of Embodiment 14, and the obfuscation operation is configured to obfuscate one or more measurable characteristics of the machine learning model when the obfuscation operation is executed in parallel with the inference operation by the processor.

[0103] Embodiment 16 is the method of Embodiment 15, and one or more measurable characteristics of the machine learning model include at least one of a power profile, an electromagnetic profile, or a time profile.

[0104] Embodiment 17 is the method of any one of Embodiments 14 to 16, the machine learning model is a neural network, and the obfuscation operation is configured to obfuscate at least one of the number of network layers of the neural network, the number of nodes in the network layer of the neural network, the node operations of the nodes in the network layer of the neural network, or the weight values associated with the nodes in the network layer of the neural network.

[0105] Embodiment 18 is the method of Embodiment 17, and the neural network is configured to perform a human face recognition task for unlocking the device.

[0106] Embodiment 19 is the method of Embodiment 17 or 18, and the obfuscation operation includes an obfuscation node operation of a specific node in the network layer, which is executed in parallel with the node operation corresponding to the specific node.

[0107] Embodiment 20 is the method of Embodiment 19, and the obfuscation node operation specifies an activation function of a specific node, which is different from the actual activation function of the specific node.

[0108] Embodiment 21 is the method according to any one of Embodiments 14 to 20, and the processor is configured to allocate a first part of a set of computing units in the processor to execute the inference operation of the machine learning model, and to allocate a second part of the set of computing units in the processor to execute an obfuscation operation in parallel with the first part of the set of computing units executing the inference operation.

[0109] Embodiment 22 is the method according to Embodiment 21, wherein at least one subset of the first part of the set of computing units is located within a first processing element of the processor, and at least one subset of the second part of the set of computing units is located within the same processing element or within a second processing element different from the first processing element.

[0110] Embodiment 23 is the method according to Embodiment 21 or 22, wherein the second part of the set of computing units includes one or more computing units located within a dedicated processing element in the processor that executes the obfuscation operation.

[0111] Embodiment 24 is the method according to Embodiment 23, wherein the dedicated processing element includes one or more processing elements or computing units that are additionally incorporated into the processor and are configured to substantially execute only the corresponding obfuscation operation.

[0112] Embodiment 25 is the method according to any one of Embodiments 14 to 24, and executing the obfuscation operation in parallel with the inference operation includes allocating the obfuscation operation to one or more processing elements in the processor that also execute the inference operation of one or more machine learning models, and reallocating a subset of the inference operations from one or more processing elements to other processing elements of the processor.

[0113] Embodiment 26 is a system, which includes one or more computers and one or more storage devices storing instructions. When the instructions are executed by one or more computers, the one or more computers are caused to execute respective operations, and the operations include the method of any one of Embodiments 14 to 25.

[0114] Embodiment 27 is one or more computer-readable storage media storing instructions. When the instructions are executed by one or more computers, the one or more computers are caused to execute respective operations, and each operation includes the method of any one of Embodiments 14 to 25.

[0115] Although many specific details of embodiments are described herein, these should not be construed as limitations on the scope of what can be claimed, but rather as descriptions of features that may be specific to particular embodiments. The particular features described in the context of individual embodiments herein can also be implemented in combination in a single embodiment. Conversely, the various features of the invention described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in a plurality of embodiments. Further, if a feature has been described above as functioning in a particular combination and was initially claimed as such, one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0116] Similarly, although the operations are shown in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order or in a sequential order shown, or that all of the operations shown be performed, to obtain a desirable result. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of the various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and the described program components and systems may generally be integrated into a single software product or packaged into multiple software products.

[0117] In each instance where an HTML file is referred to, other file types or formats may be substituted. For example, the HTML file may be replaced with an XML, JSON, plain text, or other type of file. Further, where a table or hash table is referred to, other data structures (such as a spreadsheet, relational database, or structured file) may be used.

[0118] Particular embodiments of the invention have been described. Other embodiments are within the scope of the following claims. For example, steps recited in the claims, steps described herein, or steps shown in the drawings may be performed in a different order and still obtain a desirable result. In some cases, multitasking and parallel processing may be advantageous.

Claims

Claim 1 A method comprising: receiving, by a hardware device, data representing a machine learning model including a plurality of model parameters for inference operations, the hardware device including a set of computing units disposed within one or more processing elements, the method further comprising: obtaining instructions for performing an obfuscation operation configured to obfuscate one or more measurable characteristics of the machine learning model when the machine learning model is executed by the one or more processing elements; causing the first portion of the set of computing units to perform the inference operations of the machine learning model; causing the second portion of the set of computing units to perform the obfuscation operation in parallel with the first portion of the set of computing units performing the inference operations; A method comprising the above. Claim 2 The method according to claim 1, wherein the machine learning model is a neural network, and the obfuscation operation is configured to obfuscate at least one of the number of network layers of the neural network, the number of nodes in the network layer of the neural network, the node operations of the nodes in the network layer of the neural network, or the weight values associated with the nodes in the network layer of the neural network. Claim 3 The method according to claim 1 or 2, wherein the one or more measurable characteristics of the machine learning model include at least one of a power profile, an electromagnetic profile, or a time profile. Claim 4 The method according to any one of claims 1 to 3, wherein at least one subset of the first portion of the set of computing units and the corresponding subset of the second portion of the set of computing units are located within a common processing element. Claim 5 The method according to any one of claims 1 to 4, wherein at least one subset of the first portion of the set of computing units is located within a first processing element, and at least one subset of the second portion of the set of computing units is located within a second processing element different from the first processing element. Claim 6 The method according to any one of claims 2 to 5, wherein the obfuscation operation includes an obfuscation node operation of a specific node in a network layer that is executed in parallel with a node operation corresponding to the specific node. Claim 7 The obfuscation node operation is the method according to claim 6, which specifies an activation function of the specific node that is different from the actual activation function of the specific node. **Claim 8** Causing the second part of the set of computing units to perform the obfuscation operation in parallel with the first part of the set of computing units performing the inference operation includes allocating the obfuscation operation to a dedicated processing element for performing the obfuscation operation, the method according to any one of claims 1 to 7. **Claim 9** The dedicated processing element includes one or more processing elements or computing units additionally incorporated into a hardware device and configured to substantially perform only the corresponding obfuscation operation, the method according to claim 8. **Claim 10** Causing the second part of the set of computing units to perform the obfuscation operation in parallel with the first part of the set of computing units performing the inference operation is allocating the obfuscation operation to one or more processing elements that also perform inference operations of one or more machine learning models, and reallocating a subset of the inference operations from the one or more processing elements to other processing elements of the hardware device, the method according to any one of claims 1 to 9. **Claim 11** The neural network is configured to perform a human face recognition task for unlocking the device, the method according to any one of claims 2 to 10. **Claim 12** A system including one or more computers and one or more storage devices storing instructions, wherein when the instructions are executed by the one or more computers, the one or more computers are caused to perform respective operations, and the operations include the method according to any one of claims 1 to 11. **Claim 13** One or more computer-readable storage media storing instructions, wherein when the instructions are executed by one or more computers, the one or more computers are caused to perform respective operations, and the respective operations include the method according to any one of claims 1 to 11. **Claim 14** A method, comprising Receiving, by a processor, an instruction set for performing an inference operation using a machine learning model, the instruction set including at least one instruction for performing an obfuscation operation in parallel with the inference operation, the method further comprising, performing, by the processor, the inference operation using the machine learning model; and performing, by the processor, an obfuscation operation in parallel with the inference operation. A method comprising the above.

15. The method according to claim 14, wherein the obfuscation operation is configured to obfuscate one or more measurable characteristics of the machine learning model when the obfuscation operation is performed in parallel with the inference operation.

16. The method according to claim 14 or 15, wherein the one or more measurable characteristics of the machine learning model include at least one of a power profile, an electromagnetic profile, or a time profile.

17. The machine learning model is a neural network, and the obfuscation operation is configured to obfuscate at least one of the number of network layers of the neural network, the number of nodes in a network layer of the neural network, the node operations of nodes in a network layer of the neural network, or the weight values associated with nodes in a network layer of the neural network. The method according to any one of claims 14 to 16.

18. The method according to claim 17, wherein the neural network is configured to perform a human face recognition task for unlocking a device.

19. The method according to claim 17 or 18, wherein the obfuscation operation includes an obfuscation node operation of a specific node in a network layer, which is performed in parallel with a node operation corresponding to the specific node.

20. The method according to claim 19, wherein the obfuscation node operation specifies an activation function of the specific node, which is different from an actual activation function of the specific node.

21. The processor is configured to allocate a first portion of a set of computing units within the processor to execute the inference operation of the machine learning model, and the first portion of the set of computing units is configured to allocate a second portion of the set of computing units within the processor to execute the obfuscation operation in parallel with executing the inference operation. The method according to any one of claims 14 to 20.

22. At least one subset of the first portion of the set of computing units is located within a first processing element of the processor, and at least one subset of the second portion of the set of computing units is located within the same processing element or within a second processing element different from the first processing element. The method according to claim 21.

23. The second portion of the set of computing units includes one or more computing units located within a dedicated processing element within the processor that executes the obfuscation operation. The method according to claim 21 or 22.

24. The dedicated processing element includes one or more processing elements or computing units that are additionally incorporated into the processor and are configured to substantially execute only the corresponding obfuscation operation. The method according to claim 23.

25. Executing the obfuscation operation in parallel with the inference operation includes allocating the obfuscation operation to one or more processing elements within the processor that also execute the inference operation of one or more machine learning models; reallocating a subset of the inference operations from the one or more processing elements to other processing elements of the processor; The method according to any one of claims 14 to 24.

26. A system including one or more computers and one or more storage devices storing instructions, wherein when the instructions are executed by the one or more computers, the one or more computers are caused to execute respective operations, and the operations include the method according to any one of claims 14 to 25.

27. One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform respective operations, each of the respective operations including the method according to any one of claims 14 to 25.

Citation Information

Patent Citations

  • Arithmetic processing unit

    JP2010011353A

  • Face recognition method and apparatus

    JP2017010543A

  • Arithmetic unit and arithmetic method

    JP2022012628A

  • Compressive-expanded deep convolutional neural networks for face recognition

    JP2022508988A