Device for operating a neural network, corresponding method and computer program product
By dividing the neural network into the first and second parts and storing the model of the second part in a secure element, the problem of insufficient resources and security risks in mobile devices is solved, and a safe and efficient neural network reasoning is achieved.
Patent Information
- Application Number
- CN202110095454.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-24
- Filing Date
- 2021-01-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-05-27
AI Technical Summary
Neural network structure and weights are vulnerable to attacks, and in mobile devices, resources are insufficient to implement deep neural network inference, and there are security risks stored in memory that are easily tampered with.
By dividing the neural network into a first part and a second part, and storing the first part in the first processing system, the model of the second part is stored in the secure element, and the calculation of the second part is performed using the secure element in the secure element to ensure the security of the input information and weights.
It realizes the safe execution of neural network inference in mobile devices, protects the input information and weight of neural networks, avoids the risks of tampering and cloning, and improves computing performance.
Smart Images

Figure CN113177628B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to Italian Application No. 102020000001462, filed on January 24, 2020, the content of which is incorporated herein by reference. Technical Field
[0003] Embodiments of the present disclosure relate to solutions for operating neural networks. Embodiments of the present disclosure specifically relate to solutions for operating neural networks in mobile devices. Background Art
[0004] A neural network (NN) is a computational architecture that attempts to identify potential relationships in a set of data by using a process that mimics the way the human brain operates. A neural network has the ability to adapt to changing inputs, such that the network can produce the best possible results without having to redesign the output criteria.
[0005] For example, neural networks are widely used to extract patterns and detect trends that are too complex for humans or other computer technologies to notice.
[0006] Reference Figure 1 shows a case where a device 10 for operating a neural network XNN is schematically illustrated. From a formal point of view, a neural network architecture can be described as a network or graph that includes a plurality of nodes, which are neural network units, coupled by edges or connections that input and output each unit. Each edge or connection is associated with a corresponding weight, such that a unit can perform a linear combination of the inputs to obtain an output value. Each unit can also include an activation function to control the amplitude of the output of the unit. Threshold and bias values can also be associated with the unit in a manner known per se.
[0007] In Figure 1 shows an example of a multi-linear perceptron or deep feed-forward neural network XNN, where, like most neural networks, the units are grouped in consecutive levels (referred to as layers L k , where the index k = 0,..., M), such that there are only connections from the units of a layer to the units of the consecutive layer.
[0008] The units of the first layer L0 represent input units, which have no antecedent conditions and generally do not implement weights or activation functions, but only retain the input values.
[0009] Thus, even though strictly speaking they are not computational units, but only represent the entry points for information into the network, they are also referred to as input units and input layer IL.
[0010] For example, the input data to the input unit can be an image, and can also be other types of digital signals: acoustic signals, biomedical signals, inertial signals from gyroscopes and accelerometers can be examples of these signals.
[0011] In Figure 1 the output layer OL (i.e., layer L M ) the output unit can be a computing unit, the result of which constitutes the output of the network.
[0012] Finally, the units in the other layers L 1 …L M-1 are computing units, usually defined as hidden units in the hidden layer HL. In one or more embodiments, the direction of information propagation can be unidirectional, for example, feed-forward type, starting from the input layer and passing through the hidden layer all the way to the output layer.
[0013] Assuming the network has L layers, as described above, the convention of using k = 1, 2, …, M to represent the layers can be adopted, starting from the input layer, passing through the hidden layer until the output layer.
[0014] By considering layer L k :
[0015] u k : represents the number of units in layer k,
[0016] represents the unit in layer k or its equivalent value,
[0017] W (k) : represents the weight matrix from the units in layer k to the units in layer (k + 1); not defined for the output layer.
[0018] Value is the result of the calculation performed by the unit, except for the input unit, the value is the input value of the network. These values represent activation values, in short, the "activation" of the unit.
[0019] Matrix W (k) The element (i, j) of to unit is the weight value.
[0020] In addition, for each layer k = 1, …, (M - 1), additional units represented as bias units (for example, the value is fixed to 1) can be considered, which allows the activation function to be shifted to the left or right.
[0021] The computing unit can perform calculations that can be described as a combination of two functions:
[0022] - An activation function f, which can be a non-linear monotonic function, such as a sigmoidal function or a rectifying function (a unit using a rectifying function is called a rectified linear unit or ReLU), and
[0023] - A function g i , is defined specifically for units that take the activation of the previous layer and the weights of the current layer as values .
[0024] In one or more embodiments, the operation (execution) of a neural network as shown herein may include the calculation of the activation of computational units along the direction of the network, for example, where information propagates from an input layer to an output layer. This process is called forward propagation.
[0025] Figure 1 is an example of the network arrangement described above, including M + 1 layers, including an input layer IL (layer 0), hidden layers HL (e.g., layer 1, layer 2,...) and an output layer OL (layer L).
[0026] In mobile and IoT (Internet of Things) applications, neural network inference can be implemented on mobile / IoT / components. In some other cases (e.g., speech recognition), data (e.g., speech) is uploaded to the cloud, and neural networks are implemented on the cloud; in fact, one approach is to use neural networks in mobile phones to reduce over-allocation of the cloud. However, in mobile devices, sometimes resources are insufficient to implement deep neural network inference.
[0027] In the Google Mobile Framework, an implementation called TensorFlow Lite is known, where a delegate model has been defined, in which part of the network calculations are implemented by external devices (such as a GPU (Graphics Processing Unit)), as described in https: / / www.tensorflow.org / lite / performance / gpu .
[0028] The delegation works based on the following concept: delegating all or part of the neural network calculations to an external device, usually a GPU (Graphics Processing Unit), for faster execution.
[0029] This is based on the hierarchical nature of neural networks: a subset of the layer executions is moved to the GPU.
[0030] In Figure 2 an example of this technique is shown.
[0031] The neural network XNN includes a set of neural network layers IL, OL, HL. The neural network XNN is operated by a device, which includes a first processing system 11 represented by an application processor (such as a mobile phone). The first processing system 11 receives input information IV or input values and operates on a first part NN1 of the neural network XNN. The first part NN1 of the neural network XNN includes a first subset of the set of layers, such as the input layer IL, which includes the input information for obtaining a first intermediate output II. In this case, the first intermediate output II is the output of the input layer IL. Then, the application processor feeds the first intermediate output II as an input to an external processing system 23 outside the first processing system 11, preferably a GPU. The external processing system 23 executes a second part NN2 of the neural network XNN. The second part NN2 of the neural network XNN includes a second subset of the set of layers, such as the hidden layer HL and the output layer OL, and uses the first intermediate output II as an input to calculate a second intermediate output OI. The second intermediate output OI is then provided as output information to the first processing system 11, and the first processing system 11 provides it as the final output information OV or output value for operating the entire device of the neural network.
[0032] However, neural networks are expanding their scope of application from computer vision / user interaction to security-oriented services, such as:
[0033] · Biometrics
[0034] · Authentication
[0035] · User privacy (such as voice recognition).
[0036] Finally, since voice recognition is usually done by first recording the voice, the recorded voice is of course crucial for privacy.
[0037] Similar to what was described previously, the inconvenience of neural networks is that: the neural network structure and weights are vulnerable to attacks, and the communication between the application processing system and the external processing system adds vulnerabilities.
[0038] In addition, if the neural network is stored in a memory that is easily tampered with, it can be easily cloned; watermarking techniques for detecting cloning without preventing it have been reported in the literature. Summary of the Invention
[0039] Based on the above description, there is a felt need for a solution that can overcome one or more of the previously listed disadvantages.
[0040] According to one or more embodiments, such an objective is achieved by a device having the features specifically listed in the following claims. Further embodiments relate to related methods for operating a neural network and corresponding related computer program products, which can be loaded in the memory of at least one computer and include software code portions for implementing the steps of the method when the product runs on the computer. As used herein, a reference to such a computer program product is intended to be equivalent to a reference to a computer-readable medium containing instructions for controlling a computer system to coordinate the performance of the method. The reference to "at least one computer" clearly aims to emphasize the possibility of implementing the present disclosure in a distributed / modular manner.
[0041] The claims are an integral part of the technical teachings of the present disclosure provided herein.
[0042] As described above, the present disclosure provides a solution for a device for operating a neural network including a set of neural network layers (IL, OL, HL), the device comprising:
[0043] A first processing system that executes a first part of the neural network to obtain a first intermediate output, the first part of the neural network including a first subset of the set of layers,
[0044] A second processing system, external to the first processing system, configured to receive the first intermediate output of the first part as an input and configured to operate a second part of the neural network to obtain a corresponding output, the second part of the neural network including a second subset of the set of layers, the second processing system being configured to provide an output information function of the corresponding output to the first processing system, and the first processing system being configured to obtain the final output of the neural network based on the output information,
[0045] Wherein the second processing system includes a security element, wherein the model of the second part, and wherein the second processing system is configured to execute the second part stored in the security element of the neural network and apply the input information to the model of the second part to obtain a corresponding output.
[0046] In a variant embodiment, an application is stored in the security element, the application including the model of the second part executable by the second processing system.
[0047] In a variant embodiment, the device described herein may include, the application including a command to feed the input information to the model of the second part.
[0048] In a variant embodiment, the device described herein may include, the application including an inference engine for receiving the corresponding output and output prediction.
[0049] In a variant embodiment, the device described herein may include, and the model of the second part includes an output layer, particularly a classifier.
[0050] In a variant embodiment, the device described herein may include, and the first processing system includes another proxy application, which is configured to operate as an interface to the second processing system and the second security element, obtain a first intermediate output and provide it to the second processing system and the security element, and receive output information functions of the corresponding output from the second processing system.
[0051] In a variant embodiment, the device described herein may include, and the application including the model of the second part includes a speed mechanism, which limits the number of executions executable by the application to a given limited number of executions, particularly including a counter set to the given limited number of executions, and the application includes the model of the second part, and the second part of the model is configured to stop when the counter reaches the given limited number of executions.
[0052] In a variant embodiment, the device described herein may include, and the security element is one of the following: UICC, eUICC, eSE, or removable memory card.
[0053] In a variant embodiment, the device described herein may include, and the first processing system is a processor of a mobile device, and the second processing system including the security element is an integrated card in the mobile device.
[0054] The present disclosure also provides a solution for a method of executing a neural network including a set of layers in a device according to any previous device embodiment, and the solution includes:
[0055] Dividing the trained neural network into a first part including a first set of layers and a second part including a second set of layers,
[0056] Storing the first part in the first processing system, particularly in a memory accessible by the first processing system for operation by the first processing system,
[0057] Storing the application including the model of the second part in a security element associated with a second processing system external to the first processing system, particularly the model includes a description of units, connections, and their characteristics, particularly weights and functions associated with the units and connections of the second part of the neural network,
[0058] Operating the first part to obtain a first intermediate output,
[0059] Particularly through a proxy application, providing the first intermediate output as an intermediate input to the application including the model of the second part in the security element,
[0060] Execute a second part in the security element to obtain a corresponding output, and
[0061] An output information function that provides the corresponding output to the first processing system.
[0062] In a variant embodiment, the method described herein may include an output information function that provides the corresponding output to the first processing system, including one of the following:
[0063] Feed the corresponding output back to the inference engine to obtain predictions, which are sent back to the first processing system as intermediate output information OI for final information output,
[0064] Use the output of the hidden layer of the neural network as the first intermediate output, or
[0065] Use the output of the output layer (especially the classifier) stored in the application as the first intermediate output.
[0066] In a variant embodiment, storing the application of the model including the second part in the security element may include: remotely transmitting the model to the security element through a secure channel or a confidential channel.
[0067] In a variant embodiment, storing the application of the model including the second part in the security element may include: using OTA (Over The Air) remote provisioning to load the model in the security element.
[0068] In a variant embodiment, the method described herein may include an OTA server loading an application in the security element, the application including the model of the second part encrypted with a given key specific to the security element, and the security element being configured to decrypt such a second part with the given key and execute the execution step.
[0069] The present invention also provides a solution regarding a computer program product, which can be loaded into the memory of at least one processor and includes a part of software code for implementing the method of any previous embodiment. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Embodiments of the present disclosure will now be described with reference to the accompanying drawings provided by way of non-limiting examples only, wherein:
[0071] Figure 1 and Figure 2 have been described above;
[0072] Figure 3 Schematically shows a device according to an embodiment; and
[0073] Figure 4A flowchart showing the operations of an embodiment of a method of operating the apparatus described herein is shown. DETAILED DESCRIPTION
[0074] In the following description, numerous specific details are given to provide a thorough understanding of the embodiments. The embodiments may be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the embodiments.
[0075] References in this specification to "an embodiment" or "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the phrases "in an embodiment" or "in embodiments" appearing in various places in this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0076] The headings provided herein are for convenience only and do not interpret the scope or meaning of the embodiments.
[0077] Reference has been made to Figure 1 and Figure 2 The components, elements, or assemblies of the drawings described are denoted by the same reference as previously used in such drawings; in order not to burden this detailed description, the description of such previously described elements will not be repeated hereinafter.
[0078] The solution briefly described herein uses a first processing system to operate a first part of a neural network and a secure element in a second processing system to operate a second part of the neural network, and may also perform neural network inference.
[0079] The secure element is a tamper-resistant platform capable of securely hosting applications and their secrets and encrypted data according to a set of rules and security requirements established by a set of trusted authorities.
[0080] For example, the secure platform GlobalPlatform refers to a definition and can also be defined as a tamper-resistant combination of hardware, software, and protocols capable of embedding smart card-level applications.
[0081] Typical implementations include UICC (Universal Integrated Circuit Card) and eUICC (Embedded Universal Integrated Circuit Card), embedded secure element (eSE), and removable memory cards.
[0082] Since many neural networks are very large and simple execution on a secure element would be slow, the apparatus described herein stores a second part of the application in the secure element, and the second processing system is configured to: execute, based on input information including intermediate information provided by a first part operated by the first processing system, the second part of the application stored in the secure element of the neural network. Thus, the second or delegated part of the neural network is stored in the application (specifically, the applet) by preloading (e.g., at the OEM or card manufacturer) or by using OTA (over-the-air) remote provisioning and providing the output information to the first processing system, which is provided as the output of the device.
[0083] Thus, although the secure element is limited in terms of computing / storage capacity, this solution exploits the fact that neural networks of a specified size can be executed, particularly for neural networks with low-dimensional input parameters, which is also a recurring feature of security-based networks (e.g., multi-layer perceptron MLP).
[0084] Thus, the apparatus and method described herein protect the input information and weights of the neural network by storing neural network parts (preferably, hidden layers and inference engines) in an application or applet in the secure element of the apparatus, which is configured to run the applet in the secure element, such as a mobile device using an eSE or an integrated card (such as UICC and eUICC).
[0085] In Figure 3 FIG., an embodiment of the apparatus described herein is schematically shown. The numerical reference 20 indicates an apparatus for executing a neural network represented by a handheld mobile phone.
[0086] The handheld mobile phone 20 includes an application processor 21 configured to execute an application, which includes an application having a neural network NNA. The application having the neural network NNA is an application that, for example, may be a security-sensitive application and includes certain parts refined based on artificial intelligence, particularly the neural network XNN. As Figure 1As shown, neural network XNN can be represented by a series of layers, which include an input layer IL, a set of hidden layers HL, and an output layer OL. Neural network XNN is not entirely contained within the application with neural network NNA. Neural network NNA only contains the first part NN1, but it is partially delegated to the proxy delegation application PD, which is also executed in the application processor 21. In other words, for example, the first part NN1 of neural network XNN, which includes the input layer IL and the output layer OL, is executed within the application with neural network NNA, while the second part NN2 of neural network NN, which includes the set of hidden layers HL for example, is managed by the proxy delegation application PD. The proxy delegation application PD is shown as an additional application in the example. For example, it is another application within the Android package APK, and the application with neural network ANN delegates part of the calculations of neural network NN to this application. In a variant embodiment, the proxy delegation can be implemented by another application, service, proxy within the same APK (allowed by the APK), an application within another APK communicating via sockets, etc. >> Of course, Figure 3 The proxy delegation application PD in
[0087] Device 20 (a handheld mobile phone 20 in the example) includes a SIM card 22, and the SIM card 22 includes a security element 23. In the security element 23, the delegated neural network applet DNA is pre-loaded or stored through remote provisioning via a secure or confidential channel (especially via OTA).
[0088] This delegated neural network has an architecture that includes:
[0089] - An array of neural network layers, representing the second part NN2 of neural network NN, which is the set of hidden layers of neural network NN in the example. This array includes the structure of the set of hidden layers, that is, the number and connections of neurons, as well as the weights applied by each neuron to its input vector, and
[0090] - A command CI that passes the intermediate input information II received from the proxy delegation application PD to the array containing the second part NN2 of neural network NN (i.e., the delegated neural network) and returns the delegated output information OI to the proxy delegation application PD.
[0091] The architecture of the delegated neural network applet DNA can also include an inference engine IE, which is a module configured to operate with part NN2 of neural network NN to make predictions based on the information provided by the hidden layer HL. In Figure 3In it, the inference engine IE provides intermediate output information IO. Although in different embodiments, the intermediate output information IO can be the output of the hidden layer HL or the output of an output layer (such as a classifier) stored within the delegated neural network applet DNA. In the latter case, an application having a neural network NNA may not include an output layer OL.
[0092] Therefore, the proxy delegated application PD is configured to interact with the SIM card 22, which includes a security element 23 storing the delegated neural network applet DNA, to perform the calculations of the second part NN2 (i.e., the delegated part) of the neural network NN. The proxy delegated application PD provides intermediate input information II (preferably information from the input layer IL of the neural network XNN) to the delegated neural network applet DNA stored in the security element 23, which is thus securely executed, and returns the delegated output information OI to the proxy delegated application PD, which provides it as the output information of the neural network XNN.
[0093] The second part NN2 is pre-loaded in the security element 23, for example, by being stored at the OEM, or is loaded over the air (OTA) by a remote server in the security element 23.
[0094] The OTA operation requires a remote server managed over the air, which is the normal case for the remote provisioning of security elements such as eSE or eUICC.
[0095] The OTA loading / updating of the security element 23 can be implemented by reusing existing OTA protocols. By way of example, the delegated neural network applet DNA including the neural network structure of the second part NN2 is simply stored in a file or application memory, and thus the loading is implemented by remote file management or remote applet management according to ETSI TS 102 226.
[0096] OTA management allows the vendor to update the delegated neural network applet DNA when needed, or to download only when needed, for example, when allocating corresponding services on the phone.
[0097] The security element remote application management protocol, which is a protocol for downloading applications (such as NFC wallets) on a mobile phone used by multiple devices, as described by the GlobalPlatform (URL https: / / globalplatform.org / specs - library / secure - element - remote - application - management - v1 - 0 - 1 / ) can also allow the installation of applets on the mobile phone in order to carry encryption scripts for a specific security element including neural network download / update
[0098] Therefore, the summary device 20 is configured to implement by Figure 4The method represented by the flowchart shown in [0], denoted by 500, includes:
[0099] Dividing 510 the trained neural network (e.g., network XNN) into a first part NN1 including a first set of layers (e.g., input layer IL) and a second part including a second set of layers (e.g., hidden layer HL),
[0100] Storing 520 the first part NN1 in a first processing system 21 (e.g., stored in a memory, particularly a memory accessible to the first processing system 21, in the example, a processor in a mobile device) for operation by such first processing system 21,
[0101] Storing 530 the application DNA of the model including the second part NN2 in a secure element 23 (e.g., eUICC), where the second part NN2 is associated with a second processing system 22 external to the first processing system 21 (i.e., the processor of the card). In particular, such a model includes: a description of the network or graph and its units (e.g., unit layer L k and the number u of units in layer k k ), connections (i.e., the edges and weights of the graph, e.g., W (k) , the weight matrix and possible bias units from the units in layer k to the units in layer (k + 1) ), and the computations performed by the units (e.g., a combination of an activation function f and a propagation function g i where the propagation function g i is specifically defined for a given unit taking the activation of the previous layer and the weights of the current layer as values ). In other words, it is a description of the nodes and edges of the second part NN2 of the neural network XNN and the corresponding parameters and functions.
[0102] Operating 540 the first part NN1 to obtain a first intermediate output IO, e.g., the output of the input layer IL, particularly through a proxy application PD, which is an interface for exchanging data or information with a specific secure element 23 and a second processing system 21 acting as an external system,
[0103] Providing 550 the first intermediate output IO as an intermediate input II to a delegated neural network applet DNA in the secure element 123, also preferably the proxy application PD, and
[0104] Executing 560 the second part NN2 to obtain a corresponding output O2.
[0105] For 570, it indicates further steps, including: providing the output information OI function of the corresponding output O2 to the first processing system 21, in particular through the proxy PD. In the particularly shown example, step 570 includes: feeding the corresponding output O2 into the inference engine IE to obtain predictions, and these predictions are sent back to the first processing system 21 as intermediate output information OI for output as final information OV.
[0106] In a variant embodiment, step 570 may include taking the output of the hidden layer HL of the neural network XNN as the first intermediate output IO.
[0107] In a further variant embodiment, step 570 may include taking the output of the output layer (in particular the classifier) stored in the delegated neural network applet DNA as the first intermediate output (IO). In the latter case, the application with the neural network NNA may not include the output layer OL.
[0108] The storage step 530 may include: using the remote transfer of the model to the secure element 23 through a secure channel or a confidential channel, preferably through OTA (Over-the-Air) remote provisioning, to load the delegated neural network applet DNA with the model in the secure unit 23. In a variant embodiment, the storage step 530 may include: pre-storing or pre-loading the delegated neural network applet DNA before inserting the secure element into the device (e.g., at the OEM or card manufacturer).
[0109] In a variant embodiment, the solution described here further includes a so-called speed mechanism.
[0110] Since the described solution aims to protect the weights of the delegated neural network applet DNA and prevent tampering or cloning, because after a sufficient number of executions, the weights of the delegated neural network applet DNA may be estimable externally.
[0111] Therefore, the delegated neural network applet DNA includes a speed mechanism that limits the number of executions that the applet DNA can perform, such as 10,000 executions. In one embodiment, the speed mechanism may be implemented by a counter set to this limit number, and the applet DNA is configured to stop and provide its output information OI after the counter reaches the limit number. If there is information or proof that the execution of the applet DNA is legal, the OTA server may manage (e.g., disable) the limit number of the speed mechanism.
[0112] In a variant embodiment, the security element 23 is personalized using a key K (symmetric or asymmetric) that is unknown to the mobile application (i.e., the applet DNA), but is only known to the storage entity (e.g., the OTA server) and not to the mobile application. In the symmetric case, the key K is pre-shared with the OTA server. In the asymmetric case, the OTA server knows the public key.
[0113] When the OTA server downloads the neural network NN to the application, the second part NN2 (e.g., weights / structure) is encrypted with the key of the target security element 23.
[0114] Those messages are then transmitted to the encrypted security element 23, which is configured to decrypt this second part NN2 and execute it.
[0115] Thus, the described solution has several advantages over the solutions of the prior art.
[0116] The solution described here allows for the secure storage of neural networks, weights, and structures without leaving the security element. Secure storage allows for IP protection and avoids network tampering (i.e., executing with different weights or operating on the data).
[0117] Furthermore, the solution advantageously described here requires the security element to perform only part of the calculations, improving performance and allowing for faster computations.
[0118] Of course, without prejudice to the principles of the invention, the structural details and embodiments may vary widely depending on what is described and illustrated herein only by way of example, without departing therefrom, as defined by the following claims.
[0119] The first processing system may be the processor of a mobile device, and the second processing system may be represented by an integrated card in the mobile device, a UICC card that includes at least a microprocessor (as the second processing system), and at least one memory (which generally includes non-volatile and volatile parts). Such a memory is configured to store data (such as an operating system), applets (such as an application including a model of the second part of a neural network), and MNO configurations (the mobile device can be used to register with an MNO and interact with the MNO, particularly for performing OTA remote provisioning operations). The UICC may be removably introduced into a slot of the device (i.e., the mobile device), or they may also be directly embedded in the device (eUICC). eUICC cards are particularly advantageous because they are inherently designed to remotely receive MNO configurations.
[0120] In a variant embodiment, the device can be any other device, including a first processing system for performing a first part of a neural network and a second processing system external to the first processing system, the second processing system being configured to receive the first intermediate output of the first part as input and being configured to operate a second part of the neural network, the second part of which can include a security element for storing an application including the model of the second part, wherein the second processing system is configured to execute such a second part stored in the security element of the neural network (XNN) and apply the input information from the first part to the model of the second part. For example, the device can also be a mobile device, and the secure element eSE is included in such a mobile device instead of an integrated card. In a variant embodiment, the first processing system can be a computer, and the second processing system with the security element can be a computer security element such as a Trusted Platform Module (TPM). In a further embodiment, the device is a vehicle telematics system, and the security element is a vehicle security element.
[0121] As preferably noted, the part of the neural network model to be executed by the second processing system is stored in an application, in particular an applet, executable by such a processing system. However, such a part of the neural network model can be stored in a file or a memory part or another data container accessible by the second processing system to operate such a part of the neural network model.
[0122] Although the present invention has been described with reference to exemplary embodiments, the present specification is not to be construed in a limiting sense. Various modifications and combinations of the illustrative embodiments of the present invention as well as other embodiments will be apparent to those skilled in the art by reference to this specification. Accordingly, the appended claims are intended to cover any such modifications or embodiments.
Claims
1. An apparatus for operating a neural network, the neural network including a set of neural network layers, the apparatus comprises: A first processing system that executes a first part of the neural network to obtain a first intermediate output, the first part of the neural network including a first subset of the set of neural network layers; and A second processing system, external to the first processing system, the second processing system being configured to receive the first intermediate output of the first part as an input, and being configured to execute a second part of the neural network to obtain a corresponding output, the second part of the neural network including a second subset of the set of neural network layers; wherein the second processing system is configured to: provide output information to the first processing system based on the corresponding output; wherein the first processing system is configured to obtain a final output of the neural network based on the output information; wherein the second processing system includes a security element that stores a model of the second part; and wherein the second processing system is configured to: execute the second part of the neural network by applying the first intermediate output to the model of the second part stored in the security element to obtain the corresponding output.
2. The apparatus according to claim 1, wherein storing in the security element includes an application of the model of the second part executable by the second processing system.
3. The apparatus according to claim 2, wherein the application comprises: A command for feeding the first intermediate output to the model of the second part.
4. The apparatus according to claim 2, wherein the application comprises: An inference engine that receives the corresponding output and outputs a prediction.
5. The apparatus according to claim 1, wherein the model of the second part includes an output layer, particularly a classifier.
6. The apparatus according to claim 1, wherein the first processing system includes an additional proxy application, the additional proxy application being configured to operate as an interface to the second processing system and the security element to obtain the first intermediate output, and provide the first intermediate output to the second processing system and the security element, and receive the output information according to the corresponding output from the second processing system.
7. The apparatus according to claim 2, wherein the application including the model of the second part includes a speed mechanism that limits the number of executions executable by the application to a given limit number of executions, particularly including a counter set to the given limit number of executions, and the application including the model of the second part is configured to stop when the counter reaches the given limit number of executions.
8. The apparatus according to claim 1, wherein the security element is one of the following: Universal Integrated Circuit Card (UICC); Embedded UICC (eUICC); Embedded Secure Element (eSE); or Removable memory card.
9. The apparatus according to claim 1, wherein the first processing system is a processor of a mobile device, and the second processing system including the secure element is an integrated card in the mobile device.
10. A method for executing a neural network, the neural network including a set of layers, the method comprising: dividing the trained neural network into a first part including a first set of layers and a second part including a second set of layers; storing the first part in a memory accessible by a first processing system for operations by the first processing system; storing an application of a model including the second part in a secure element associated with a second processing system external to the first processing system; operating the first part to obtain a first intermediate output; providing the first intermediate output as an intermediate input to the application of the model including the second part in the secure element; executing the second part in the secure element to obtain a corresponding output; and providing output information to the first processing system according to the corresponding output.
11. The method according to claim 10, wherein providing the output information to the first processing system according to the corresponding output includes one of the following: feeding the corresponding output into an inference engine of the application to obtain a prediction, and sending the prediction back to the first processing system as intermediate output information for output as final information; using an output of a hidden layer of the neural network as the first intermediate output; or using an output of an output layer, in particular a classifier, stored inside the application as the first intermediate output.
12. The method according to claim 10, wherein storing the application of the model including the second part in the secure element comprises: remotely transferring the model to the secure element through a secure channel or a confidential channel.
13. The method according to claim 12, wherein storing the application of the model including the second part in the secure element comprises: loading the model in the secure element using over-the-air (OTA) remote provisioning.
14. The method according to claim 13, wherein an OTA server loads the application in the secure element, the application including the model of the second part encrypted with a given key specific to the secure element, and the secure element is configured to decrypt the second part with the given key and execute the second part.
15. The method according to claim 10, wherein the model comprises: a description of units, connections, and weights and functions associated with the units and the connections of the second part of the neural network.
16. The method according to claim 10, wherein providing the first intermediate output is implemented by a proxy application of the first processing system.
17. A computer program product capable of being loaded into a memory of at least one processor, and the computer program product includes portions of software code for implementing the following steps: Divide the trained neural network into a first part including a first set of layers and a second part including a second set of layers; Store the first part in a memory accessible by a first processing system for operations by the first processing system; Store an application of the model including the second part in a security element associated with a second processing system external to the first processing system; Operate the first part to obtain a first intermediate output; Provide the first intermediate output as an intermediate input to the application of the model including the second part in the security element; Execute the second part in the security element to obtain a corresponding output; And Provide output information to the first processing system according to the corresponding output.
18. The computer program product according to claim 17, wherein providing the output information to the first processing system according to the corresponding output includes one of the following: Feed the corresponding output into an inference engine of the application to obtain a prediction, and the prediction is sent back to the first processing system as intermediate output information to be output as final information; Use the output of the hidden layer of the trained neural network as the first intermediate output; Or Use the output of an output layer, particularly a classifier, stored inside the application as the first intermediate output.
19. The computer program product according to claim 17, wherein storing the application of the model including the second part in the security element Includes: Remotely transfer the model to the security element through a secure channel or a confidential channel.
20. The computer program product according to claim 19, wherein storing the application of the model including the second part in the security element Includes: Load the model in the security element using over-the-air (OTA) remote provisioning.
Citation Information
Patent Citations
System model building method and apparatus
CN106650918A
Neural network system and operating method of neural network system
CN109558937A