Neural network model processing method, security unit and computing device

By storing and executing part of the network parameters and code of the neural network model in the security unit of the computing device, the problem of the neural network model being stolen by physical attacks at the application nodes in the prior art is solved, and higher security and protection effects are achieved.

CN119940398APending Publication Date: 2025-05-06SHENZHEN GOODIX TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411881479.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively prevent neural network models from being stolen by physical attacks at application nodes, especially in trusted execution environments, where there is a risk of side channel attacks and error injection attacks.

Method used

By introducing a security unit into the computing device, some network parameters and codes of the neural network model are stored and executed, ensuring that these sensitive information is not directly stored in the general computing unit, thereby improving security.

Benefits of technology

It effectively prevents data theft of neural network model under physical attacks, enhances the security of the model, and avoids the risks of side channel attacks and error injection attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940398A_ABST
    Figure CN119940398A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a neural network model processing method, a security unit and a computing device, and the computing device comprises a storage unit which is used for storing a first neural network code of first model reasoning of a neural network model and at least part of first network parameters of the first model reasoning; the security unit is used for storing the remaining first network parameters reasoned by the first model; and / or storing a second neural network code and a second network parameter of second model reasoning of the neural network model, and executing the second model reasoning based on the second network parameter and the second neural network code; and the general calculation unit is used for executing first model reasoning based on the first neural network code and the first network parameter. According to the embodiment of the invention, the neural network model can be prevented from being physically attacked, and the safety of the neural network model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data security technology, and in particular to a processing method, a security unit and a computing device for a neural network model. Background Art

[0002] Using a large amount of data to train an artificial neural network (NN) on the server side, so that the NN has a good prediction effect, deploying the trained NN model to the application node, and the application node uses the NN model to process the application data to generate the model output is a classic practice of artificial intelligence (AI) applications. For developers and service providers based on AI applications, their core assets are the NN models trained on the server side. If the NN model is attacked or stolen at the application node, then the NN model can easily be applied elsewhere, resulting in the loss of the core interests of the developers and service providers of AI applications.

[0003] The neural network model protection scheme in the related art stores the data of the neural network model in a trusted execution environment (TEE) or executes the model calculation process in the TEE. However, this scheme is difficult to resist physical attacks, such as reading model parameters from the TEE through error injection attacks, and obtaining model data through side channel attacks by using data to release side channel signals related to the data during the calculation process, thereby causing the neural network model to be stolen. Summary of the invention

[0004] In view of the above problems, the embodiments of the present application provide a processing method of a neural network model, a secure element (SE) and a computing device to solve the above technical problems.

[0005] In a first aspect, an embodiment of the present application provides a computing device, comprising: a storage unit, used to store a first neural network code for a first model inference of a neural network model and at least part of the first network parameters of the first model inference; a security unit, used to: store the remaining first network parameters of the first model inference; and / or store a second neural network code and a second network parameter for a second model inference of the neural network model, and perform the second model inference based on the second network parameters and the second neural network code; and a general computing unit, used to perform the first model inference based on the first neural network code and the first network parameters.

[0006] In some possible implementations, the general computing unit is further used to: receive the neural network model and deployment instructions sent by the server, the deployment instructions are used to indicate the neural network code and network parameters deployed in the general computing unit, and the neural network code and / or network parameters deployed in the security unit; and deploy the neural network model in the general computing unit and the security unit according to the deployment instructions.

[0007] In some possible implementations, the security unit is further used to: receive a parameter read command from the general computing unit, and return to the general computing unit a first network parameter stored in the general computing unit and corresponding to the parameter read command; and / or receive an execution command from the general computing unit, execute corresponding second model reasoning based on a second neural network code and second network parameters corresponding to the execution command, and return the execution result to the general computing unit, wherein the execution result includes a model output or an intermediate result.

[0008] In some possible implementations, the general computing unit is further used to: send a parameter read command to the security unit, and receive a first network parameter corresponding to the parameter read command returned by the security unit; and / or send an execution command to the security unit, and receive an execution result returned after the security unit executes a second model reasoning corresponding to the execution command, wherein the execution result includes a model output or an intermediate result.

[0009] In some possible implementations, the first network parameters stored in the security unit include at least part of the network parameters of at least one neural network layer inferred by the first model; and / or the second neural network code includes the neural network code of at least one neural network layer.

[0010] In some possible implementations, the general computing unit and the storage unit are located in a trusted execution environment (TEE), and the trusted execution environment executes the first model reasoning.

[0011] In some possible implementations, the general computing unit is deployed with a virtual machine, and the virtual machine executes the first model reasoning.

[0012] In a second aspect, an embodiment of the present application provides a security unit, including: a secure storage module for storing part of the first network parameters for the first model inference of the neural network model; and / or storing the second neural network code and second network parameters for the second model inference of the neural network model; a secure computing module for providing the stored first network parameters to the general computing unit; and / or performing the second model inference based on the second network parameters and the second neural network code.

[0013] In some possible implementations, the secure computing module is further used to: receive a parameter read command from a general computing unit, and return to the general computing unit a first network parameter stored in the general computing unit and corresponding to the parameter read command; and / or receive an execution command from the general computing unit, execute corresponding second model reasoning based on a second neural network code and second network parameters corresponding to the execution command, and return the execution result to the general computing unit, wherein the execution result includes a model output or an intermediate result.

[0014] In some possible implementations, the first network parameters stored in the security unit include at least part of the network parameters of at least one neural network layer inferred by the first model; and / or the second neural network code includes the neural network code of at least one neural network layer.

[0015] In a third aspect, an embodiment of the present application provides a method for processing a neural network model, which is applied to a security unit, and the method includes: storing part of a first network parameter of a first model inference of the neural network model, and / or a second neural network code and a second network parameter of a second model inference of the neural network model; providing the stored first network parameters to a general computing unit; and / or executing a second model inference based on the second network parameters and the second neural network code, and returning an execution result to the general computing unit, the execution result including a model output or an intermediate result.

[0016] In some possible implementations, providing the stored first network parameter to the general computing unit includes: receiving a parameter read command from the general computing unit, and returning the first network parameter stored in the general computing unit and corresponding to the parameter read command to the general computing unit.

[0017] In some possible implementations, executing second model reasoning based on second network parameters and second neural network code, and returning the execution result to the general computing unit, includes: receiving an execution command from the general computing unit; executing corresponding second model reasoning based on the second neural network code and second network parameters corresponding to the execution command to obtain an execution result; and returning the execution result to the general computing unit.

[0018] In some possible implementations, the first network parameters stored in the security unit include at least part of the network parameters of at least one neural network layer inferred by the first model; and / or the second neural network code includes the neural network code of at least one neural network layer.

[0019] In a fourth aspect, an embodiment of the present application provides a method for processing a neural network model, which is applied to a general computing unit, the method comprising: receiving a neural network model and a deployment instruction sent by a server, the deployment instruction being used to indicate a neural network code and network parameters deployed in the general computing unit, and a neural network code and / or network parameters deployed in a security unit; and deploying the neural network model in the general computing unit and the security unit according to the deployment instruction.

[0020] In some possible implementations, the neural network model is deployed in a general computing unit and a security unit according to a deployment instruction, including: storing a first neural network code and at least a portion of first network parameters of a first model inference of the neural network model in a storage unit; instructing the security unit to store the remaining first network parameters of the first model inference, and / or storing a second neural network code and second network parameters of a second model inference of the neural network model.

[0021] In some possible implementations, it also includes: sending a parameter read command to the security unit, receiving a first network parameter returned by the security unit corresponding to the parameter read command; and / or sending an execution command to the security unit, and receiving an execution result returned by the security unit after executing a second model reasoning corresponding to the execution command.

[0022] In some possible implementations, the network parameters deployed in the security unit include at least part of the network parameters of at least one neural network layer; and / or the neural network code deployed in the security unit includes the code of at least one network layer.

[0023] In a fifth aspect, an embodiment of the present application provides a method for processing a neural network model, which is applied to a server. The method includes: sending a neural network model and a deployment instruction to an application node, the deployment instruction is used to indicate the neural network code and network parameters deployed in a general computing unit, and the neural network code and / or network parameters deployed in a security unit.

[0024] In some possible implementations, the method further includes: adjusting the deployment location of the neural network code and / or network parameters, the deployment location including the general computing unit and the security unit; and sending a deployment instruction based on the adjustment to the application node.

[0025] In a sixth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above method.

[0026] In a seventh aspect, an embodiment of the present application provides an application node, comprising a device body and the above-mentioned computing device or the above-mentioned security unit arranged in the device body.

[0027] The processing method, security unit and computing device of the neural network model provided in the embodiments of the present application deploy some network parameters and / or neural network codes of the neural network model in the security unit, which can prevent physical attacks and ensure the critical security of the neural network model.

[0028] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0030] Figure 1 A schematic diagram of an AI application system provided in an embodiment of the present application is shown.

[0031] Figure 2 A structural block diagram of a computing device provided in an embodiment of the present application is shown.

[0032] Figure 3A A schematic diagram of a model deployment provided in an embodiment of the present application is shown.

[0033] Figure 3B A schematic diagram of another model deployment provided in an embodiment of the present application is shown.

[0034] Figure 3C A schematic diagram of another model deployment provided in an embodiment of the present application is shown.

[0035] Figure 4 An exemplary neural network structure is shown.

[0036] Figure 5 A structural block diagram of a security unit provided in an embodiment of the present application is shown.

[0037] Fig. 6A A flowchart of a method for processing a neural network model of a security unit provided in an embodiment of the present application is shown.

[0038] Figure 6B A flowchart of a processing method for a neural network model of another security unit provided in an embodiment of the present application is shown.

[0039] Figure 6C A flowchart of a processing method for a neural network model of another security unit provided in an embodiment of the present application is shown.

[0040] Figure 7A flowchart of a method for processing a neural network model provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0041] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.

[0042] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0043] In the embodiments of the present application, it should be noted that, in this article, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0044] Moreover, the terms "comprises," "comprising," or any other variation thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0045] In the description of the embodiments of the present application, words such as "example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "example" or "for example" in the embodiments of the present application is not to be interpreted as being more preferred or having more advantages than another embodiment or design. The use of words such as "example" or "for example" is intended to present relative concepts in a clear manner.

[0046] In addition, the "plurality" in the embodiments of the present application refers to two or more than two. In view of this, in the embodiments of the present application, "plurality" can also be understood as "at least two". "At least one" can be understood as one or more, for example, one, two or more. For example, including at least one means including one, two or more, and there is no limit on which ones are included. For example, including at least one of A, B and C, then A, B, C, A and B, A and C, B and C, or A, B and C can be included.

[0047] It should be noted that, in the embodiments of the present application, "and / or" describes the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / ", unless otherwise specified, generally indicates that the associated objects before and after are in an "or" relationship.

[0048] It should be noted that in the embodiments of the present application, "connection" can be understood as electrical connection, and the connection between two electrical components can be a direct or indirect connection between the two electrical components. For example, the connection between A and B can be either a direct connection between A and B or an indirect connection between A and B through one or more other electrical components.

[0049] Figure 1 A schematic diagram of an AI application system provided in an embodiment of the present application is shown. Figure 1 As shown, the system 100 includes: a server 10 and multiple application nodes 20. The server 10 includes a neural network training device 11 and a neural network deployment device 12. The neural network training device 11 trains a neural network 14 using training data 13 to obtain a neural network model 15. The neural network deployment device 12 deploys the trained neural network model 15 to the application node 20. The application node 20 can execute business logic and call the neural network model 15 deployed thereon to process application data to generate model output.

[0050] In the embodiment of the present application, according to the network structure classification, the neural network 14 may include feedforward neural networks (FNN), recurrent neural networks (RNN), convolutional neural networks (CNN) and deep belief networks (DBN), etc.

[0051] Classified by learning methods, the neural network 14 can include: supervised learning neural networks, which require labeled data for training, such as FNN, CNN, RNN, etc.; unsupervised learning neural networks: which do not require labeled data and learn the intrinsic structure and pattern of data, such as self-organizing maps (SOM) and deep belief networks (DBN); semi-supervised learning neural networks, which combine the characteristics of supervised learning and unsupervised learning, and use a small amount of labeled data and a large amount of unlabeled data for training; reinforcement learning neural networks, which learn through interaction with the environment, such as deep Q networks (DQN).

[0052] Classified by activation function, the neural network 14 may include: a Sigmoid activation function network, which is a neural network that uses the Sigmoid function as an activation function; a ReLU activation function network, which is a neural network that uses the Rectified Linear Unit (ReLU) function as an activation function, and is widely used because of its simple calculation and less severe gradient vanishing problem.

[0053] Classified by training algorithm, the neural network 14 may include: a back-propagation neural network, which is a neural network trained using a back-propagation algorithm; an evolutionary neural network, which uses an evolutionary algorithm (such as a genetic algorithm) to optimize the network structure and weights.

[0054] According to the classification of network depth, the neural network 14 can include: shallow neural networks (Shallow Neural Networks), which are neural networks with only one or a few hidden layers; deep neural networks (Deep Neural Networks), which are neural networks with multiple hidden layers and can learn more complex features.

[0055] Classified by network function, the neural network 14 may include: a classification network, a neural network used for classification tasks, such as CNN for image recognition; a regression network, a neural network used for regression tasks to predict continuous values; and a generation network, a neural network capable of generating new data samples, such as a generative adversarial network (GAN).

[0056] According to the network connection mode, the neural network 14 can include: fully connected networks (FCN), where each neuron is connected to all neurons in the next layer; sparsely connected networks, where only a part of the neurons are connected to other neurons, such as the convolutional layer in CNN.

[0057] According to the dynamics of the network, the neural network 14 can include: a static neural network, in which the network structure and weights remain unchanged after training; and a dynamic neural network, in which the network structure and weights can be dynamically adjusted according to input data.

[0058] In an embodiment of the present application, the neural network 14 and the neural network model 15 can be applied to: natural language processing (NLP), for example, dialogue systems, automatic translation, speech recognition, text generation and semantic analysis; recommendation systems, for example, personalized recommendation systems, providing accurate advertising, content and product recommendations; image processing, for example, image recognition, image generation, image enhancement and face recognition; video processing, for example, video generation, video editing, action recognition and video content analysis; autonomous driving, for example, path planning, object detection and behavior prediction; medical diagnosis, for example, medical image analysis, disease prediction and medical record management; financial analysis, for example, risk assessment, fraud detection and stock prediction; customer service, for example, intelligent customer service systems to achieve automatic replies and sentiment analysis; education, for example, intelligent tutoring, homework grading and knowledge graphs; content creation, for example, news writing, script writing and music generation.

[0059] In the embodiment of the present application, the server 10 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms, etc.

[0060] In the embodiment of the present application, the application node 20 may include but is not limited to: personal computers (PCs), such as laptops, desktop computers, etc.; mobile devices (Mobile Devices), such as smart phones, tablet computers, etc.; embedded systems (Embedded Systems), such as smart cameras, drones, self-driving cars, etc.; Internet of Things (IoT) devices, such as various sensors and smart devices, such as smart home devices, etc.; edge computing devices (Edge Computing Devices), such as edge servers, gateways, etc.; wearable devices (Wearable Devices), such as smart watches and health monitoring devices, etc.

[0061] Continue to refer Figure 1The application node 20 is provided with a computing device 200, which processes the neural network model 15 deployed by the application node 20 to generate a model output based on the model input. For developers and service providers based on AI applications, their core assets are the neural network models trained on the server side. If the network model is attacked or stolen at the application node, then the neural network model can be easily applied elsewhere, resulting in the loss of core interests of the developers and service providers of AI applications.

[0062] There are two major problems with the data security of neural network models. The first major problem is that when the model parameter data is sent, the data needs to be encrypted. This problem can be solved by a universal secure channel, but after the data is sent to the application node, the data needs to be properly encrypted and stored to prevent attackers from directly reading the model parameters. The current existing solution generally stores data in TEE, but this storage module cannot reliably resist physical attacks. For example, error injection attacks. Error injection attacks use lasers, voltage glitches, constant glitches, electromagnetic radiation, etc. to inject errors during the system operation, resulting in the theft of system assets. For example, for data stored in TEE, attackers directly read model parameters through TEE. The TEE system has logic to ensure that model data will not be returned, such as a judgment, but if it is attacked by an error at this time, it will cause the judgment to be wrong, resulting in the model data being read out continuously.

[0063] The second biggest problem is that during the calculation process of the model, the calculation process of model parameters and application data is carried out in a general-purpose CPU or a TEE with some security. The operation process may also be compromised by physical attacks, such as side channel attacks. Side channel attacks use the data to release side channel signals related to the data during the calculation process, such as power consumption, electromagnetic radiation, etc. Attackers can use this information to obtain core data.

[0064] To this end, an embodiment of the present application provides a computing device, which may be an application node, or may be arranged on an application node. Figure 2 As shown, the computing device 200 provided in the embodiment of the present application includes a general computing unit 210, a security unit 220 and a storage unit 230. In a specific implementation, the communication between the general computing unit 210 and the security unit 230 can be encrypted communication, and the security unit 220 and the storage unit 230 can use encrypted storage technology.

[0065] In the embodiment of the present application, Figure 2The general computing unit 210, security unit 220 and storage unit 230 shown may be discrete devices or at least partially integrated together. In a specific implementation, the general computing unit 210 may be integrated with the storage unit 230 as a system on chip, and the security unit 220 is a chip independent of the system on chip. In other specific implementations, the general computing unit 210, the storage unit 230 and the security unit 220 are integrated as a system on chip, in which case the security unit 220 is also referred to as an embedded security unit or an integrated security unit. In still other specific implementations, the general computing unit 210 and the security unit 220 are integrated as a system on chip, and the storage unit 230 is an independent memory. The storage unit 230 may be a non-volatile memory.

[0066] In the embodiment of the present application, a part of the neural network model is deployed in the general computing unit 210, and another part of the neural network model is deployed in the security unit 220. Thus, the part deployed in the security unit 220 can be protected from physical attacks, while avoiding the limitation of storage space and computing speed caused by completely deploying the neural network model in the security unit 220.

[0067] In some possible implementations, reference Figure 3A As shown, the storage unit 230 stores the first neural network code of the first model inference of the neural network model and part of the first network parameters of the first model inference; the security unit 220 stores the remaining first network parameters of the first model inference. In this embodiment, the first model inference is the complete inference of the neural network model, the first neural network code is the complete code of the neural network model, and the first network parameters are the complete network parameters of the neural network model. The model output can be obtained by performing the first model inference based on the model input. Some network parameters of the neural network model are stored in the security unit 220, which can prevent the partial network parameters from being subjected to physical attacks.

[0068] Continue to refer Figure 3A As shown, when the neural network model is called to process the model input to obtain the model output, the general computing unit 210 performs the first model reasoning, and in the process of performing the first model reasoning, sends a parameter read command to the security unit 220 to obtain the first network parameter that is not stored in the storage unit 230. The security unit receives the parameter read command and returns the first network parameter stored in itself and corresponding to the parameter read command to the general computing unit 210. The general computing unit 210 receives the first network parameter corresponding to the parameter read command returned by the security unit 220, and performs the first model reasoning based on the received first network parameter.

[0069] In a specific implementation, the first network parameters stored in the security unit 220 may include at least part of the network parameters of at least one network layer. The embodiment of the present application does not limit the specific network parameters stored in the security unit 220. In practical applications, the network parameters stored in the security unit 220 may be selected based on factors such as the network structure of the neural network model and the importance of the network parameters. For example, if a network layer has many network parameters, some of the network parameters of the network layer may be stored in the security unit 220; if a network layer has small or more critical network parameters, all of the network parameters of the network layer may be stored in the security unit 220.

[0070] In one example, reference Figure 4 In the exemplary network layer structure shown, the storage unit 230 stores the neural network codes of the regularization layer 1, the convolution layer 2, the pooling layer 3, the convolution layer 4, the pooling layer 5, the flattening layer 6, the fully connected layer 7 and the fully connected layer 8, as well as the network parameters of the other network layers except the convolution layer 4. The security unit 220 stores the network parameters of the convolution layer 4. Figure 4 , the general computing unit 210 executes the business logic. In this process, if it is determined to call the neural network model, the neural network model is called based on the model input, and the regularization layer 1, the convolution layer 2, and the pooling layer 3 are executed based on the network parameters stored in the storage unit 230. When executing the convolution layer 4, a parameter read command is sent to the security unit 220 to obtain the network parameters of the convolution layer 4. The security unit receives the parameter read command and returns the network parameters of the convolution layer 4 to the general computing unit 210. The general computing unit 210 receives the network parameters of the convolution layer 4 returned by the security unit 220, executes the convolution layer 4, and executes the pooling layer 5, the flattening layer 6, the fully connected layer 7, and the fully connected layer 8 based on the network parameters stored in the storage unit 230 to obtain the model output.

[0071] In another example, refer to Figure 4 In the exemplary network layer structure shown, the storage unit 230 stores the neural network codes of the regularization layer 1, the convolution layer 2, the pooling layer 3, the convolution layer 4, the pooling layer 5, the flattening layer 6, the fully connected layer 7 and the fully connected layer 8, as well as the network parameters of the other network layers except the convolution layer 4 and the fully connected layer 7. The security unit 220 stores the network parameters of the convolution layer 4 and the fully connected layer 7. Figure 4, the general computing unit 210 executes the business logic. During this process, if it is determined to call the neural network model, the neural network model is called based on the model input, and the regularization layer 1, the convolution layer 2, and the pooling layer 3 are executed based on the network parameters and the neural network code stored in the storage unit 230. When executing the convolution layer 4, a first parameter read command is sent to the security unit 220, and the first parameter read command indicates to obtain the network parameters of the convolution layer 4. The security unit receives the first parameter read command and returns the network parameters of the convolution layer 4 to the general computing unit 210. The general computing unit 210 receives the network parameters of the convolution layer 4 returned by the security unit 220, executes the convolution layer 4, and then executes the pooling layer 5 and the flattening layer 6 based on the network parameters and the neural network code stored in the storage unit 230. Subsequently, when executing the fully connected layer 7, a second parameter read command is sent to the security unit 220, and the second parameter read command indicates to obtain the network parameters of the fully connected layer 7. The security unit 220 receives the second parameter read command and returns the network parameters of the fully connected layer 7 to the general computing unit 210. The general computing unit 210 receives the network parameters of the fully connected layer 7 returned by the security unit 220, executes the fully connected layer 7, and then executes the fully connected layer 8 to obtain the model output.

[0072] In other examples, part of the network parameters of a network layer are stored in the storage unit 230, and the remaining network parameters are stored in the security unit 220. In this case, when executing the network layer, the general computing unit 210 obtains the complete network parameters of the network layer from the storage unit 230 and the security unit 220. The process of the general computing unit 210 obtaining the network parameters from the security unit 220 is similar to the above example and will not be described in detail here.

[0073] In some possible implementations, reference Figure 3B As shown, the storage unit 230 stores the first neural network code and the first network parameters of the first model inference of the neural network model; the security unit 220 stores the second neural network code and the second network parameters of the second model inference of the neural network model, and performs the second model inference based on the second network parameters and the second neural network code. The storage unit 230 stores the complete neural network code and network parameters of the first model inference, and the security unit 220 stores the complete neural network code and network parameters of the second model inference. In this embodiment, the first model inference and the second model inference are performed based on the model input to generate the model output. Through this embodiment, part of the model inference of the neural network model is performed in the security unit 220, which can prevent this part of the calculation process from being physically attacked.

[0074] Continue to refer Figure 3BAs shown, when calling the neural network model to process the model input to obtain the model output, the general computing unit 210 performs the first model reasoning, and in the process of performing the first model reasoning, sends an execution command to the security unit 220 to perform the second model reasoning in the security unit 220. The security unit 220 receives the execution command, performs the corresponding second model reasoning, and returns the execution result generated by performing the second model reasoning to the general computing unit 210. The general computing unit 210 receives the execution result returned by the security unit 220. The execution result may include an intermediate result or a model output. If the execution result is an intermediate result, the general computing unit 210 continues to perform the first model reasoning based on the intermediate result to obtain the model output. If the execution result is a model output, the general computing unit 210 executes the business logic based on the model output.

[0075] In a specific implementation, the second neural network code stored and executed by the security unit 220 may include a neural network code of at least one network layer. The embodiment of the present application does not limit the specific neural network code stored and executed by the security unit 220. In practical applications, the neural network code stored in the security unit 220 may be selected based on factors such as the network structure of the neural network model and the importance of the network layer.

[0076] In one example, reference Figure 4 In the exemplary network layer structure shown, the storage unit 230 stores the neural network codes and network parameters of the regularization layer 1, the convolution layer 2, the pooling layer 3, the convolution layer 4, the pooling layer 5, the flattening layer 6 and the fully connected layer 7. The security unit 220 stores the neural network codes and network parameters of the fully connected layer 8. Figure 4 , the general computing unit 210 executes the business logic. During this process, if it is determined to call the neural network model, the neural network model is called based on the model input, and based on the network parameters and neural network code stored in the storage unit 230, the regularization layer 1, convolution layer 2, pooling layer 3, convolution layer 4, pooling layer 5, flattening layer 6 and fully connected layer 7 are executed, and then an execution command is sent to the security unit 220 to execute the fully connected layer 8. The execution command may carry the intermediate result obtained by executing the fully connected layer 7. The security unit 220 receives the execution command, calculates the fully connected layer 8 based on the neural network code and network parameters of the fully connected layer 8 stored in itself, and returns the execution result to the general computing unit 210. The general computing unit 210 receives the execution result returned by the security unit 220. In Figure 4In the example, the fully connected layer 8 is the last layer of the neural network model, the execution result obtained by the security unit 220 executing the fully connected layer 8 is the model output, and the general computing unit 210 can execute the business logic based on the model output. In this example, the first model reasoning includes executing the regularization layer 1, the convolution layer 2, the pooling layer 3, the convolution layer 4, the pooling layer 5, the flattening layer 6 and the fully connected layer 7, and the second model reasoning includes executing the fully connected layer 8.

[0077] In another example, refer to Figure 4 In the exemplary network layer structure shown, the storage unit 230 stores the neural network codes and network parameters of the regularization layer 1, the convolution layer 2, the pooling layer 3, the pooling layer 5, the flattening layer 6 and the fully connected layer 8. The security unit 220 stores the neural network codes and network parameters of the convolution layer 4 and the fully connected layer 7. Figure 4 , the general computing unit 210 executes the business logic. During this process, if it is determined to call the neural network model, the neural network model is called based on the model input, and the regularization layer 1, the convolution layer 2, and the pooling layer 3 are executed based on the network parameters and the neural network code stored in the storage unit 230. Subsequently, the general computing unit 210 sends a first execution command to the security unit 220. The first execution command carries the execution result of the pooling layer 3, and the first execution command instructs the security unit to execute the convolution layer 4. The security unit 220 receives the first execution command sent by the general computing unit, calculates the convolution layer 4 based on the neural network code and network parameters of the convolution layer 4 stored in itself, and returns the execution result to the general computing unit 210. The general computing unit 210 receives the execution result of the convolution layer 4 returned by the security unit 220. Subsequently, the general computing unit 210 executes the pooling layer 5 and the flattening layer 6 based on the network parameters and the neural network code stored in the storage unit 230. Further, the general computing unit 210 sends a second execution command to the security unit 220, the second execution command carries the execution result of the flattening layer 6, and the second execution command instructs the security unit to execute the fully connected layer 7. The security unit 220 receives the second execution command sent by the general computing unit, calculates the fully connected layer 7 based on the neural network code and network parameters of the fully connected layer 7 stored in itself, and returns the execution result to the general computing unit 210. The general computing unit 210 receives the execution result of the fully connected layer 7 returned by the security unit 220. Subsequently, the general computing unit 210 executes the fully connected layer 8 based on the network parameters and neural network code stored in the storage unit 230 to obtain the model output. In this example, the first model reasoning includes executing the regularization layer 1, the convolution layer 2, the pooling layer 3, the pooling layer 5, the flattening layer 6 and the fully connected layer 8, and the second model reasoning includes executing the convolution layer 4 and the fully connected layer 7.

[0078] In some possible implementations, reference Figure 3CAs shown, the storage unit 230 stores the first neural network code of the first model inference of the neural network model and part of the first network parameters of the first model inference; the security unit 220 stores the remaining first network parameters of the first model inference, as well as the second neural network code and second network parameters of the second model inference of the neural network model. The security unit 220 performs the second model inference based on the second network parameters and the second neural network code. The general computing unit 210 performs the first model inference based on the first neural network code and the first network parameters. Through this implementation, it is possible to prevent some network parameters and some computing processes from being physically attacked.

[0079] Continue to refer Figure 3C As shown, when the neural network model is called to process the model input to obtain the model output, the general computing unit 210 performs the first model reasoning, and in the process of performing the first model reasoning, sends a parameter reading command to the security unit 220 to obtain the first network parameter not stored in the storage unit 230. The security unit receives the parameter reading command and returns the first network parameter stored by itself and corresponding to the parameter reading command to the general computing unit 210. The general computing unit 210 receives the first network parameter corresponding to the parameter reading command returned by the security unit 220, and performs the first model reasoning based on the received first network parameter. In the process of performing the first model reasoning, the general computing unit 210 also sends an execution command to the security unit 220 to perform the second model reasoning in the security unit 220. The security unit 220 receives the execution command, performs the corresponding second model reasoning, and returns the execution result generated by performing the second model reasoning to the general computing unit 210. The general computing unit 210 receives the execution result returned by the security unit 220. The execution result may include an intermediate result or a model output. If the execution result is an intermediate result, the general computing unit 210 continues to perform the first model reasoning based on the intermediate result to obtain the model output. If the execution result is a model output, the general computing unit 210 executes the business logic based on the model output.

[0080] In one example, reference Figure 4 In the exemplary network layer structure shown, the storage unit 230 stores the neural network codes of the regularization layer 1, the convolution layer 2, the pooling layer 3, the convolution layer 4, the pooling layer 5, the flattening layer 6 and the fully connected layer 8, as well as the network parameters of the other network layers except the convolution layer 4. The security unit 220 stores the network parameters of the convolution layer 4, as well as the neural network code and network parameters of the fully connected layer 7. Figure 4, the general computing unit 210 executes the business logic. During this process, if it is determined to call the neural network model, the neural network model is called based on the model input, and the regularization layer 1, the convolution layer 2, and the pooling layer 3 are executed based on the network parameters and the neural network code stored in the storage unit 230. Subsequently, the general computing unit 210 sends a parameter read command to the security unit 220 to obtain the network parameters of the convolution layer 4. The security unit 220 receives the parameter read command sent by the general computing unit 210 and returns the network parameters of the convolution layer 4 stored in itself to the general computing unit 210. The general computing unit 210 receives the network parameters of the convolution layer 4 returned by the security unit 220 and executes the convolution layer 4. Subsequently, the general computing unit 210 executes the pooling layer 5 and the flattening layer 6 based on the network parameters and the neural network code stored in the storage unit 230. Further, the general computing unit 210 sends an execution command to the security unit 220, the execution command carries the execution result of the flattening layer 6, and the execution command instructs the security unit 220 to execute the fully connected layer 7. The security unit 220 receives the execution command sent by the general computing unit, performs calculations of the fully connected layer 7 based on the neural network code and network parameters of the fully connected layer 7 stored in the security unit 220, and returns the execution result to the general computing unit 210. The general computing unit 210 receives the execution result of the fully connected layer 7 returned by the security unit 220. Subsequently, the general computing unit 210 executes the fully connected layer 8 based on the network parameters and neural network code stored in the storage unit 230 to obtain the model output.

[0081] In some possible implementations, the general computing unit 210 is also used to: receive a neural network model and a deployment instruction sent by a server, the deployment instruction is used to indicate the neural network code and network parameters deployed in the general computing unit, and the neural network code and / or network parameters deployed in the security unit; and deploy the neural network model in the general computing unit and the security unit according to the deployment instruction. In a specific implementation, the network parameters deployed in the security unit 220 include at least part of the network parameters of at least one neural network layer executed by the general computing unit 210. The neural network code deployed in the security unit 220 includes the neural network code of at least one neural network layer.

[0082] In some possible implementations, the general computing unit 210 and the storage unit 230 are located in a trusted execution environment (TEE), and the trusted execution environment executes the first model reasoning. In this implementation, the trusted execution environment can communicate with the security unit 220, and the trusted execution environment can obtain network parameters from the security unit 220, instruct the security unit 220 to execute the second model reasoning, and receive the execution result of the second model reasoning. In order to ensure the security of the terminal device, a terminal device security framework represented by ARM (advanced RISC machines) TrustZone has emerged (where the full name of RISC in English is reduced instruction set computer). Under the ARM TrustZoneR framework, system-level security is obtained by dividing the software and hardware resources of the system on chips (SoC) into two worlds. These two worlds are the normal world and the secure world (secure world) (also called security domain and non-security domain), and these two worlds correspond to the rich execution environment (REE) and the trusted execution environment (TEE), respectively. REE and TEE run on the same physical device, each running a set of operating systems. REE runs client applications (CA) with low security requirements; TEE runs trusted applications (TA) whose security needs to be guaranteed, providing a secure execution environment for authorized trusted applications TA. CA and TA communicate through the communication mechanism provided by ARM TrustZone, just like a client and server.

[0083] In some possible implementations, the general computing unit 210 is deployed with a virtual machine, and the virtual machine performs the first model reasoning. In this implementation, the virtual machine can communicate with the security unit 220, and the virtual machine can obtain network parameters from the security unit 220, instruct the security unit 220 to perform the second model reasoning, and receive the execution result of the second model reasoning.

[0084] The present application also provides a security unit that can implement Figure 2 The functions of the security unit 220 are shown. In some embodiments, Figure 5 As shown, the security unit 500 may include a security storage module 510 and a security computing module 520. The security storage module 510 may have permissions that can be accessed only by the security computing module 520.

[0085] In some embodiments, the secure storage module 510 may store part of the first network parameters of the first model inference of the neural network model, and the secure computing module 520 may provide the first network parameters stored by itself to the general computing unit. In this embodiment, the first model inference of the neural network model is its complete model inference, and the first network parameters are the complete parameters of the neural network model. In a specific implementation, the secure computing module 520 may receive a parameter read command from the general computing unit, and return the first network parameters stored in the secure storage module 510 and corresponding to the parameter read command to the general computing unit. In some possible embodiments, in the security unit 500, the first network parameters stored in the secure storage module 510 include at least part of the network parameters of at least one neural network layer of the first model inference.

[0086] In one example, reference Figure 4 The exemplary network layer structure shown in FIG. 5 shows, the security unit 500 secure storage module 510 stores the network parameters of the convolution layer 4. Figure 4 When executing convolution layer 4, the general computing unit sends a parameter read command to the security unit 500 to obtain the network parameters of convolution layer 4. The security computing module 520 of the security unit 500 receives the parameter read command, reads the network parameters of convolution layer 4 stored in the security storage module 510, and returns the network parameters of convolution layer 4 to the general computing unit. The general computing unit receives the network parameters of convolution layer 4 returned by the security computing module 520 and executes convolution layer 4. This example mainly illustrates the process of the security unit 500 providing network parameters. The execution of other network layers by the general computing unit can refer to the previous examples of this specification, which will not be repeated here.

[0087] In another example, refer to Figure 4 The exemplary network layer structure shown in FIG. 5 shows, the secure storage module 510 stores the network parameters of the convolutional layer 4 and the fully connected layer 7. Figure 4, when executing convolution layer 4, the general computing unit sends a first parameter read command to the security unit 500, and the first parameter read command indicates to obtain the network parameters of convolution layer 4. The security computing module 520 of the security unit 500 receives the first parameter read command, reads the network parameters of convolution layer 4 stored in the security storage module 510, and returns the network parameters of convolution layer 4 to the general computing unit. The general computing unit receives the network parameters of convolution layer 4 returned by the security computing module 520 and executes convolution layer 4. When executing fully connected layer 7, the general computing unit sends a second parameter read command to the security unit 500, and the second parameter read command indicates to obtain the network parameters of fully connected layer 7. The security computing module 520 of the security unit 500 receives the second parameter read command, reads the network parameters of fully connected layer 7 stored in the security storage module 510, and returns the network parameters of fully connected layer 7 to the general computing unit. The general computing unit receives the network parameters of fully connected layer 7 returned by the security unit 500 and executes fully connected layer 7. The general computing unit can refer to the examples mentioned above in this specification for executing other network layers, which will not be repeated here. This example mainly illustrates the process of the security unit 500 providing network parameters. For the general computing unit executing other network layers, please refer to the previous examples in this specification, which will not be described in detail here.

[0088] In some embodiments, the secure storage module 510 may store the second neural network code and the second network parameters of the second model inference of the neural network model, and the secure computing module 520 may perform the second model inference based on the second network parameters and the second neural network code. In this embodiment, the second model inference is a partial model inference of the neural network model. In a specific implementation, the secure computing module 520 may receive an execution command from a general computing unit, perform corresponding second model inference based on the second neural network code and the second network parameters corresponding to the execution command stored in the secure storage module 510, and return the execution result to the general computing unit, the execution result including the model output or the intermediate result. In some possible embodiments, in the security unit 500, the second neural network code stored in the secure storage module 510 includes the neural network code of at least one neural network layer.

[0089] In one example, reference Figure 4 The exemplary network layer structure shown in FIG. 5 shows, the security unit 500 stores the neural network code and network parameters of the fully connected layer 8. Figure 4, the general computing unit executes regularization layer 1, convolution layer 2, pooling layer 3, convolution layer 4, pooling layer 5, flattening layer 6 and fully connected layer 7, and then sends an execution command to the security unit 500 to execute the fully connected layer 8. The execution command can carry the intermediate result obtained by executing the fully connected layer 7. The security computing module 520 of the security unit 500 receives the execution command, calculates the fully connected layer 8 based on the neural network code and network parameters of the fully connected layer 8 stored in the security storage module 510, and returns the execution result to the general computing unit. The general computing unit receives the execution result returned by the security computing module 520. Figure 4 In the example, the fully connected layer 8 is the last layer of the neural network model, and the execution result obtained by the security computing module 520 executing the fully connected layer 8 is the model output, and the general computing unit can execute the business logic based on the model output. In this example, the first model reasoning includes executing the regularization layer 1, the convolution layer 2, the pooling layer 3, the convolution layer 4, the pooling layer 5, the flattening layer 6 and the fully connected layer 7, and the second model reasoning includes executing the fully connected layer 8.

[0090] In another example, refer to Figure 4 The exemplary network layer structure shown in FIG. 5 shows, the secure storage module 510 of the security unit 500 stores the neural network code and network parameters of the convolutional layer 4 and the fully connected layer 7. Figure 4, the general computing unit executes the business logic. During this process, if it is determined to call the neural network model, the neural network model is called based on the model input to execute the regularization layer 1, the convolution layer 2, and the pooling layer 3. Subsequently, the general computing unit sends a first execution command to the security unit 500. The first execution command carries the execution result of the pooling layer 3, and the first execution command instructs the security unit to execute the convolution layer 4. The security computing module 520 of the security unit 500 receives the first execution command sent by the general computing unit, calculates the convolution layer 4 based on the neural network code and network parameters of the convolution layer 4 stored in the security storage module 510, and returns the execution result to the general computing unit. The general computing unit receives the execution result of the convolution layer 4 returned by the security computing module 520. Subsequently, the general computing unit 210 executes the pooling layer 5 and the flattening layer 6, and sends a second execution command to the security unit 220. The second execution command carries the execution result of the flattening layer 6, and the second execution command instructs the security unit to execute the fully connected layer 7. The security computing module 520 of the security unit 500 receives the second execution command sent by the general computing unit, performs calculations of the fully connected layer 7 based on the neural network code and network parameters of the fully connected layer 7 stored in the security storage module 510, and returns the execution result to the general computing unit. The general computing unit receives the execution result of the fully connected layer 7 returned by the security computing module 520. In this example, the first model reasoning includes executing the regularization layer 1, the convolution layer 2, the pooling layer 3, the pooling layer 5, the flattening layer 6 and the fully connected layer 8, and the second model reasoning includes executing the convolution layer 4 and the fully connected layer 7.

[0091] In some embodiments, the secure storage module 510 may store part of the first network parameters of the first model inference of the neural network model, and store the second neural network code and the second network parameters of the second model inference of the neural network model. In this embodiment, the first model inference is a partial model inference of the neural network model, and the second model inference is another partial model inference of the neural network model. The secure computing module 520 may provide the first network parameters stored in the secure storage module 510 to the general computing unit, and perform the second model inference based on the second network parameters and the second neural network code stored in the secure storage module 510, and return the execution result to the general computing unit, and the execution result includes the model output or the intermediate result.

[0092] In a specific implementation, the secure computing module 520 may receive a parameter read command from the general computing unit, and return to the general computing unit the first network parameter stored in the secure storage module 510 and corresponding to the parameter read command. The secure computing module 520 may also receive an execution command from the general computing unit, execute corresponding second model reasoning based on the second neural network code and the second network parameter corresponding to the execution command, and return the execution result to the general computing unit, the execution result including the model output or the intermediate result.

[0093] In one example, reference Figure 4 In the exemplary network layer structure shown, the secure storage module 510 of the security unit 500 stores the network parameters of the convolutional layer 4, and the neural network code and network parameters of the fully connected layer 7. Figure 4 , the general computing unit executes the business logic. During this process, if it is determined to call the neural network model, the neural network model is called based on the model input, and the regularization layer 1, the convolution layer 2, and the pooling layer 3 are executed. Then the general computing unit sends a parameter read command to the security unit 500 to obtain the network parameters of the convolution layer 4. The security computing module 520 of the security unit 500 receives the parameter read command sent by the general computing unit 210, and returns the network parameters of the convolution layer 4 stored in the security storage module 510 to the general computing unit. The general computing unit receives the network parameters of the convolution layer 4 returned by the security computing module 520 and executes the convolution layer 4. Then, the general computing unit executes the pooling layer 5 and the flattening layer 6, and sends an execution command to the security unit 220. The execution command carries the execution result of the flattening layer 6, and the execution command instructs the security unit 500 to execute the fully connected layer 7. The security computing module 520 receives the execution command sent by the general computing unit, performs the calculation of the fully connected layer 7 based on the neural network code and network parameters of the fully connected layer 7 stored in the security storage module 510, and returns the execution result to the general computing unit. The general computing unit receives the execution result of the fully connected layer 7 returned by the security unit 500. Subsequently, the general computing unit executes the fully connected layer 8 to obtain the model output.

[0094] The present application embodiment provides a method for processing a neural network model, which can be implemented by the security unit 220 and the security unit 500. Fig. 6A , Figure 6B and Figure 6C An implementation of this method will be described.

[0095] In some embodiments, reference Fig. 6A As shown, the method includes the following steps.

[0096] Step S601A, the security unit stores part of the first network parameters inferred by the first model of the neural network model.

[0097] In this embodiment, in the computing device 200, the storage unit 230 stores the first neural network code of the first model inference of the neural network model and part of the first network parameters of the first model inference. The security unit 220 can store part of the first network parameters of the first model inference of the neural network model. In this implementation, the first model inference is the complete inference of the neural network model, the first neural network code is the complete code of the neural network model, the first network parameters are the complete network parameters of the neural network model, part of the first network parameters are stored in the storage unit 230, and the other part is stored in the security unit 220. The model output can be obtained by performing the first model inference based on the model input. Some network parameters of the neural network model are stored in the security unit 220, which can prevent the part of the network parameters from being subjected to physical attacks.

[0098] In a specific implementation, the security unit 220 can receive part of the first network parameters inferred by the first model of the neural network model sent by the general computing unit 210, and store part of the first network parameters inferred by the first model of the neural network model in the security unit 220.

[0099] In a specific implementation, in the computing device 200, the first network parameters stored in the security unit 220 may include at least part of the network parameters of at least one network layer. The embodiment of the present application does not limit the specific network parameters stored in the security unit 220. In practical applications, the network parameters stored in the security unit 220 may be selected based on factors such as the network structure of the neural network model and the importance of the network parameters. For example, if a network layer has many network parameters, some of the network parameters of the network layer may be stored in the security unit 220; if a network layer has small or more critical network parameters, all of the network parameters of the network layer may be stored in the security unit 220.

[0100] Step S602A: The security unit provides the stored first network parameter to the general computing unit.

[0101] In this embodiment, in the computing device 200, when the general computing unit 210 calls the neural network model to process the model input to obtain the model output, the general computing unit 210 performs the first model reasoning, and in the process of performing the first model reasoning, sends a parameter reading command to the security unit 220 to obtain the first network parameter that is not stored in the storage unit 230. The security unit 220 receives the parameter reading command and returns the first network parameter stored in itself and corresponding to the parameter reading command to the general computing unit 210. The general computing unit 210 receives the first network parameter corresponding to the parameter reading command returned by the security unit 220, and performs the first model reasoning based on the received first network parameter.

[0102] In some embodiments, reference Figure 6BAs shown, the method includes the following steps. Through this embodiment, part of the model reasoning of the neural network model is performed in the security unit 220, which can prevent the part of the calculation process from being physically attacked.

[0103] Step S601B, the security unit stores the second neural network code and second network parameters inferred by the second model of the neural network model.

[0104] In this embodiment, in the computing device 200, the storage unit 230 stores the first neural network code and the first network parameters of the first model inference of the neural network model, and the security unit 220 stores the second neural network code and the second network parameters of the second model inference of the neural network model. The storage unit 230 stores the complete neural network code and network parameters of the first model inference, and the security unit 220 stores the complete neural network code and network parameters of the second model inference.

[0105] In a specific implementation, the security unit 220 may receive the second neural network code and the second network parameters inferred by the second model of the neural network model sent by the general computing unit 210, and store the second neural network code and the second network parameters inferred by the second model of the neural network model in the security unit 220. The second neural network code is configured to be executable by the security unit 220, such as an application running in the security unit 220.

[0106] Step S602B: The security unit executes second model reasoning based on the second network parameters and the second neural network code to obtain an execution result, which includes a model output or an intermediate result.

[0107] Step S603B: the security unit returns the execution result to the general computing unit.

[0108] In this embodiment, the first model reasoning and the second model reasoning are performed based on the model input to generate the model output. Through this embodiment, part of the model reasoning of the neural network model is performed in the security unit 220, which can prevent the part of the computing process from being physically attacked.

[0109] In this embodiment, when the general computing unit 210 calls the neural network model to process the model input to obtain the model output, it performs the first model reasoning, and in the process of performing the first model reasoning, sends an execution command to the security unit 220 to perform the second model reasoning in the security unit 220. The security unit 220 receives the execution command, performs the corresponding second model reasoning based on the second neural network code and the second network parameters stored in itself, and returns the execution result generated by performing the second model reasoning to the general computing unit 210. The general computing unit 210 receives the execution result returned by the security unit 220. The execution result may include an intermediate result or a model output. If the execution result is an intermediate result, the general computing unit 210 continues to perform the first model reasoning based on the intermediate result to obtain the model output. If the execution result is a model output, the general computing unit 210 executes the business logic based on the model output.

[0110] In a specific implementation, the second neural network code stored in the security unit 220 may include a neural network code of at least one network layer. The embodiment of the present application does not limit the specific neural network code stored in the security unit 220. In practical applications, the neural network code stored in the security unit 220 may be selected based on factors such as the network structure of the neural network model and the importance of the network layer.

[0111] In some embodiments, reference Figure 6C As shown, the method includes the following steps.

[0112] Step S601C, the security unit stores part of the first network parameters inferred by the first model of the neural network model, the second neural network code and the second network parameters inferred by the second model of the neural network model.

[0113] In this embodiment, in the computing device 200, the storage unit 230 stores the first neural network code and part of the first network parameters inferred by the first model of the neural network model; the security unit 220 stores the remaining first network parameters inferred by the first model, and the second neural network code and second network parameters inferred by the second model of the neural network model. Through this implementation, it is possible to prevent some network parameters and some computing processes from being physically attacked.

[0114] In a specific implementation, the security unit 220 may receive the remaining first network parameters of the first model inference of the neural network model, and the second neural network code and the second network parameters of the second model inference of the neural network model sent by the general computing unit 210, and store the remaining first network parameters of the first model inference, and the second neural network code and the second network parameters of the second model inference of the neural network model in the security unit 220. The second neural network code is configured to be executable by the security unit 220, such as an application running in the security unit 220.

[0115] In a specific implementation, the second neural network code stored in the security unit 220 may include a neural network code of at least one network layer. The embodiment of the present application does not limit the specific neural network code stored and executed by the security unit 220. In practical applications, the neural network code stored in the security unit 220 may be selected based on factors such as the network structure of the neural network model and the importance of the network layer.

[0116] In a specific implementation, the first network parameters stored in the security unit 220 may include at least part of the network parameters of at least one network layer. The embodiment of the present application does not limit the specific network parameters stored in the security unit 220. In practical applications, the network parameters stored in the security unit 220 may be selected based on factors such as the network structure of the neural network model and the importance of the network parameters. For example, if a network layer has many network parameters, some of the network parameters of the network layer may be stored in the security unit 220; if a network layer has small or more critical network parameters, all of the network parameters of the network layer may be stored in the security unit 220.

[0117] Step S602C: the security unit provides the stored first network parameter to the general computing unit.

[0118] In this embodiment, when the general computing unit 210 calls the neural network model to process the model input to obtain the model output, it performs the first model reasoning, and in the process of performing the first model reasoning, sends a parameter reading command to the security unit 220 to obtain the first network parameter that is not stored in the storage unit 230. The security unit receives the parameter reading command and returns the first network parameter stored in itself and corresponding to the parameter reading command to the general computing unit 210. The general computing unit 210 receives the first network parameter corresponding to the parameter reading command returned by the security unit 220, and performs the first model reasoning based on the received first network parameter.

[0119] Step S603C: The security unit executes second model reasoning based on the second network parameters and the second neural network code to obtain an execution result, which includes a model output or an intermediate result.

[0120] Step S604C: the security unit returns the execution result to the general computing unit.

[0121] In this embodiment, during the process of executing the first model reasoning, the general computing unit 210 also sends an execution command to the security unit 220 to execute the second model reasoning in the security unit 220. The security unit 220 receives the execution command, executes the corresponding second model reasoning, and returns the execution result generated by executing the second model reasoning to the general computing unit 210. The general computing unit 210 receives the execution result returned by the security unit 220. The execution result may include an intermediate result or a model output. If the execution result is an intermediate result, the general computing unit 210 continues to perform the first model reasoning based on the intermediate result to obtain the model output. If the execution result is a model output, the general computing unit 210 executes the business logic based on the model output.

[0122] It should be understood that the execution order of step S602C and step S603C is not limited in this embodiment. Based on the deployment of network parameters and neural network code, step S602C can be executed after step S603C, or S603C can be executed after step S602C.

[0123] The present application embodiment provides a method for processing a neural network model, which can be implemented by a server and an application node, for example Figure 1 The server 10 and application node 20 are shown. The server is responsible for distributing the trained neural network model to the application node. The application node includes a computing device, which can be Figure 2 The computing device 200 shown. The computing device includes a security unit, which may be Figure 5 The security unit 500 shown in FIG. 5 is a schematic diagram of a security unit 500. In a specific implementation, the communication between the server and the computing device is encrypted communication, and the communication channel may be WIFI, mobile communication, etc. Figure 7 As shown, the method includes the following steps.

[0124] Step S701, the server sends a neural network model and deployment instructions to the application node, where the deployment instructions are used to indicate the neural network code and network parameters deployed in the general computing unit, as well as the neural network code and / or network parameters deployed in the security unit.

[0125] In the embodiment of the present application, in the application node including the computing device 200, the server can instruct to deploy a part of the neural network model in the general computing unit 210, and deploy another part of the neural network model in the security unit 220. In this way, the part deployed in the security unit 220 is protected from physical attacks, while avoiding the limitation of storage space and computing speed caused by completely deploying the neural network model in the security unit 220.

[0126] In some possible implementations, the server instructs that the first neural network code of the first model inference of the neural network model and part of the first network parameters of the first model inference be deployed in the general computing unit 210, and the remaining first network parameters of the first model inference be deployed in the security unit 220. In this implementation, the first model inference is the complete inference of the neural network model, the first neural network code is the complete code of the neural network model, and the first network parameters are the complete network parameters of the neural network model. The model output can be obtained by performing the first model inference based on the model input. Some network parameters of the neural network model are stored in the security unit 220, which can prevent the partial network parameters from being subjected to physical attacks.

[0127] In some possible implementations, the server may instruct to deploy the first neural network code and the first network parameters of the first model inference of the neural network model in the general computing unit 210, and deploy the second neural network code and the second network parameters of the second model inference of the neural network model in the security unit 220. The complete neural network code and network parameters of the first model inference are deployed in the general computing unit 210, and the complete neural network code and network parameters of the second model inference are deployed in the security unit 220. Through this implementation, part of the model inference of the neural network model is executed in the security unit 220, which can prevent the part of the computing process from being physically attacked.

[0128] In some possible implementations, the server may instruct to deploy the first neural network code and part of the first network parameters of the first model inference of the neural network model in the general computing unit 210, and deploy the remaining first network parameters of the first model inference, as well as the second neural network code and second network parameters of the second model inference of the neural network model in the security unit 220. Through this implementation, it is possible to prevent some network parameters and some computing processes from being physically attacked.

[0129] In a specific implementation, the server may configure the first neural network code as an application that can be executed by the general computing unit 210 , and configure the second neural network code as an application that can be executed by the security unit 220 .

[0130] Step S702: The general computing unit of the application node receives the neural network model and deployment instructions sent by the server.

[0131] In this embodiment, in the system 100, the application node 20 and the server 10 can communicate through a trusted channel. In a specific implementation, the server can encrypt the neural network model and the deployment instruction, and the general computing unit 210 can decrypt and verify the data sent by the server. Encryption, decryption and verification are existing technologies, and this embodiment will not be described in detail.

[0132] Step S703: The general computing unit deploys the neural network model in the general computing unit and the security unit according to the deployment instruction.

[0133] In this embodiment, if the server instructs that the first neural network code of the first model inference of the neural network model and part of the first network parameters of the first model inference be deployed in the general computing unit 210, and the remaining first network parameters of the first model inference be deployed in the security unit 220, the general computing unit 210 stores the first neural network code of the first model inference of the neural network model and part of the first network parameters of the first model inference in the storage unit 230, and instructs the security unit 220 to store the remaining first network parameters of the first model inference.

[0134] If the server can instruct the deployment of the first neural network code and the first network parameters of the first model inference of the neural network model in the general computing unit 210, and the deployment of the second neural network code and the second network parameters of the second model inference of the neural network model in the security unit 220, then the general computing unit 210 stores the first neural network code and the first network parameters of the first model inference of the neural network model in the storage unit 230, and instructs the security unit 220 to store the second neural network code and the second network parameters of the second model inference.

[0135] If the server instructs the general computing unit 210 to deploy the first neural network code and part of the first network parameters of the first model inference of the neural network model, and the security unit 220 to deploy the remaining first network parameters of the first model inference, as well as the second neural network code and second network parameters of the second model inference of the neural network model, the general computing unit 210 stores the first neural network code and part of the first model inference of the neural network model in the storage unit 230, and instructs the security unit 220 to store the remaining first network parameters of the first model inference, as well as the second neural network code and second network parameters of the second model inference.

[0136] After the neural network model is deployed in the general computing unit 210 and the security unit 220 according to the deployment instructions, the general computing unit 210 can execute the business logic and call the neural network model when it is determined to call the neural network model in the execution of the business logic.

[0137] In some possible implementations, the server may also adjust the deployment location of the neural network code and / or network parameters, including the general computing unit and the security unit; and send a deployment instruction based on the adjustment to the application node. In a specific implementation, the deployment location of the neural network code and / or network parameters may be adjusted when the neural network model is updated or periodically. The security enhancement mechanism of this implementation can increase the difficulty of attackers to attack the system and improve the security of the neural network model.

[0138] In this embodiment, after the neural network model is deployed on the application node or computing device, the neural network model can be used to generate a model output based on the model input.

[0139] In some possible implementations, when the neural network model is called to process the model input to obtain the model output, the general computing unit 210 performs the first model reasoning, and in the process of performing the first model reasoning, sends a parameter read command to the security unit 220 to obtain the first network parameter that is not stored in the storage unit 230. The security unit receives the parameter read command and returns the first network parameter stored in itself and corresponding to the parameter read command to the general computing unit 210. The general computing unit 210 receives the first network parameter corresponding to the parameter read command returned by the security unit 220, and performs the first model reasoning based on the received first network parameter.

[0140] In some possible implementations, when calling a neural network model to process a model input to obtain a model output, the general computing unit 210 performs a first model reasoning, and in the process of performing the first model reasoning, sends an execution command to the security unit 220 to perform a second model reasoning in the security unit 220. The security unit 220 receives the execution command, performs the corresponding second model reasoning, and returns the execution result generated by performing the second model reasoning to the general computing unit 210. The general computing unit 210 receives the execution result returned by the security unit 220. The execution result may include an intermediate result or a model output. If the execution result is an intermediate result, the general computing unit 210 continues to perform the first model reasoning based on the intermediate result to obtain the model output. If the execution result is a model output, the general computing unit 210 executes the business logic based on the model output.

[0141] In some possible implementations, when the neural network model is called to process the model input to obtain the model output, the general computing unit 210 performs the first model reasoning, and in the process of performing the first model reasoning, sends a parameter reading command to the security unit 220 to obtain the first network parameter not stored in the storage unit 230. The security unit receives the parameter reading command and returns the first network parameter stored by itself and corresponding to the parameter reading command to the general computing unit 210. The general computing unit 210 receives the first network parameter corresponding to the parameter reading command returned by the security unit 220, and performs the first model reasoning based on the received first network parameter. In the process of performing the first model reasoning, the general computing unit 210 also sends an execution command to the security unit 220 to perform the second model reasoning in the security unit 220. The security unit 220 receives the execution command, performs the corresponding second model reasoning, and returns the execution result generated by performing the second model reasoning to the general computing unit 210. The general computing unit 210 receives the execution result returned by the security unit 220. The execution result may include an intermediate result or a model output. If the execution result is an intermediate result, the general computing unit 210 continues to perform the first model reasoning based on the intermediate result to obtain the model output. If the execution result is a model output, the general computing unit 210 executes the business logic based on the model output.

[0142] An embodiment of the present application also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above method.

[0143] The embodiment of the present application also provides an application node, including a device body and the above-mentioned computing device or the above-mentioned security unit arranged in the device body. Application nodes may include but are not limited to: Personal Computers (PCs), such as laptops, desktop computers, etc.; Mobile Devices (Mobile Devices), such as smart phones, tablet computers, etc.; Embedded Systems (Embedded Systems), such as smart cameras, drones, self-driving cars, etc.; Internet of Things (IoT) devices, such as various sensors and smart devices, such as smart home devices, etc.; Edge Computing Devices, such as edge servers, gateways, etc.; Wearable Devices, such as smart watches and health monitoring devices, etc.

[0144] An embodiment of the present application also provides a server, which is used to: send a neural network model and deployment instructions to an application node, the deployment instructions are used to indicate the neural network code and network parameters deployed in a general computing unit, and the neural network code and / or network parameters deployed in a security unit.

[0145] In some possible implementations, the server is further used to: adjust the deployment location of the neural network code and / or network parameters, the deployment location including the general computing unit and the security unit; and send a deployment instruction based on the adjustment to the application node.

[0146] In specific implementations, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms, etc.

[0147] The above are only preferred embodiments of the present application, and are not intended to limit the present application in any form. Although the present application has been disclosed as above with preferred embodiments, it is not intended to limit the present application. Any technical personnel in the field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present application. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.

Claims

1. A computing device, characterized in that: include: A storage unit, configured to store a first neural network code inferred by a first model of a neural network model and at least a portion of first network parameters inferred by the first model; A security unit, used to: store the remaining first network parameters inferred by the first model; and / or storing a second neural network code and a second network parameter for second model inference of the neural network model, and performing the second model inference based on the second network parameter and the second neural network code; A general computing unit is used to perform the first model inference based on the first neural network code and the first network parameters.

2. The computing device according to claim 1, wherein: The general computing unit is further used for: Receiving a neural network model and a deployment instruction sent by a server, wherein the deployment instruction is used to indicate a neural network code and network parameters deployed in a general computing unit, and a neural network code and / or network parameters deployed in a security unit; The neural network model is deployed in the general computing unit and the security unit according to the deployment instruction.

3. The computing device according to claim 1, wherein: The safety unit is further used for: receiving a parameter reading command from the general computing unit, and returning to the general computing unit a first network parameter stored in the general computing unit and corresponding to the parameter reading command; and / or Receive an execution command from the general computing unit, execute corresponding second model reasoning based on a second neural network code and a second network parameter corresponding to the execution command, and return an execution result to the general computing unit, wherein the execution result includes a model output or an intermediate result.

4. The computing device according to claim 1, wherein: The general computing unit is further used for: Sending a parameter read command to the security unit, and receiving a first network parameter corresponding to the parameter read command returned by the security unit; and / or An execution command is sent to the security unit, and an execution result returned after the security unit executes the second model reasoning corresponding to the execution command is received, wherein the execution result includes a model output or an intermediate result.

5. The application node according to claim 1, characterized in that: The first network parameters stored by the security unit include at least part of the network parameters of at least one neural network layer inferred by the first model; and / or The second neural network code includes neural network code of at least one neural network layer.

6. The computing device according to any one of claims 1 to 5, characterized in that: The general computing unit and the storage unit are located in a trusted execution environment, and the trusted execution environment executes the first model reasoning.

7. The computing device according to any one of claims 1 to 5, characterized in that: The general computing unit is deployed with a virtual machine, and the virtual machine executes the first model reasoning.

8. A safety unit, characterized in that: include: A secure storage module, used to store part of the first network parameters inferred by the first model of the neural network model; and / or storing a second neural network code and a second network parameter inferred by a second model of the neural network model; A secure computing module, configured to provide the first network parameter stored in the general computing unit; and / or performing the second model inference based on the second network parameters and the second neural network code.

9. The safety unit according to claim 8, characterized in that The secure computing module is further used for: receiving a parameter reading command from the general computing unit, and returning to the general computing unit a first network parameter stored in the general computing unit and corresponding to the parameter reading command; and / or Receive an execution command from the general computing unit, execute corresponding second model reasoning based on a second neural network code and a second network parameter corresponding to the execution command, and return an execution result to the general computing unit, wherein the execution result includes a model output or an intermediate result.

10. The security unit according to claim 8, characterized in that The first network parameters stored by the security unit include at least part of the network parameters of at least one neural network layer inferred by the first model; and / or The second neural network code includes neural network code of at least one neural network layer.

11. A method for processing a neural network model, characterized in that: Applied to a security unit, the method comprises: Storing part of the first network parameters inferred by the first model of the neural network model, and / or the second neural network code and second network parameters inferred by the second model of the neural network model; Providing the first stored network parameters to a general computing unit; and / or executing the second model inference based on the second network parameters and the second neural network code, and returning the execution result to the general computing unit, wherein the execution result includes the model output or the intermediate result.

12. The method according to claim 11, characterized in that The providing the first stored network parameter to the general computing unit includes: A parameter reading command from the general computing unit is received, and a first network parameter stored in the general computing unit and corresponding to the parameter reading command is returned to the general computing unit.

13. The method according to claim 11, characterized in that The performing the second model reasoning based on the second network parameter and the second neural network code, and returning the execution result to the general computing unit, comprises: receiving an execution command from the general computing unit; Execute corresponding second model reasoning based on the second neural network code and the second network parameter corresponding to the execution command to obtain an execution result; The execution result is returned to the general computing unit.

14. The method according to claim 11, characterized in that The first network parameters include at least part of the network parameters of at least one neural network layer inferred by the first model; and / or The second neural network code includes neural network code of at least one neural network layer.

15. A method for processing a neural network model, characterized in that: Applied to a general computing unit, the method comprises: Receiving a neural network model and a deployment instruction sent by a server, wherein the deployment instruction is used to indicate a neural network code and network parameters deployed in a general computing unit, and a neural network code and / or network parameters deployed in a security unit; The neural network model is deployed in the general computing unit and the security unit according to the deployment instruction.

16. The method according to claim 15, characterized in that The deploying the neural network model in the general computing unit and the security unit according to the deployment instruction includes: Storing in a storage unit a first neural network code of a first model inference of a neural network model and at least a portion of first network parameters inferred by the first model inference; Instruct the security unit to store the remaining first network parameters inferred by the first model, and / or store the second neural network code and second network parameters inferred by the second model of the neural network model.

17. The method according to claim 16, characterized in that Also includes: Sending a parameter reading command to the security unit, and receiving a first network parameter corresponding to the parameter reading command returned by the security unit; and / or An execution command is sent to the security unit, and an execution result returned by the security unit after executing the second model reasoning corresponding to the execution command is received.

18. The method according to claim 15, characterized in that The network parameters deployed in the security unit include at least part of the network parameters of at least one neural network layer; and / or The neural network code deployed in the security unit includes code of at least one network layer.

19. A method for processing a neural network model, characterized in that: Applied to a server, the method comprises: A neural network model and a deployment instruction are sent to an application node, wherein the deployment instruction is used to indicate the neural network code and network parameters deployed in a general computing unit, and the neural network code and / or network parameters deployed in a security unit.

20. The method of claim 19, wherein: Also includes: Adjusting the deployment location of the neural network code and / or network parameters, the deployment location including the general computing unit and the security unit; Sending a deployment indication based on the adjustment to the application node.

21. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 11-20.

22. An application node, characterized in that: The device comprises a device body and a computing device as described in any one of claims 1 to 7 or a security unit as described in any one of claims 8 to 10 arranged on the device body.

Citation Information

Cited By

  • Processing method for neural network model, and secure element and computing apparatus

    WO2026129814A1