Convolutional neural network reasoning method, device and equipment based on equipment-side trusted execution environment

By receiving and correcting the intermediate data of the convolutional neural network in a trusted execution environment on the device side, and performing partial inference calculations in client applications, the problems of loss of control and secure memory limitation during device side model deployment are solved, and the security and accuracy of the model are improved.

CN119990310AActive Publication Date: 2025-05-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510030708.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-13
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

When deploying neural network models on the device side, service providers may lose control of proprietary models, and the security memory constraints lead to a decrease in model inference accuracy.

Method used

The convolutional neural network inference method based on the device-side trusted execution environment is adopted. By receiving and correcting the intermediate data of the convolutional neural network in the trusted execution environment, and performing partial inference calculations in the client application, the security and accuracy of the model are ensured.

Benefits of technology

It effectively protects the privacy information of the convolutional neural network, solves the problem of secure memory limitation, and performs model security inference with no precision loss on the device side.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990310A_ABST
    Figure CN119990310A_ABST
Patent Text Reader

Abstract

The invention provides a convolutional neural network reasoning method, device and equipment based on a trusted execution environment of an equipment end, and the method is executed in the trusted execution environment running in the equipment end, and comprises the steps: receiving intermediate data outputted by a current target convolutional layer in a first model after obfuscation calculation; correcting the intermediate data and transmitting the corrected intermediate data to the client application, so as to input the corrected intermediate data into a second model to obtain result data when determining that the target convolutional layer is the last convolutional layer; and receiving result data and inputting the result data into the last full connection layer in the convolutional neural network to obtain reasoning result data. According to the method and the device, on the basis of realizing convolutional neural network security reasoning based on hardware, the problem of limited security memory existing in equipment ends such as edge equipment can be effectively solved, convolutional neural network security reasoning without precision loss can be executed at the equipment end, and privacy information of the convolutional neural network can be effectively protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of model reasoning technology, and in particular to a convolutional neural network reasoning method, apparatus and device based on a device-side trusted execution environment. Background Art

[0002] In recent years, machine learning has been widely used, among which the progress of deep neural networks has played an important role. The large-scale popularization of mobile and IoT devices has also triggered the demand for inference of neural network models on the device side. Although deploying models on the device side can well protect users' privacy data and reduce inference latency, it also introduces new problems. If the model is deployed on the device side, the service provider may lose control of the proprietary model. Attackers can steal key information on the device by cracking the root permissions of smart devices, which may lead to the leakage or abuse of proprietary models and infringe the intellectual property rights of service providers.

[0003] In order to enhance the model privacy protection capability in the device-side reasoning scenario, the isolated computing environment provided by the Trusted Execution Environment (TEE) is used. For example, TrustZone is a specific implementation of TEE. By dividing the system hardware into the secure world (Secure World) and the normal world (Normal World), the computing environment and data are isolated, thereby providing a trusted and untrusted operating environment on one processor at the same time. Even if an attacker has root privileges on the device, he cannot obtain data in the secure world.

[0004] However, the secure memory in the secure world is very limited, and the neural network model cannot be fully loaded into the memory to perform model inference. Quantization and pruning techniques can reduce resource requirements, but usually at the expense of model inference accuracy. Summary of the invention

[0005] In view of this, the embodiments of the present application provide a convolutional neural network reasoning method, apparatus and device based on a device-side trusted execution environment to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of the present application provides a convolutional neural network reasoning method based on a trusted execution environment on a device side, the method being executed in a trusted execution environment running on a device side, the method comprising:

[0007] During the inference process, receiving intermediate data outputted by the current target convolutional layer in the first model corresponding to the convolutional neural network after the obfuscation calculation in the client application; wherein the first model includes multiple convolutional layers connected sequentially in the convolutional neural network, and the target convolutional layer represents any convolutional layer in the first model;

[0008] Correcting the intermediate data according to the correction parameters corresponding to the target convolution layer to obtain corrected intermediate data corresponding to the target convolution layer, and determining whether the pre-acquired input data corresponding to the target convolution layer has been obfuscated, and if so, transmitting the corrected intermediate data corresponding to the target convolution layer to the client application, so that the client application determines whether the target convolution layer is the last convolution layer in the first model, and if so, the client application inputs the corrected intermediate data into the second model corresponding to the convolution neural network, and sends out the result data output by the second model, wherein the second model corresponding to the convolution neural network includes sequentially connected pooling layers and multiple fully connected layers except for each convolution layer and the last fully connected layer in the convolution neural network;

[0009] Receive the result data sent by the client application, and input the result data into the last fully connected layer of the convolutional neural network pre-stored locally to obtain the inference result data corresponding to the output of the last fully connected layer.

[0010] In some embodiments of the present application, the first model after obfuscation calculation in the client application runs in a rich execution environment in the client application, so that the client application inputs the input data corresponding to the current target convolutional layer in the first model into the target convolutional layer in the rich execution environment, so that the target convolutional layer performs a convolution operation on the input data based on the obfuscation weight corresponding to the target convolutional layer and outputs the intermediate data corresponding to the target convolutional layer.

[0011] In some embodiments of the present application, the convolutional neural network reasoning method based on the device-side trusted execution environment also includes:

[0012] If it is determined that the pre-acquired input data corresponding to the target convolutional layer has not been obfuscated, random noise is added to the corrected intermediate data corresponding to the target convolutional layer to complete the obfuscation of the corrected intermediate data;

[0013] The corrected intermediate data after the obfuscation process is transmitted to the client application.

[0014] In some embodiments of the present application, if the client application determines that the target convolutional layer is not the last convolutional layer in the first model, the client application uses the corrected intermediate data corresponding to the target convolutional layer as the input data corresponding to the next convolutional layer corresponding to the target convolutional layer, and inputs the input data corresponding to the next convolutional layer into the next convolutional layer, so that the next convolutional layer performs a convolution operation on the input data corresponding to the next convolutional layer based on the confusion weight corresponding to the next convolutional layer, and then outputs the intermediate data corresponding to the next convolutional layer;

[0015] Correspondingly, after transmitting the corrected intermediate data corresponding to the target convolutional layer to the client application and before receiving the result data sent by the client application, the method further includes:

[0016] During the inference process, receiving intermediate data outputted by a next convolutional layer corresponding to the target convolutional layer in the first model after the obfuscation calculation;

[0017] The intermediate data corresponding to the next convolution layer is corrected according to the correction parameters corresponding to the next convolution layer to obtain the corrected intermediate data corresponding to the next convolution layer, and it is determined whether the pre-acquired input data corresponding to the next convolution layer has been obfuscated. If so, the corrected intermediate data corresponding to the next convolution layer is transmitted to the client application.

[0018] In some embodiments of the present application, the first model and the second model corresponding to the convolutional neural network constitute the first half model of the convolutional neural network, and the last fully connected layer corresponding to the convolutional neural network constitutes the second half model of the convolutional neural network; and the first model after the obfuscation calculation is generated by performing obfuscation calculation on the first model based on a preset obfuscation model generation algorithm in advance, and the first model after the obfuscation calculation includes the obfuscation weights and correction parameters corresponding to each convolution layer;

[0019] The first model and the second model after obfuscation calculation are sent to the client application in advance for storage.

[0020] In some embodiments of the present application, before receiving the intermediate data of the current target convolution layer output in the first model after the obfuscation calculation corresponding to the convolutional neural network in the client application during the inference process, it also includes:

[0021] Receive and locally store the correction parameters corresponding to the last fully connected layer in the convolutional neural network and each convolution layer in the first model after confusion calculation.

[0022] Another aspect of the present application provides a convolutional neural network reasoning device based on a trusted execution environment on a device side, wherein the device is arranged in a trusted execution environment running on a device side, and includes:

[0023] An intermediate data receiving module, used to receive the intermediate data output by the current target convolution layer in the first model after the obfuscation calculation corresponding to the convolutional neural network in the client application during the inference process; wherein the first model includes multiple convolutional layers connected in sequence in the convolutional neural network, and the target convolutional layer represents any convolutional layer in the first model;

[0024] A correction and data transmission module, used to correct the intermediate data according to the correction parameters corresponding to the target convolution layer to obtain corrected intermediate data, and determine whether the pre-acquired input data corresponding to the target convolution layer has been obfuscated. If so, the corrected intermediate data is transmitted to the client application, so that the client application determines whether the target convolution layer is the last convolution layer in the first model. If so, the client application inputs the corrected intermediate data into the second model corresponding to the convolutional neural network and sends out the result data output by the second model, wherein the second model corresponding to the convolutional neural network includes sequentially connected pooling layers and multiple fully connected layers in the convolutional neural network except for each convolution layer and the last fully connected layer;

[0025] The hierarchical reasoning module is used to receive the result data sent by the client application, and input the result data into the last fully connected layer of the convolutional neural network pre-stored locally, to obtain the reasoning result data corresponding to the output of the last fully connected layer.

[0026] A third aspect of the present application provides a convolutional neural network reasoning device, comprising: a client application deployed in a rich execution environment and a trusted application deployed in a trusted execution environment;

[0027] The trusted application is used to execute the aforementioned convolutional neural network reasoning method based on the device-side trusted execution environment.

[0028] In some embodiments of the present application, the client application is used to perform the following:

[0029] Determining in the rich execution environment whether the target convolutional layer is the last convolutional layer in the first model;

[0030] If yes, input the corrected intermediate data into the second model corresponding to the convolutional neural network, and send out the result data output by the second model;

[0031] If not, the corrected intermediate data corresponding to the target convolution layer is used as the input data corresponding to the next convolution layer corresponding to the target convolution layer, and the input data corresponding to the next convolution layer is input into the next convolution layer, so that the next convolution layer performs a convolution operation on the input data corresponding to the next convolution layer based on the confusion weight corresponding to the next convolution layer, and then outputs the intermediate data corresponding to the next convolution layer, and sends the intermediate data corresponding to the next convolution layer to the trusted application.

[0032] In some embodiments of the present application, before determining in the rich execution environment whether the target convolutional layer is the last convolutional layer in the first model, the client application further executes the following:

[0033] The first model and the second model after the obfuscated calculation are received and locally stored in the rich execution environment.

[0034] The fourth aspect of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the convolutional neural network reasoning method based on the device-side trusted execution environment is implemented.

[0035] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the convolutional neural network reasoning method based on a device-side trusted execution environment.

[0036] The convolutional neural network inference method based on a device-side trusted execution environment provided in the present application is executed in a trusted execution environment running on the device side, and the method includes: receiving intermediate data output by the current target convolutional layer in the first model after obfuscation calculation corresponding to the convolutional neural network in the client application during the inference process; wherein the first model includes multiple convolutional layers connected in sequence in the convolutional neural network, and the target convolutional layer represents any convolutional layer in the first model; correcting the intermediate data according to the correction parameters corresponding to the target convolutional layer to obtain the corrected intermediate data corresponding to the target convolutional layer, and judging whether the pre-acquired input data corresponding to the target convolutional layer has been obfuscated, and if so, transmitting the corrected intermediate data corresponding to the target convolutional layer to the client application, so that the client application judges whether the target convolutional layer is the first model , if yes, the client application inputs the corrected intermediate data into the second model corresponding to the convolutional neural network, and sends the result data output by the second model, wherein the second model corresponding to the convolutional neural network includes sequentially connected pooling layers and multiple fully connected layers in the convolutional neural network except for each convolutional layer and the last fully connected layer; receiving the result data sent by the client application, and inputting the result data into the last fully connected layer in the convolutional neural network pre-stored locally, to obtain the inference result data corresponding to the output of the last fully connected layer, which can effectively solve the problem of limited secure memory on the device side such as edge devices on the basis of hardware-based secure inference of convolutional neural networks, can perform secure inference of convolutional neural networks without precision loss on the device side, and can effectively protect the privacy information of convolutional neural networks.

[0037] Additional advantages, purposes, and features of the present application will be partially described in the following description, and will become partially apparent to those skilled in the art after studying the following, or may be learned from the practice of the present application. The purposes and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the specification and the drawings.

[0038] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present application are not limited to the above specific description, and the above and other purposes that can be achieved by the present application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings described herein are used to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. The components in the drawings are not drawn to scale, but are only for the purpose of illustrating the principles of the present application. In order to facilitate the illustration and description of some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:

[0040] Figure 1 This is a first flow chart of a convolutional neural network reasoning method based on a device-side trusted execution environment in one embodiment of the present application.

[0041] Figure 2 This is a second flow chart of a convolutional neural network reasoning method based on a device-side trusted execution environment in one embodiment of the present application.

[0042] Figure 3 This is a schematic diagram of the structure of a convolutional neural network reasoning device based on a device-side trusted execution environment in one embodiment of the present application.

[0043] Figure 4 Schematic diagram of the structure of a convolutional neural network inference device in one embodiment of the present application.

[0044] Figure 5 A flowchart of a convolutional neural network inference process executed by a client application in a convolutional neural network inference device in one embodiment of the present application.

[0045] Figure 6 This is a schematic overview of a TrustZone-based device-side model security reasoning method in an application example of the present application.

[0046] Figure 7 This is a flowchart of a specific implementation of a device-side model security reasoning method based on TrustZone in an application example of the present application.

[0047] Figure 8 This is a schematic diagram of the structure of the model security reasoning process in an application example of this application.

[0048] Fig. 9 This is a flowchart of a specific implementation of the model security reasoning process in an application example of the present application. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the implementation modes and the accompanying drawings. Here, the illustrative implementation modes and descriptions of the present application are used to explain the present application, but are not intended to limit the present application.

[0050] It should also be noted here that in order to avoid obscuring the present application due to unnecessary details, only the structures and / or processing steps closely related to the scheme according to the present application are shown in the accompanying drawings, while other details that are not very relevant to the present application are omitted.

[0051] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0052] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0053] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0054] In recent years, machine learning has been widely used, among which the progress of deep neural networks has played an important role. Deep neural networks are applied in various fields, including image classification, anomaly detection, and health monitoring. The large-scale popularization of mobile and IoT devices has also triggered the demand for neural network model inference on the device side. Traditionally, to use cloud-based model inference services, users need to transfer their data from the device side to a remote cloud server through the network, and return the results after the inference is completed, which is also called machine learning as a service (MLaaS). There are two main problems with cloud-based model inference services. The first is data privacy. Deep learning technology is increasingly used in sensitive environments. Users will lose control of their data after uploading private data such as images and health information collected on the device side to a remote cloud server. The second is real-time responsiveness. Sending data to a remote cloud server for model inference may introduce additional queuing, transmission, and inference delays. However, in special scenarios such as autonomous driving and elderly alarm systems, even an additional 0.1s delay is difficult to accept.

[0055] In the machine learning as a service scenario, the service provider uses a pre-inferred neural network model to provide inference services to users. The client holds private data x, and the neural network model that implements function f is deployed on a remote server. The client hopes to obtain the result of the neural network model's model inference on x. However, the client wants to keep its private data x private, that is, x cannot be uploaded to the cloud server without encryption; at the same time, the neural network model is the intellectual property of the service provider and cannot be deployed to the client without encryption. In general, in the model inference stage, the input data and the neural network model are held by different roles, and the privacy of the data and the security of the model must be considered. Model security inference needs to achieve inference calculations on client input data while meeting these security requirements.

[0056] In recent years, secure reasoning methods based on encryption technology and hardware isolation have become a hot topic in academia. For example, homomorphic encryption (HE) is a privacy-enhancing technology based on encryption that allows calculations to be performed on encrypted data. The output of the calculation is still in encrypted form, and the result obtained by decrypting the encrypted output is consistent with the result of directly calculating the unencrypted data. Service providers can perform reasoning without decrypting the data and return the encrypted result to the user, who then decrypts it. This method ensures the confidentiality of user data, but HE only supports addition or multiplication operations. Even though the fully homomorphic encryption (FHE) scheme can support both addition and multiplication calculations, FHE has extremely high requirements for hardware performance, and in practical applications, there are problems such as noise accumulation and high overhead in the calculation process. Another countermeasure is a secure reasoning method based on hardware isolation, which deploys the model on the local device for model reasoning. Reasoning the neural network model on the device helps protect the user's privacy because the user's private data never leaves the device. Therefore, many studies have attempted to transfer complex neural network models from the cloud to device-side deployment. By deploying deep neural network models on mobile devices, the dependence on network bandwidth and cloud server resources can be reduced, thereby improving the performance of mobile applications and minimizing the communication costs and delays of applications.

[0057] However, while deploying models on the device side can well protect the user's privacy data and reduce inference latency, it also introduces a new problem. If the model is deployed on the device side, the service provider may lose control of the proprietary model. Attackers can steal key information on the device by cracking the root permissions of smart devices, which may lead to the leakage or abuse of proprietary models and infringe the intellectual property rights of service providers. In addition to the possible leakage of the weights of proprietary models, the inference data of the model also has the risk of privacy leakage. The information of the inference data may be leaked during the inference process, which is called membership inference attack (MIA). After the inference is completed, the deep neural network model will "remember" the features in the inference data. This feature memory ability enables the model to identify samples that show similar patterns in other data. However, the model's memory usually contains more specific information in the inference data set that is not related to the target pattern (category information that the model needs to classify). This feature brings risks to the privacy of the inference data.

[0058] In order to enhance the privacy protection capability on the device side, protection can be implemented at the hardware level, among which the trusted execution environment (TEE) has been widely studied. TEE technologies represented by ARM TrustZone and Intel SGX provide users with a trusted isolated computing environment. However, these technologies are limited by memory and computing power, and are less efficient when processing large-scale neural networks. Taking TrustZone as an example, its secure memory is usually only about 10MB, which is difficult to meet the corresponding resource requirements for deep learning tasks. In summary, with the development of mobile and IoT devices, the demand for secure model reasoning on the device side is increasing. How to ensure model security while protecting user privacy has become an important issue that needs to be solved in the field of machine learning.

[0059] Based on this, in order to solve the problem of limited secure memory on devices such as edge devices on the basis of secure reasoning of model execution on the device side, the embodiments of the present application respectively provide a convolutional neural network reasoning method based on a device-side trusted execution environment, a convolutional neural network reasoning device based on a device-side trusted execution environment for executing the convolutional neural network reasoning method based on a device-side trusted execution environment, a system, an electronic device, a computer-readable storage medium and a computer program product. According to the characteristics of the neural network model reasoning process, with the help of the trusted execution environment provided by the hardware, a model security reasoning method with no precision loss is implemented, which can protect the weight privacy and reasoning data privacy of proprietary models, and promote the safe deployment and rapid implementation of artificial intelligence models on wearable mobile devices.

[0060] The details are described in detail through the following examples.

[0061] Based on this, the embodiment of the present application provides a convolutional neural network reasoning method based on a device-side trusted execution environment that can be implemented by a convolutional neural network reasoning device based on a device-side trusted execution environment. The method is executed in a trusted execution environment running on the device side, see Figure 1 The convolutional neural network reasoning method based on the device-side trusted execution environment specifically includes the following contents:

[0062] Step 100: During the inference process, intermediate data outputted by the current target convolutional layer in the first model after obfuscation calculation corresponding to the convolutional neural network in the client application is received; wherein the first model includes multiple convolutional layers connected in sequence in the convolutional neural network, and the target convolutional layer represents any convolutional layer in the first model.

[0063] It can be understood that the convolutional neural network reasoning method based on the device-side trusted execution environment provided in the embodiment of the present application includes an reasoning process of reasoning convolutional neural network. In practical applications, it may include one or more reasoning processes of reasoning convolutional neural network, that is, the convolutional neural network reasoning method based on the device-side trusted execution environment provided in the embodiment of the present application can be executed separately for each different input data to be reasoned.

[0064] In one or more embodiments of the present application, a convolutional neural network (CNN) is a deep learning model that can effectively extract features of data such as images and videos. It simulates the feature extraction ability of the human visual system through convolution operations and performs well in various visual tasks. The main principle of CNN is to extract local features from data using convolution operations. The convolution layer extracts features such as edges and textures of the data layer by layer by applying a set of inferable filters (convolution kernels), and then combines the pooling layer for downsampling. The layer-by-layer superposition of features enables CNN to learn high-level semantic information from the underlying low-level features. The convolution operation has the characteristics of translation invariance and local receptive field, which enables CNN to perform well on data with spatial structure such as images.

[0065] Step 200: Correct the intermediate data according to the correction parameters corresponding to the target convolution layer to obtain the corrected intermediate data corresponding to the target convolution layer, and determine whether the pre-acquired input data corresponding to the target convolution layer has been obfuscated. If so, transmit the corrected intermediate data corresponding to the target convolution layer to the client application, so that the client application determines whether the target convolution layer is the last convolution layer in the first model. If so, the client application inputs the corrected intermediate data into the second model corresponding to the convolutional neural network and sends out the result data output by the second model, wherein the second model corresponding to the convolutional neural network includes sequentially connected pooling layers and multiple fully connected layers in the convolutional neural network except for each convolution layer and the last fully connected layer.

[0066] Among them, the intermediate data refers to the input data corresponding to the current target convolution layer in the first model input by the client application into the target convolution layer. If the target convolution layer is the first convolution layer in the first model, the input data corresponding to the target convolution layer is the local inference data, which is set according to the application scenario of the convolution neural network. For example, if the convolution neural network is ultimately used for face recognition, the inference data is the face image data with a label for uniquely identifying a person. If the target convolution layer is not the first convolution layer in the first model, the input data corresponding to the target convolution layer is the corrected intermediate data corresponding to the intermediate data output by the previous convolution layer corresponding to the target convolution layer.

[0067] It can be understood that the client application (CA) executes: determining whether the target convolution layer is the last convolution layer in the first model; if so, the client application inputs the corrected intermediate data into the second model corresponding to the convolutional neural network, and sends out the result data output by the second model and other contents.

[0068] In addition, the convolutional neural network inference device based on the device-side trusted execution environment for executing the convolutional neural network inference method based on the device-side trusted execution environment can also be implemented using a trusted application (TA), and the trusted application and the client application are set in the same device, the client application can be deployed in the rich execution environment REE of the device, and the trusted application is deployed in the trusted execution environment TEE of the device.

[0069] In step 200, the result data refers to the data output by the second model after the corrected intermediate data is input into the second model corresponding to the convolutional neural network.

[0070] Step 300: Receive the result data sent by the client application, and input the result data into the last fully connected layer of the convolutional neural network pre-stored locally to obtain the inference result data corresponding to the output of the last fully connected layer.

[0071] In one or more embodiments of the present application, TrustZone can be used to implement a trusted execution environment and a rich execution environment on the device side. TrustZone technology is a hardware architecture designed by ARM for embedded devices, which builds a security framework to resist various possible attacks. TrustZone technology provides system-level isolation for upper-layer applications by implementing interrupt isolation, isolation of RAM and ROM inside the chip, isolation of external RAM, isolation of peripherals, etc. It has been applied to billions of processors to protect the security environment and sensitive data of various applications, and application scenarios include identity authentication, payment, and content protection. TrustZone isolates data and code by dividing the system hardware into a secure world (Secure World) and a normal world (Normal World), thereby providing both trusted and untrusted operating environments on one processor. TrustZone's isolation design ensures that even if the operating system or application is attacked, the data in the secure world is still protected.

[0072] In the scenario of the convolutional neural network reasoning method based on the device-side trusted execution environment provided by the embodiment of the present application, TrustZone is mainly responsible for the functions of private data storage and trusted computing. For example, the secure storage function it provides is used to encrypt and store model weight information and other private data; the trusted computing function it provides is used to perform reasoning calculations and other calculations on some layers of the neural network model. Since TrustZone implements a trusted execution environment at the hardware level, its security can be fully trusted.

[0073] From the above description, it can be seen that the convolutional neural network reasoning method based on the device-side trusted execution environment provided in the embodiment of the present application can effectively solve the problem of limited secure memory on device sides such as edge devices on the basis of hardware-based implementation of convolutional neural network security reasoning, can perform convolutional neural network security reasoning without precision loss on the device side, and can effectively protect the privacy information of the convolutional neural network.

[0074] In order to further improve the reasoning resource availability and reasoning validity of the first model after obfuscation calculation in the client application, in a convolutional neural network reasoning method based on a device-side trusted execution environment provided in an embodiment of the present application, the first model after obfuscation calculation in the client application in the convolutional neural network reasoning method based on a device-side trusted execution environment runs in a rich execution environment in the client application, so that the client application inputs the input data corresponding to the current target convolutional layer in the first model into the target convolutional layer in the rich execution environment, so that the target convolutional layer performs a convolution operation on the input data based on the obfuscation weight corresponding to the target convolutional layer, and then outputs the intermediate data corresponding to the target convolutional layer.

[0075] It is understandable that the rich execution environment may also be referred to as a rich execution environment (REE), which refers to an environment in which an operating system is running, and may run general-purpose OS (Operating System) such as Android and IOS. REE is an open environment that is vulnerable to attack, so it is necessary to adopt the first model after obfuscation calculation, perform most of the reasoning calculations in the rich execution environment (REE), and perform correction and a small amount of reasoning calculations on the intermediate data (IR) in the trusted execution environment (TEE). In one or more embodiments of the present application, the first model after obfuscation calculation may also be referred to as an obfuscated model, and the intermediate data may also be referred to as an intermediate result.

[0076] In order to further effectively protect the privacy information of the convolutional neural network, in a convolutional neural network reasoning method based on a device-side trusted execution environment provided in an embodiment of the present application, see Figure 2, step 200 in the convolutional neural network reasoning method based on the device-side trusted execution environment specifically includes the following contents:

[0077] Step 210: Correct the intermediate data according to the correction parameters corresponding to the target convolution layer to obtain corrected intermediate data corresponding to the target convolution layer.

[0078] Step 220: Determine whether the pre-acquired input data corresponding to the target convolutional layer has been obfuscated;

[0079] If yes, execute step 230 ; if no, execute step 240 .

[0080] Step 230: Transmit the corrected intermediate data corresponding to the target convolutional layer to the client application, so that the client application determines whether the target convolutional layer is the last convolutional layer in the first model. If so, the client application inputs the corrected intermediate data into the second model corresponding to the convolutional neural network, and issues the result data output by the second model, and then the convolutional neural network reasoning device based on the device-side trusted execution environment executes step 300; if not, the client application uses the corrected intermediate data corresponding to the target convolutional layer as the input data corresponding to the next convolutional layer corresponding to the target convolutional layer, and inputs the input data corresponding to the next convolutional layer into the next convolutional layer, so that the next convolutional layer performs a convolution operation on the input data corresponding to the next convolutional layer based on the confusion weight corresponding to the next convolutional layer, and then outputs the intermediate data corresponding to the next convolutional layer, and then the convolutional neural network reasoning device based on the device-side trusted execution environment executes the following step 310.

[0081] Step 240: adding random noise to the corrected intermediate data corresponding to the target convolutional layer to complete the obfuscation processing of the corrected intermediate data, and then executing step 250.

[0082] It is understandable that in order to prevent attackers from inferring the values ​​of model weights using input data and intermediate results (IR), random noise needs to be added before copying the IR to the untrusted normal world to hide the true values ​​of the intermediate results.

[0083] Specifically, if the normal world obtains the input and the corrected non-zero intermediate results, the attacker can infer the value of the weights (solving the linear equation) through sufficient queries, so the non-zero IR should be protected before being copied to the untrusted normal world by adding the activation function result f(X l ) to add random noise. If the input of the l-1th layer is not confused, its corrected output X l, a one-time random value e is added to IR. This scheme supports any activation function that can capture nonlinear properties and improve the accuracy of the model. It is represented as e l The random noise is added to the intermediate results. l is treated as confidential data and stored privately in secure storage. The output of the intermediate result protection process is X l ′=f(X l )+e l This process is only used for convolution operations with unobfuscated inputs, because if there is only obfuscated input, even if the correct intermediate result is obtained, the value of the weight cannot be solved.

[0084] Step 250: Transmit the corrected intermediate data after the obfuscation process to the client application, and then execute step 300 or the following step 310.

[0085] In order to further improve the execution reliability and effectiveness of the convolutional neural network reasoning method based on the device-side trusted execution environment, in a convolutional neural network reasoning method based on the device-side trusted execution environment provided in an embodiment of the present application, if the client application judges that the target convolutional layer is not the last convolutional layer in the first model, the client application uses the corrected intermediate data corresponding to the target convolutional layer as the input data corresponding to the next convolutional layer corresponding to the target convolutional layer, and inputs the input data corresponding to the next convolutional layer into the next convolutional layer, so that the next convolutional layer performs a convolution operation on the input data corresponding to the next convolutional layer based on the confusion weight corresponding to the next convolutional layer, and then outputs the intermediate data corresponding to the next convolutional layer;

[0086] For the corresponding Figure 2 After step 230 and step 250, and before step 300, the convolutional neural network reasoning method based on the device-side trusted execution environment further specifically includes the following contents:

[0087] Step 310: During the inference process, intermediate data outputted by the next convolutional layer corresponding to the target convolutional layer in the first model after the confusion calculation is received.

[0088] Step 320: Correct the intermediate data corresponding to the next convolution layer according to the correction parameters corresponding to the next convolution layer to obtain the corrected intermediate data corresponding to the next convolution layer, and determine whether the pre-acquired input data corresponding to the next convolution layer has been obfuscated; if so, transmit the corrected intermediate data corresponding to the next convolution layer to the client application; if not, add random noise to the corrected intermediate data corresponding to the next convolution layer to complete the obfuscation of the corrected intermediate data, and transmit the corrected intermediate data corresponding to the next convolution layer that has been obfuscated to the client application.

[0089] CNN models are generally composed of convolutional layers, activation layers, pooling layers, and fully connected layers. Different combinations of these layers can extract multi-level features of data, making the network suitable for learning and classifying complex patterns. The following is an introduction to the main layers:

[0090] (1) Convolutional layer: The convolution kernel is used to extract features from the input data through a sliding window operation. The parameters of the convolutional layer include the size, number, step size, and padding of the convolution kernel, which determine the effect of feature extraction. The data features extracted by the convolution kernel will form a feature map, which reflects the spatial pattern of the data.

[0091] (1) Normalization layer: After completing the convolution calculation, the data is usually normalized and converted into a standard distribution range to ensure the stability of the network.

[0092] (1) Activation layer: The nonlinear features of CNN rely on activation functions, such as ReLU, Sigmoid, and Tanh. ReLU is the most commonly used activation function, which improves the inference speed and stability of the model by outputting negative numbers as zero.

[0093] (1) Pooling layer: The pooling layer is used to reduce the dimension of the feature map, extract the core features of the data, and reduce the computational complexity. Common pooling methods include average pooling and maximum pooling, which can effectively suppress overfitting.

[0094] (1) Fully connected layer: maps the features extracted by the previous convolutional layer to the final classification or regression results. The fully connected layer linearly combines all inputs and outputs the final prediction result through an activation function. It is suitable for the output layer of the task.

[0095] When performing model inference on the device side, the efficiency of CNN enables it to perform complex image classification tasks with limited resources.

[0096] Based on this, in order to further improve the execution reliability and effectiveness of the convolutional neural network reasoning method based on the device-side trusted execution environment, in a convolutional neural network reasoning method based on a device-side trusted execution environment provided in an embodiment of the present application, the first model and the second model corresponding to the convolutional neural network constitute the first half model M of the convolutional neural network. pre , the last fully connected layer corresponding to the convolutional neural network constitutes the second half of the convolutional neural network model M post ; and the first model after the obfuscation calculation is generated by performing obfuscation calculation on the first model based on a preset obfuscation model generation algorithm in advance, which can be written as the obfuscation model M p ' re , and the first model after confusion calculation contains the confusion weights W corresponding to each convolution layer l ′ and correction parameter r l , where l represents the lth convolutional layer.

[0097] The first model and the second model after obfuscation calculation are sent to the client application in advance for storage.

[0098] For the corresponding Figure 2 , the convolutional neural network reasoning method based on the device-side trusted execution environment further specifically includes the following contents before step 100:

[0099] Step 010: Receive and locally store the correction parameters corresponding to the last fully connected layer in the convolutional neural network and each convolutional layer in the first model after confusion calculation.

[0100] From the software level, the present application also provides a convolutional neural network reasoning device based on a device-side trusted execution environment for executing all or part of the convolutional neural network reasoning method based on a device-side trusted execution environment, see Figure 3 The convolutional neural network reasoning device based on the device-side trusted execution environment specifically includes the following contents:

[0101] The intermediate data receiving module 10 is used to receive the intermediate data output by the current target convolution layer in the first model after the obfuscation calculation corresponding to the convolutional neural network in the client application during the inference process; wherein the first model includes multiple convolutional layers connected in sequence in the convolutional neural network, and the target convolutional layer represents any convolutional layer in the first model;

[0102] A correction and data transmission module 20, used to correct the intermediate data according to the correction parameters corresponding to the target convolution layer to obtain corrected intermediate data, and determine whether the pre-acquired input data corresponding to the target convolution layer has been obfuscated. If so, the corrected intermediate data is transmitted to the client application, so that the client application determines whether the target convolution layer is the last convolution layer in the first model. If so, the client application inputs the corrected intermediate data into the second model corresponding to the convolutional neural network and sends out the result data output by the second model, wherein the second model corresponding to the convolutional neural network includes sequentially connected pooling layers and multiple fully connected layers in the convolutional neural network except for each convolution layer and the last fully connected layer;

[0103] The hierarchical reasoning module 30 is used to receive the result data sent by the client application, and input the result data into the last fully connected layer of the convolutional neural network pre-stored locally to obtain the reasoning result data corresponding to the output of the last fully connected layer.

[0104] The embodiment of the convolutional neural network inference device based on the device-side trusted execution environment provided in the present application can be specifically used to execute the processing flow of the embodiment of the convolutional neural network inference method based on the device-side trusted execution environment in the above-mentioned embodiment. Its functions will not be repeated here, and reference can be made to the detailed description of the embodiment of the convolutional neural network inference method based on the device-side trusted execution environment.

[0105] The part of the convolutional neural network reasoning device based on the device-side trusted execution environment that performs convolutional neural network reasoning based on the device-side trusted execution environment can be completed in the edge device. The specific selection can be based on the processing capability of the edge device and the limitations of the user's usage scenario. This application is not limited to this. If all operations are completed in the edge device, the edge device may also include a processor for specific processing of convolutional neural network reasoning based on the device-side trusted execution environment.

[0106] The above-mentioned edge device may have a communication module (i.e., a communication unit) that can communicate with a remote client device to realize data transmission between the client device. The client device may also communicate with a server, which may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0107] The server and the client device may communicate with each other using any suitable network protocol, including network protocols that have not yet been developed on the date of filing this application. The network protocols may include, for example, TCP / IP, UDP / IP, HTTP, HTTPS, etc. Of course, the network protocols may also include, for example, RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer) protocols used on top of the above protocols.

[0108] From the above description, it can be seen that the convolutional neural network reasoning device based on the device-side trusted execution environment provided in the embodiment of the present application can effectively solve the problem of limited secure memory on the device side such as edge devices on the basis of hardware-based implementation of convolutional neural network security reasoning, can perform convolutional neural network security reasoning without precision loss on the device side, and can effectively protect the privacy information of the convolutional neural network.

[0109] Based on the above embodiments, the present application also provides a convolutional neural network reasoning device, see Figure 4 , the convolutional neural network reasoning device specifically includes the following contents:

[0110] Client applications deployed in a rich execution environment and trusted applications deployed in a trusted execution environment;

[0111] The trusted application is used to execute the convolutional neural network reasoning method based on the device-side trusted execution environment provided in the above embodiment. In order to further improve the memory capacity and reasoning effectiveness of the convolutional neural network reasoning, in one embodiment of the convolutional neural network reasoning device provided in this application, see Figure 5 , the client application in the convolutional neural network inference device is used to perform the following:

[0112] Step 400: Determine in the rich execution environment whether the target convolutional layer is the last convolutional layer in the first model;

[0113] If yes, execute step 500 ; if no, execute step 600 .

[0114] Step 500: Input the corrected intermediate data into the second model corresponding to the convolutional neural network, and send out the result data output by the second model so that the trusted application executes step 300.

[0115] Step 600: Using the corrected intermediate data corresponding to the target convolution layer as the input data corresponding to the next convolution layer corresponding to the target convolution layer, and inputting the input data corresponding to the next convolution layer into the next convolution layer, so that the next convolution layer performs a convolution operation on the input data corresponding to the next convolution layer based on the confusion weight corresponding to the next convolution layer, and then outputs the intermediate data corresponding to the next convolution layer, and sends the intermediate data corresponding to the next convolution layer to the trusted application, so that the trusted application executes step 310.

[0116] And, in order to further improve the effectiveness and reliability of convolutional neural network reasoning, in one embodiment of the convolutional neural network reasoning device provided in this application, see Figure 5 , the client application in the convolutional neural network inference device is also used to perform the following before step 400:

[0117] Step 020: Receive and locally store the first model and the second model after obfuscation calculation in the rich execution environment.

[0118] To further illustrate the above embodiments, the present application also provides a convolutional neural network reasoning method using TrustZone as a device-side trusted execution environment and a rich execution environment, that is, a device-side model security reasoning method based on TrustZone, see Figure 6 ,The device-side model security reasoning method based on TrustZone is divided into three stages, namely the ,preparation stage, the initialization stage and the reasoning stage.

[0119] 1. Preparation stage: neural network model division to obtain the first half of the indirect safety reasoning model M pre and the second half of the direct security reasoning model M post ; Then, the model M is obtained by the confusion model generation algorithm pre The obfuscated version of model M p ' re .

[0120] Specifically, the original model with good inference is first divided into the first half model M pre and the second half of the model M post Next, the confusion model generation algorithm proposed in the application example of this application is used to obtain the model M pre The obfuscated version of model M p ' re .

[0121] 2. Initialization phase: Next, the model M p ' re and Model M postDeployed in the normal world and the secure world respectively, and store private data securely;

[0122] Specifically, the operating environment is built on the local device. First, the service provider provides the client application (CA) and the trusted application (TA) code, which contains the basic environment required to run the model; then the obfuscated model M p ' re And the second half of the model M post All layers except the last fully connected layer are deployed in the normal world. To ensure the safety of loading and running the program on the edge device, the secure world will hash the initial state of the application and feed this signature back to the service provider to ensure trusted operation. After the authentication is completed, the service provider sends the second half of the model M to the secure world. post The last fully connected layer in , and provides correction parameters for intermediate result correction.

[0123] 3. Inference phase: The user initiates an inference request, M p ' re The model starts reasoning and sends the reasoning results to the model M in the secure world through shared memory. post Continue to execute; finally, filter the inference results and return them to the user for viewing.

[0124] Specifically, users provide input and obtain inference results. First, users pass input data to the client application on the edge device; the client application interacts with the trusted application and uses the obfuscated model M p ' re Reasoning and correction of model M post The reasoning and result filtering are performed to complete secure model reasoning; finally, the client application feeds back the reasoning results to the user.

[0125] Through this method, users can perform model security reasoning without precision loss on the device side, effectively protecting the privacy information of the model. The specific implementation process of the device-side model security reasoning method based on TrustZone is as follows: Figure 7 shown.

[0126] The TrustZone-based device-side model security reasoning method performs most of the reasoning calculations in the rich execution environment (REE) and corrects the intermediate results (IR) and a small amount of reasoning calculations in the trusted execution environment (TEE). The structural diagram of the model security reasoning process is shown in the figure below. Figure 8As shown in Figure 1, the model security reasoning process consists of four parts: I) De-noise the IR using the correction parameter r. II) Hide the real intermediate result with random noise e. III) De-noise the IR using r and e. IV) Model partition execution. The obfuscated model weights W′ are stored in the untrusted storage space, while only r, e and pre-computed u are saved in the secure storage. The first three parts can be summarized as obfuscated model reasoning and correction.

[0127] Starting from the user's request for model reasoning, the complete reasoning operation process is realized through the model security reasoning system. Fig. 9 As shown. From the process, we can see that by initiating a model inference request from the CA on the REE side, the preparation for model inference is completed by reading files from the file system, including the loading of input data and obfuscation model. Then, the obfuscated model inference and correction are performed, and most of the calculations are performed on the REE side. Only the intermediate results are corrected through TEE, which effectively reduces the amount of calculation on the TEE side and alleviates the demand for secure memory. After the convolution layer with obfuscated weights is executed, the subsequent unobfuscated pooling layer and fully connected layer are deployed in REE and TEE respectively using partitioned execution. Specifically, the first few layers are executed in the resource-rich REE, and only the last layer of the model is executed in TEE. After all layers of the model are executed, the inference results are filtered using the topK method, and some inference results are returned to the REE side for users to view.

[0128] In general, the application example of this application proposes a device-side model security reasoning method and device based on TrustZone, which effectively protects the privacy information of the model. The main work content is as follows. (1) Protecting the weight information of the neural network model: Protecting the confidentiality of the proprietary model by generating an obfuscated model. (2) Optimizing performance: Optimizing performance by minimizing the model reasoning calculations running in the secure world on the edge device.

[0129] Specifically, the core technical contents of the application examples of this application are described as follows:

[0130] (1) Neural network model division

[0131] Since deep neural networks follow a layered architecture, the neural network model can be divided. In this method, the deep neural network model is divided into two parts, with the first fully connected layer layer partition is the dividing point, and the first half of the model M is obtained pre and the second half of the model M post The model M pre Performs convolution, activation, and pooling operations, including most of the calculations in the reasoning process; Model M post Perform calculations on the fully connected layer.

[0132] By dividing the neural network model, model reasoning can be performed more flexibly, making full use of the computing resources on the device while ensuring safety.

[0133] (2) Confusion Model Generation

[0134] An algorithm for generating an obfuscated version of a neural network model is proposed to optimize the performance of secure reasoning and improve memory efficiency. In order to prevent users from restoring neural network models or illegally reselling reasoning services, the obfuscated model needs to meet the following three goals: (1) the prediction results of the input are significantly changed; (2) the constructed obfuscated parameters are difficult to identify;

[0135] This method focuses on the weights of the convolutional layer, because in the CNN model, more than 90% of the computational operations come from convolution operations. Experiments have shown that even if only part of the weights of the convolutional layer are modified, the accuracy of the model will drop significantly to an unacceptable level. In addition, although noise can also be added to the weights of the fully connected layer, due to the large number of weights in the fully connected layer, the amount of computation is greater, and the use of noise correction schemes may affect the real-time performance of reasoning.

[0136] In order to cope with the memory and computational performance limitations of the secure world, this method will prioritize modifying key weights that contribute more to the model. This idea is inspired by the pruning method in model compression. It is proposed in this paper that small-norm filters contribute less to model performance. Therefore, in order to significantly reduce the inference accuracy of the obfuscated model, weights that are rarely deleted in the pruning method should be modified first. Considering the limited secure memory and inference efficiency requirements, this method chooses to modify only a small number of larger weights instead of all weights. Experiments have shown that even if only part of the weights are modified, the input prediction label will change significantly. In order to achieve the second security goal (i.e., obfuscated weights are difficult to identify), the obfuscated weights need to be carefully designed. Specifically, given an obfuscated model, it is difficult for an attacker to infer which weights are "encrypted" or "unencrypted".

[0137] The proportion of processed filter weights is set to p%. In order to make these carefully designed weights difficult to identify, the processed p% weights should be consistent with the remaining 1-p% weight distribution in the same convolution kernel. The process of generating the confusion model is described in detail in Algorithm 1 as shown in Table 1. The algorithm involves the noise amplification factor k, the number of convolutional layers L in the first half of the model, the original weight W of the model, and the weights W1,…,W corresponding to each convolutional layer from the 1st to the Lth layer in the model. L , and the subset W with larger weights top .

[0138] Table 1

[0139]

[0140]

[0141] Based on the weight mean m and weight standard deviation s of the 1-p% part of these convolution kernel weights, random noise η = N (m, k·s 2 ), where k is the noise amplification factor, the original weight w i Replace with w i ′=-w i +η i , ensuring that the weights follow a similar distribution while making the inference performance of the confused model worse.

[0142] The weight modification scheme that combines sign reversal with noise perturbation aims to significantly reduce the inference performance of the convolutional neural network and ensure its ∈-indistinguishability. The noise amplification factor k is an important parameter that controls the size of the noise variance and thus determines the degree of impact of the modification on the weights. A reasonable k value should ensure that the model performance is significantly reduced while not significantly deviating from the original weight distribution.

[0143] (3) Confusion model reasoning and correction

[0144] In order to reduce the computational pressure of the secure world and optimize the reasoning performance, this method transfers most of the convolution operations and the calculation of the fully connected layer to the normal world (enriched execution environment). In addition, in order to prevent attackers from using input data and intermediate results (IR) to infer the value of the model weights, it is necessary to add random noise before copying the IR to the untrusted normal world to hide the true value of the intermediate result. The following introduces a three-stage correction framework that can effectively protect the reasoning process and correct the prediction results.

[0145] In the first stage, for input data that has not been obfuscated, the correction parameter r in the secure memory is used to implement the intermediate result correction. In the second stage, when the input data has not been obfuscated, a one-time random value e is added to the intermediate result IR to hide the true value of the weight, making it sufficiently uncertain to prevent attackers from inferring the weight through the IR. In the third stage, for the obfuscated IR, e and the pre-calculated u are used for denoising correction.

[0146] Take a CNN with two convolutional layers as an example ( Figure 8), the convolution calculation is performed in the ordinary world (Rich Execution Environment REE), and the intermediate result correction is completed in the secure world (Trusted Execution Environment TEE). Through the interaction between the two, the secure reasoning of the model is completed. This method can be extended to any CNN model structure, and even the fully connected layer can adopt a similar noise addition and correction strategy. By adding random noise in the reasoning stage and correcting it when necessary, the true model weights are effectively hidden while ensuring the accuracy of model reasoning. Even if attackers can obtain intermediate results, they still cannot infer the model weights from them due to the presence of random noise, thereby ensuring the security of reasoning. This method ensures the efficiency and security of reasoning. Even in an environment with limited memory and computing resources, the security mechanism of TrustZone can still ensure data privacy and model integrity. The following is a detailed introduction to the secure reasoning correction scheme.

[0147] Intermediate result correction (first stage). When only the model weights are obfuscated, the obfuscated weights W constructed by the client application in the ordinary world l '=W l +r l and the input X of the lth layer l Perform convolution operation between them to get the obfuscated output result Y l '=X l ·(W l +r l )+bias l , and correct the IR in the secure world. l and the confusion result Y l 'Pass it to the secure world and then use the correction parameter r l Perform correction and obtain the correction result X l+1 =Y l '-X l ·r l =X l ·(W l +r l )+bias l -X l ·r l =X l ·W l +bias l .

[0148] Intermediate result protection (second stage). If the normal world obtains the input and the corrected non-zero intermediate result, the attacker can infer the value of the weights (solve the linear equation) through sufficient queries, so the non-zero IR should be protected before being copied to the untrusted normal world by l) to add random noise. If the input of the l-1th layer is not confused, its corrected output X l , a one-time random value e is added to IR. This scheme supports any activation function that can capture nonlinear properties and improve the accuracy of the model. It is represented as e l The random noise is added to the intermediate results. l is treated as confidential data and stored privately in secure storage. The output of the intermediate result protection process is X′ l =f(X l )+e l This process is only used for convolution operations with unobfuscated inputs, because if there is only obfuscated input, even if the correct intermediate result is obtained, the value of the weight cannot be solved.

[0149] Intermediate result correction (third stage). When both weights and inputs are obfuscated, the obfuscated weight W constructed by the client application in the ordinary world l ′=W l +r l and the input X′ of the lth layer l =(X l +e l ) to perform convolution operation and obtain the obfuscated output result Y l '=X′ l ·W l '+bias l =(X l +e l )·(W l +r l )+bias l , and correct the IR in the secure world. l and the confusion result Y l 'Pass it to the secure world and then use the correction parameter r l and e l Correction is performed and the result is X l+1 =X l ·W l +bias l =(X l +e l )·(W l +r l )+bias l -e l ·(W l +r l )-(X l +e l -e l )·r l =Y l '-ul -(X′ l -e l )·r l , where u l =e l ·(W l +r l ). l The calculation of u is independent of the input data, so u can be pre-calculated l To save calibration time in the secure world.

[0150] The above reasoning correction process continues until no more obfuscated weights or inputs are involved. The remaining calculations (such as fully connected layers) will be partitioned through the model to fully utilize the resources of the rich execution environment, thereby completing the remaining reasoning process safely and efficiently.

[0151] (4) Model partition execution

[0152] Different layers of the model have different memories of input data. Studies have shown that the front-layer features close to the input have stronger transferability on new datasets, while the back-layer features are more focused on specific tasks. This means that the front layers are able to capture more general information (such as the ambient tones in an image), while the back layers memorize more detailed task features (such as facial features). Therefore, when the model is accessed by an untrusted third party, the memory information in the model may be used to infer sensitive features in the reasoning data, resulting in privacy leakage risks.

[0153] Putting the last few layers of the neural network model into TEE for model reasoning can effectively prevent MIA attacks, because the last layer has a higher probability of leaking private information about the reasoning data. post Without loss of generality, the layer first Layer to layer partition -1 layer performs obfuscation model reasoning in REE and intermediate result correction in TEE. CA in REE runs layer partition Layer to layer last -1 layer, the TA in the TEE runs the last layer during inference last .

[0154] Execute the model M by partitioning post , while ensuring security, making better use of the computing resources on the device side.

[0155] (5) Filter and return the results

[0156] After completing the calculations of all layers, TA performs top-k filtering on the output of the last layer to control the returned data and mitigate MIA attacks, thereby further preventing the leakage of member information. The filtered prediction results are returned to CA through shared memory as the final output for users to view.

[0157] That is to say, the application example of this application proposes a method for secure model reasoning for mobile and IoT devices. This method deploys a neural network model on edge devices (such as smart watches and bracelets) that support TrustZone technology, so that user privacy data does not leave the local device, and at the same time uses the secure hardware on the device to ensure that the proprietary model of the service provider is fully protected. Through the isolated environment and secure storage provided by TrustZone, the private weights of the model can be protected to avoid leakage to untrusted device users, thereby safeguarding the intellectual property rights of the model owner. In addition, by introducing a model partitioning strategy, member inference attacks (MIA) can be effectively prevented, further improving the privacy security of model inference data. The application example of this application focuses on the reasoning of convolutional neural network (CNN) models, which is a commonly used model type in many mainstream machine learning tasks.

[0158] The embodiment of the present application also provides an electronic device, which may include a processor, a memory, a receiver and a transmitter, wherein the processor is used to execute the convolutional neural network reasoning method based on the device-side trusted execution environment mentioned in the above embodiment, wherein the processor and the memory may be connected via a bus or other means, such as by bus connection. The receiver may be connected to the processor and the memory via wired or wireless means.

[0159] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0160] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the program instructions / modules corresponding to the convolutional neural network reasoning method based on the device-side trusted execution environment in the embodiment of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, the convolutional neural network reasoning method based on the device-side trusted execution environment in the above method embodiment is implemented.

[0161] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0162] The one or more modules are stored in the memory, and when executed by the processor, perform the convolutional neural network reasoning method based on the device-side trusted execution environment in the embodiment.

[0163] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit, which may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.

[0164] As an implementation method, the functions of the receiver and the transmitter in the present application can be considered to be implemented through a transceiver circuit or a dedicated chip for transceiver, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit or a general chip.

[0165] As another implementation method, it is possible to use a general-purpose computer to implement the server provided in the embodiment of the present application, that is, to store the program code for implementing the functions of the processor, receiver, and transmitter in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.

[0166] The embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the convolutional neural network reasoning method based on the device-side trusted execution environment are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0167] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the aforementioned convolutional neural network reasoning method based on a device-side trusted execution environment.

[0168] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0169] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0170] In the present application, features described and / or illustrated for one embodiment may be used in the same manner or in a similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.

[0171] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the embodiments of the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A convolutional neural network reasoning method based on a device-side trusted execution environment, characterized in that: The method is executed in a trusted execution environment running on a device, and the method includes: During the inference process, receiving intermediate data outputted by the current target convolutional layer in the first model corresponding to the convolutional neural network after the obfuscation calculation in the client application; wherein the first model includes multiple convolutional layers connected sequentially in the convolutional neural network, and the target convolutional layer represents any convolutional layer in the first model; Correcting the intermediate data according to the correction parameters corresponding to the target convolution layer to obtain corrected intermediate data corresponding to the target convolution layer, and determining whether the pre-acquired input data corresponding to the target convolution layer has been obfuscated, and if so, transmitting the corrected intermediate data corresponding to the target convolution layer to the client application, so that the client application determines whether the target convolution layer is the last convolution layer in the first model, and if so, the client application inputs the corrected intermediate data into the second model corresponding to the convolution neural network, and sends out the result data output by the second model, wherein the second model corresponding to the convolution neural network includes sequentially connected pooling layers and multiple fully connected layers except for each convolution layer and the last fully connected layer in the convolution neural network; Receive the result data sent by the client application, and input the result data into the last fully connected layer of the convolutional neural network pre-stored locally to obtain the inference result data corresponding to the output of the last fully connected layer.

2. The convolutional neural network reasoning method based on a device-side trusted execution environment according to claim 1, characterized in that: The first model after obfuscation calculation in the client application runs in the rich execution environment of the client application, so that the client application inputs the input data corresponding to the current target convolutional layer in the first model into the target convolutional layer in the rich execution environment, so that the target convolutional layer performs a convolution operation on the input data based on the obfuscation weight corresponding to the target convolutional layer and outputs the intermediate data corresponding to the target convolutional layer.

3. The convolutional neural network reasoning method based on a device-side trusted execution environment according to claim 1, characterized in that: Also includes: If it is determined that the pre-acquired input data corresponding to the target convolutional layer has not been obfuscated, random noise is added to the corrected intermediate data corresponding to the target convolutional layer to complete the obfuscation of the corrected intermediate data; The corrected intermediate data after the obfuscation process is transmitted to the client application.

4. The convolutional neural network reasoning method based on a device-side trusted execution environment according to claim 1, characterized in that: If the client application determines that the target convolution layer is not the last convolution layer in the first model, the client application uses the corrected intermediate data corresponding to the target convolution layer as input data corresponding to a next convolution layer corresponding to the target convolution layer, and inputs the input data corresponding to the next convolution layer into the next convolution layer, so that the next convolution layer performs a convolution operation on the input data corresponding to the next convolution layer based on the confusion weight corresponding to the next convolution layer, and then outputs the intermediate data corresponding to the next convolution layer; Correspondingly, after transmitting the corrected intermediate data corresponding to the target convolutional layer to the client application and before receiving the result data sent by the client application, the method further includes: During the inference process, receiving intermediate data outputted by a next convolutional layer corresponding to the target convolutional layer in the first model after the obfuscation calculation; The intermediate data corresponding to the next convolution layer is corrected according to the correction parameters corresponding to the next convolution layer to obtain the corrected intermediate data corresponding to the next convolution layer, and it is determined whether the pre-acquired input data corresponding to the next convolution layer has been obfuscated. If so, the corrected intermediate data corresponding to the next convolution layer is transmitted to the client application.

5. The convolutional neural network reasoning method based on a device-side trusted execution environment according to claim 1, characterized in that: The first model and the second model corresponding to the convolutional neural network constitute the first half model of the convolutional neural network, and the last fully connected layer corresponding to the convolutional neural network constitutes the second half model of the convolutional neural network; and the first model after the obfuscation calculation is generated by performing obfuscation calculation on the first model based on a preset obfuscation model generation algorithm in advance, and the first model after the obfuscation calculation includes obfuscation weights and correction parameters corresponding to each convolution layer; The first model and the second model after obfuscation calculation are sent to the client application in advance for storage.

6. The convolutional neural network reasoning method based on a device-side trusted execution environment according to claim 5, characterized in that: Before receiving the intermediate data of the current target convolution layer output in the first model after the obfuscation calculation corresponding to the convolutional neural network in the client application during the inference process, it also includes: Receive and locally store the correction parameters corresponding to the last fully connected layer in the convolutional neural network and each convolution layer in the first model after confusion calculation.

7. A convolutional neural network reasoning device based on a device-side trusted execution environment, characterized in that: The device is set in a trusted execution environment running on the device side, including: An intermediate data receiving module, used to receive the intermediate data output by the current target convolution layer in the first model after the obfuscation calculation corresponding to the convolutional neural network in the client application during the inference process; wherein the first model includes multiple convolutional layers connected in sequence in the convolutional neural network, and the target convolutional layer represents any convolutional layer in the first model; A correction and data transmission module, used to correct the intermediate data according to the correction parameters corresponding to the target convolution layer to obtain corrected intermediate data, and determine whether the pre-acquired input data corresponding to the target convolution layer has been obfuscated. If so, the corrected intermediate data is transmitted to the client application, so that the client application determines whether the target convolution layer is the last convolution layer in the first model. If so, the client application inputs the corrected intermediate data into the second model corresponding to the convolutional neural network and sends out the result data output by the second model, wherein the second model corresponding to the convolutional neural network includes sequentially connected pooling layers and multiple fully connected layers in the convolutional neural network except for each convolution layer and the last fully connected layer; The hierarchical reasoning module is used to receive the result data sent by the client application, and input the result data into the last fully connected layer of the convolutional neural network pre-stored locally, to obtain the reasoning result data corresponding to the output of the last fully connected layer.

8. A convolutional neural network inference device, characterized in that: include: Client applications deployed in a rich execution environment and trusted applications deployed in a trusted execution environment; The trusted application is used to execute the convolutional neural network reasoning method based on the device-side trusted execution environment as described in any one of claims 1 to 6.

9. The convolutional neural network inference device according to claim 8, characterized in that: The client application is used to perform the following: Determining in the rich execution environment whether the target convolutional layer is the last convolutional layer in the first model; If yes, input the corrected intermediate data into the second model corresponding to the convolutional neural network, and send out the result data output by the second model; If not, the corrected intermediate data corresponding to the target convolution layer is used as the input data corresponding to the next convolution layer corresponding to the target convolution layer, and the input data corresponding to the next convolution layer is input into the next convolution layer, so that the next convolution layer performs a convolution operation on the input data corresponding to the next convolution layer based on the confusion weight corresponding to the next convolution layer, and then outputs the intermediate data corresponding to the next convolution layer, and sends the intermediate data corresponding to the next convolution layer to the trusted application.

10. The convolutional neural network inference device according to claim 9, characterized in that: Before determining in the rich execution environment whether the target convolutional layer is the last convolutional layer in the first model, the client application further executes the following: The first model and the second model after the obfuscated calculation are received and locally stored in the rich execution environment.

Citation Information

Patent Citations

  • Deep neural network reasoning method with privacy protection

    CN114003961A

  • Method and device for realizing reasoning of neural network model

    CN118194346A

  • Method for realizing safety model reasoning and related equipment

    CN118278522A

  • Scene-based dynamic security protection method, device and program product for data medium station

    CN118779893A

  • Quantitative neural network security reasoning method and device based on trusted execution environment

    CN118966352A