Neural network model reasoning method based on model compiling and NPU cooperation
By encrypting the neural network model weights during the model compilation and operation stage, the security risks faced by model weights after the deployment of the end-side SoC are solved, and the model is highly secure and stable.
Patent Information
- Application Number
- CN202510116325.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, after the weight parameters of the neural network algorithm model are deployed to the end-side SoC, they are easily acquired by malicious attackers through reverse analysis, resulting in the model facing security risks.
The neural network model trained by the deep learning framework is compiled and encrypted through the model compiler, and an encrypted model file is generated and stored in memory. The NPU loads the encrypted model file at runtime, decrypts the model weights and performs inferences to ensure that the model weights are encrypted during the compilation and operation stages.
It effectively avoids the risk of leakage of model weights during deployment and operation, enhances the security of the model, and reduces the security risks caused by attacks.
Smart Images

Figure CN120012836A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular to a neural network model reasoning method based on model compilation and NPU collaboration. Background Art
[0002] With the widespread application of deep learning technology, deep learning accelerator NPU plays an important role in accelerating the operation of neural network algorithm models. The existing mainstream architecture is to use deep learning frameworks for model training, and then deploy the trained neural network algorithm model to the end-side SoC (System on Chip). Common deep learning frameworks include pytorch, tensorflow, onnx, etc. Due to the model format compatibility problem between deep learning frameworks and end-side SoC, end-side SoC usually cannot directly run the model format trained by the deep learning framework. Instead, it is necessary to first use the model compilation tool to complete the compilation of the deep learning framework model format to the end-side SoC model format. The compiled model file is stored in the memory of the end-side SoC for the NPU inference integrated in the end-side SoC.
[0003] However, in this process, the protection of the weight parameters of the neural network algorithm model faces many difficulties. Malicious attackers can exploit loopholes in the existing protection mechanism and obtain the weight parameters of the neural network algorithm model by reverse analyzing the model files stored in the memory of the end-side SoC, thereby easily obtaining the customer's key algorithm model, causing the neural network algorithm model to face many security risks. Summary of the invention
[0004] In response to the above problems and technical requirements, this application proposes a neural network model reasoning method based on model compilation and NPU collaboration. The technical solution of this application is as follows:
[0005] A neural network model reasoning method based on model compilation and NPU collaboration, the neural network model reasoning method comprising:
[0006] After compiling the neural network model trained by the deep learning framework using the model compiler to obtain the model weight plaintext, the model weight plaintext is encrypted to obtain the model weight ciphertext, and an encrypted model file is generated and written into the memory; the generated encrypted model file includes an operator information data block and a weight parameter data block, the operator information data block carries the operator information, and the operator information indicates the weight storage information of the model weight ciphertext carried by the weight parameter data block in the memory;
[0007] When performing neural network model inference, the NPU is used to load the operator information data block in the encrypted model file into the NPU and parse it to obtain the operator information. The model weight ciphertext is read from the memory according to the operator information and loaded into the NPU for decryption to obtain the model weight plaintext before performing model inference.
[0008] The beneficial technical effects of this application are:
[0009] The present application discloses a neural network model reasoning method based on model compilation and NPU collaboration. After the neural network model is deployed to the end-side SoC, the method starts to intervene from the model compilation stage. The model compiler not only completes the model format conversion but also encrypts the model weight plaintext to obtain the model weight ciphertext, thereby generating an encrypted model file stored in the memory. Even if the encrypted model file is accidentally leaked, the model weight plaintext will not be leaked. In the running stage, the NPU completes the model reasoning after real-time decryption based on the configuration. By encrypting and decrypting the model weight plaintext throughout the entire process from compilation to reasoning, a solid security line of defense is built for the model data, which can effectively avoid the risk of model data leakage caused by various potential attack methods and protect the security of the neural network model in all directions.
[0010] In this method, the key used for NPU decryption is written into the key security storage area and NPU is only allowed to read it on demand. The model weight ciphertext obtained by NPU decryption is not written into the memory but cached in the NPU on-chip cache. By limiting CPU access, it can resist security attacks on memory analysis, thus having high security and greatly reducing security risks. Even in complex and changeable security threat environments, it can still ensure the high security of weight data when the model is running, providing a strong guarantee for the stable and safe functioning of the model.
[0011] This method carefully designs the collaborative operation process of the CPU and NPU. The CPU is responsible for parsing the header file data block to extract the decryption configuration information and configure the encryption and decryption control registers of the NPU. The NPU accurately loads and parses the relevant information in the encrypted model file based on these configurations to finally realize model reasoning. The entire process is closely linked to form a strict protection system, so that the weight data is in a safe and controllable state in each link, effectively avoiding the possibility of external attacks to obtain weight data, thereby effectively protecting the user's model data from leakage and effectively safeguarding the user's model intellectual property rights. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of the terminal-side SoC system architecture in one embodiment of the present application.
[0013] Figure 2 The figure is a schematic diagram of the model file structure of an encrypted model file in an embodiment.
[0014] Figure 3 It is a flowchart of a neural network model inference method in one embodiment of the present application.
[0015] Figure 4 It is a schematic diagram of the internal structure of the NPU in one embodiment of the present application. DETAILED DESCRIPTION
[0016] The specific implementation of the present application is further described below in conjunction with the accompanying drawings.
[0017] This application discloses a neural network model reasoning method based on model compilation and NPU collaboration. The method is applied to the end-side SoC. The end-side SoC integrates CPU, NPU, model compiler and memory. Please refer to Figure 1 The system architecture diagram is shown.
[0018] After the neural network model trained by the deep learning framework is deployed to the end-side SoC, the model compiler first compiles the neural network model and converts the neural network model into a model format supported by the end-side SoC, thereby obtaining the relevant operator weight parameters in the neural network model that can be accelerated on the NPU, which is referred to as model weight plaintext in this application. This application optimizes the compilation process performed by the model compiler. The model compiler performs encryption operations while completing the model compilation work, further encrypts the model weight plaintext to obtain the model weight ciphertext, and generates an encrypted model file based on the obtained model weight ciphertext and writes it into the memory of the end-side SoC.
[0019] Please refer to Figure 1 As shown in the schematic diagram, the generated encrypted model file includes at least an operator information data block and a weight parameter data block, and the weight parameter data block carries the compiled and encrypted model weight ciphertext. The operator information data block carries operator information, and the operator information indicates the weight storage information of the model weight ciphertext carried by the weight parameter data block in the memory.
[0020] The neural network model includes multiple operators, which are built into a hierarchical structure. The model weight plaintext of the entire neural network model actually includes the model weight plaintext of each operator. Correspondingly, the model weight ciphertext obtained by the model compiler from encrypting the model weight plaintext of the neural network model actually includes the model weight ciphertext of each operator. Therefore, the weight parameter data block in the encrypted model file includes N fields, each field carries the model weight ciphertext of an operator in the neural network model, and the integer parameter N ≥ 2, such as Figure 2 It is shown that the weight parameter data block includes model weight ciphertext 1 of operator 1, model weight ciphertext 2 of operator 2, ... model weight ciphertext N of operator N.
[0021] The operator information data block in the corresponding encrypted model file includes N operator information units, each of which carries the operator information of an operator. The operator information of each operator includes the weight storage information of the operator's model weight ciphertext in the memory. The weight storage information of each operator includes the weight storage address of the operator's model weight ciphertext in the memory and the weight data size of the model weight ciphertext. Figure 2 It is shown that the operator information data block includes operator information 1 of operator 1, operator information 2 of operator 2, ... operator information N of operator N. Taking operator 3 as an example, operator information 3 indicates the weight storage address of model weight ciphertext 3 of operator 3 in the memory and the data size of model weight ciphertext 1.
[0022] In addition, neural network models generally include multiple different types of operators. Different types of operators implement different functions. Therefore, the operator information of each operator also includes the operator type of the operator. Common operator types include convolution operators, pooling operators, relu operators, etc.
[0023] In addition, each operator in the neural network model is used to perform neural network calculations on its own input data. Therefore, the operator information of each operator also includes input storage information allocated to the input data of the operator. The input storage information allocated to each operator includes the data size of the input data of the operator and the input data storage address allocated to the input data.
[0024] The subsequent decryption operation performed by the NPU on the model weight ciphertext needs to correspond to the encryption processing performed by the model compiler on the model weight plaintext. In another embodiment, in order to adapt to the usage requirements in different scenarios, the encryption processing performed by the model compiler on the model weight plaintext is not always a fixed encryption method. In this embodiment, the encrypted model file also includes a header file data block, and the header file data block carries encryption and decryption configuration information. The encryption and decryption configuration information corresponds to the encryption processing operation performed by the model compiler on the model weight plaintext, so that the NPU can perform corresponding decryption operations based on the encryption and decryption configuration information to decrypt the model weight ciphertext. Among them, the decryption configuration information carried by the header file data block includes the encryption and decryption algorithm type and the encryption and decryption mode, and a form of symmetric algorithm is realized by the combination of the encryption and decryption algorithm type and the encryption and decryption mode. Among them, the encryption and decryption algorithm type is any one of the SM4 symmetric algorithm and the AES symmetric algorithm, and the encryption and decryption mode is any one of the ECB mode, the CBC mode, the CFB mode and the OFB mode.
[0025] Furthermore, the decryption configuration information carried by the header file data block also includes a decryption enable signal. When the decryption enable signal is valid, it indicates to start the model weight decryption operation, otherwise it indicates not to start the model weight decryption operation so as to be compatible with traditional methods.
[0026] The encrypted model file compiled by the model compiler carries the above-mentioned information. The specific data structure and the length and encoding form of each field can be pre-customized. This application does not limit this. After the data structure of the encrypted model file is determined in advance, the subsequent CPU and NPU can parse the data structure of the encrypted model file to obtain various information carried by the encrypted model file. The encrypted model file compiled by the model compiler carries the model weight ciphertext stored in the memory, so that even if the encrypted model file is accidentally leaked, it is difficult for the attacker to directly read the model weight plaintext. This is the first line of defense to ensure the security of the model data.
[0027] Then, when performing neural network model inference on the neural network model, the NPU loads the operator information data block in the encrypted model file into the NPU, and parses the operator information data block inside the NPU to obtain the operator information, and then reads the model weight ciphertext from the memory according to the operator information and loads it into the NPU to decrypt the model weight plaintext before performing model inference. In this process, the operator information data block is parsed within the NPU chip, and the model weight ciphertext is also decrypted and inferred in real time within the NPU chip, so that it is always in a relatively closed and secure environment during the model inference process. The outside world cannot spy on the weight data through memory analysis methods, which greatly improves the security of the weight data and can effectively resist various security attacks from memory analysis, ensuring the high security of the model during operation.
[0028] The key used by the NPU to decrypt the model weight ciphertext is stored in the key security storage area, and the NPU obtains the key by accessing the key security storage area. The key is not updated in real time, but the key and random number (which can be used as IV) are written to the key security storage area inside the chip through a proprietary key writing tool when the chip leaves the factory, and the CPU does not have access to the key security storage area, ensuring the security of the key, fundamentally guaranteeing the reliability of the decryption operation and the confidentiality of the model weights.
[0029] When the encrypted model file includes a header file data block that carries encryption and decryption configuration information, before using the NPU for neural network model inference, please refer to Figure 3 The flowchart shown in the figure also first uses the CPU to parse the header file data block in the encrypted model file to obtain the encryption and decryption configuration information, and writes the encryption and decryption configuration information into the encryption and decryption control register of the NPU to complete the decryption configuration of the NPU, so that the NPU uses the obtained key to decrypt the model weight ciphertext according to the encryption and decryption configuration information in the encryption and decryption control register to obtain the model weight plaintext:
[0030] When the NPU detects that the decryption enable signal stored in the encryption and decryption control register is valid, it enables the corresponding encryption and decryption algorithm according to the encryption and decryption algorithm type stored in the encryption and decryption control register, and enables the corresponding encryption and decryption mode according to the encryption and decryption mode stored in the encryption and decryption control register, and then uses the acquired key to decrypt the model weight ciphertext according to the encryption and decryption algorithm type in the encryption and decryption configuration information and the encryption and decryption mode in the encryption and decryption configuration information.
[0031] When the NPU detects that the decryption enable signal stored in the encryption and decryption control register is invalid, the NPU does not start the model weight decryption operation. At this time, it no longer reads other encryption and decryption configuration information stored in the encryption and decryption control register. It can be determined that the model weight plaintext is read from the memory and loaded into the NPU chip according to the operator information, and the model inference is performed directly.
[0032] Based on the feature that the operator information data block carries the operator information of multiple operators, the NPU can obtain the operator information of each operator by parsing the operator information data block. Then, when the NPU reads the model weight ciphertext from the memory according to the operator information and decrypts it, for any integer parameter 1≤n≤N, the model weight ciphertext of the nth operator is read from the memory according to the operator information of any nth operator and loaded into the NPU, and then the model weight ciphertext of the nth operator is decrypted to obtain the model weight plaintext of the nth operator.
[0033] When the NPU decrypts the model weight plaintext obtained by decryption, it obtains the input data of the nth operator according to the input storage information in the operator information of the nth operator, performs a neural network calculation (usually a convolution operation) on the input data of the nth operator and the model weight plaintext of the nth operator according to the calculation method corresponding to the operator type of the nth operator to obtain a calculation result, and writes the calculation result into the input storage information of the next layer of operators of the nth operator.
[0034] For the internal structure of NPU, please refer to Figure 4 , NPU includes a data loading unit, a parsing unit, a decryption unit, a parameter cache unit and a neural network acceleration unit, then the specific process of NPU executing the above method is:
[0035] (1) First, use the data loading unit to load the operator information data block into the NPU and transmit it to the parsing unit.
[0036] (2) Using the parsing unit to parse the operator information data block, the operator information of any operator is obtained, including the operator type, input storage information, and weight storage information.
[0037] (3) The parsing unit transmits the weight storage information to the data loading unit. The data loading unit reads the model weight ciphertext of the nth operator from the memory according to the weight storage information in the operator information of any nth operator and loads it into the NPU to enter the decryption unit.
[0038] (4) The decryption unit decrypts the model weight ciphertext of the nth operator according to the encryption and decryption configuration information stored in the encryption and decryption control register to obtain the model weight plaintext of the nth operator. In theory, the model weight plaintext of the nth operator should be directly provided to the neural network acceleration unit for neural network calculation, but the operators in the neural network model are not just simple serial structures, so the decryption reasoning of the operator is not executed serially, so the NPU also includes a parameter cache unit. After the decryption unit decrypts the model weight plaintext of the nth operator, it first writes it into the parameter cache unit in the NPU, and records the plaintext cache information of the model weight plaintext of the nth operator in the parameter cache unit and provides it to the neural network acceleration unit. The plaintext cache information includes the storage address and storage length of the model weight plaintext. The model weight plaintext can be read according to the plaintext cache information. The model weight plaintext obtained by decryption here is no longer written into the memory, but cached in the NPU chip, avoiding being read by the CPU, thereby avoiding the leakage of the model weight plaintext.
[0039] (5) The neural network acceleration unit obtains the model weight plaintext of the nth operator. In one case, the decryption unit decrypts the model weight plaintext of the nth operator and directly provides it to the neural network acceleration unit. In this case, the neural network acceleration unit directly receives the model weight plaintext of the nth operator transmitted by the decryption unit. In another case, the decryption unit decrypts the model weight plaintext of the nth operator and first writes it into the parameter cache unit inside the NPU. In this case, when the neural network acceleration unit performs model inference on the nth operator, it reads the model weight plaintext of the nth operator from the parameter cache unit according to the plaintext cache information of the nth operator.
[0040] The data loading unit is used to read the input data of the nth operator from the memory according to the input storage information in the operator information of the nth operator and provide it to the neural network acceleration unit. The neural network acceleration unit performs neural network calculation on the input data of the nth operator and the model weight plaintext according to the operator type of the nth operator parsed by the parsing unit to obtain the calculation result of the nth operator. The data loading unit is then used to write the calculation result of the nth operator into the memory according to the input storage information of the next layer of operators of the nth operator.
[0041] The above is only a preferred embodiment of the present application, and the present application is not limited to the above embodiments. It is understood that other improvements and changes directly derived or associated by those skilled in the art without departing from the spirit and concept of the present application should be considered to be included in the protection scope of the present application.
Claims
1. A neural network model inference method based on model compilation and NPU collaboration, characterized in that: The neural network model inference method comprises: After compiling the neural network model trained by the deep learning framework using a model compiler to obtain a model weight plaintext, the model weight plaintext is encrypted to obtain a model weight ciphertext, and an encrypted model file is generated and written into the memory; the generated encrypted model file includes an operator information data block and a weight parameter data block, the operator information data block carries operator information, and the operator information indicates the weight storage information of the model weight ciphertext carried by the weight parameter data block in the memory; When performing neural network model inference, the NPU is used to load the operator information data block in the encrypted model file into the NPU and parse it to obtain the operator information. According to the operator information, the model weight ciphertext is read from the memory and loaded into the NPU for decryption to obtain the model weight plaintext before performing model inference.
2. The neural network model inference method according to claim 1, characterized in that: The generated encrypted model file also includes a header file data block, the header file data block carries encryption and decryption configuration information, and the encryption and decryption configuration information corresponds to the encryption processing operation performed by the model compiler on the model weight plaintext; The neural network model reasoning method also includes: The CPU is used to parse the header file data block in the encrypted model file to obtain the encryption and decryption configuration information, and the encryption and decryption configuration information is written into the encryption and decryption control register of the NPU to complete the decryption configuration of the NPU, so that the NPU decrypts the model weight ciphertext according to the encryption and decryption configuration information in the encryption and decryption control register.
3. The neural network model inference method according to claim 2, characterized in that: The NPU decrypts the model weight ciphertext according to the encryption and decryption configuration information in the encryption and decryption control register, including: The NPU accesses the key security storage area to obtain the key, and uses the key to decrypt the model weight ciphertext according to the encryption and decryption configuration information in the encryption and decryption control register; wherein the CPU does not have access rights to the key security storage area.
4. The neural network model inference method according to claim 1, characterized in that: The weight parameter data block in the encrypted model file includes N fields, each field carries the model weight ciphertext of an operator in the neural network model; the operator information data block in the encrypted model file includes N operator information units, each operator information unit carries the operator information of an operator, and the operator information of each operator includes the weight storage information of the model weight ciphertext of the operator in the memory; wherein the integer parameter N≥2; According to the operator information, the model weight ciphertext is read from the memory and loaded into the NPU for decryption to obtain the model weight plaintext, including: According to the operator information of any nth operator, the model weight ciphertext of the nth operator is read from the memory and loaded into the NPU, and then decrypted to obtain the model weight plaintext of the nth operator, with an integer parameter of 1≤n≤N.
5. The neural network model inference method according to claim 4, characterized in that: The operator information of each operator also includes the operator type of the operator and the input storage information allocated to the input data of the operator. The model reasoning using NPU also includes: The input data of the nth operator is obtained according to the input storage information in the operator information of the nth operator, and a neural network calculation is performed on the input data of the nth operator and the model weight plaintext of the nth operator according to the calculation method corresponding to the operator type of the nth operator to obtain a calculation result, and the calculation result is written into the input storage information of the next layer operator of the nth operator.
6. The neural network model inference method according to claim 5, characterized in that: The NPU includes a data loading unit, a parsing unit, a decryption unit, and a neural network acceleration unit. The method executed by the NPU includes: Use the data loading unit to load the operator information data block into the NPU; Utilizing the parsing unit to parse the operator information data block to obtain the operator information of any individual operator; The data loading unit is used to read the model weight ciphertext of the nth operator from the memory according to the weight storage information in the operator information of any nth operator and load it into the NPU to enter the decryption unit, with an integer parameter of 1≤n≤N; Decrypting the model weight ciphertext of the nth operator according to the information in the encryption and decryption control register using the decryption unit to obtain the model weight plaintext of the nth operator; The data loading unit is used to read the input data of the nth operator from the memory according to the input storage information in the operator information of the nth operator and provide it to the neural network acceleration unit. The neural network acceleration unit is used to perform neural network calculation on the input data of the nth operator and the model weight plaintext according to the operator type of the nth operator obtained by parsing the parsing unit to obtain a calculation result.
7. The neural network model inference method according to claim 6, characterized in that: The NPU also includes a parameter cache unit, and the method executed by the NPU also includes: The model weight plaintext of the nth operator obtained by the decryption unit is written into the parameter cache unit in the NPU, and the plaintext cache information of the model weight plaintext of the nth operator in the parameter cache unit is provided to the neural network acceleration unit; when the neural network acceleration unit uses the neural network acceleration unit to perform model inference on the nth operator, the model weight plaintext of the nth operator is read from the parameter cache unit according to the plaintext cache information of the nth operator.
8. The neural network model inference method according to claim 2, characterized in that: The decryption configuration information carried by the header file data block in the encrypted model file includes the encryption and decryption algorithm type and the encryption and decryption mode; decrypting the model weight ciphertext according to the encryption and decryption configuration information in the encryption and decryption control register includes: decrypting the model weight ciphertext according to the encryption and decryption algorithm type in the encryption and decryption configuration information using the encryption and decryption mode in the encryption and decryption configuration information.
9. The neural network model inference method according to claim 8, characterized in that: The decryption configuration information carried by the header file data block in the encrypted model file also includes a decryption enable signal, and the neural network model inference method also includes: When the NPU detects that the decryption enable signal stored in the encryption and decryption control register is valid, it reads the model weight ciphertext from the memory according to the operator information, loads it into the NPU, decrypts it to obtain the model weight plaintext, and then performs model inference; When the NPU detects that the decryption enable signal stored in the encryption and decryption control register is invalid, it reads the model weight plaintext directly from the memory according to the operator information and loads it into the NPU for model inference.
10. The neural network model inference method according to claim 8, characterized in that: The encryption and decryption algorithm type is any one of the SM4 symmetric algorithm and the AES symmetric algorithm, and the encryption and decryption mode is any one of the ECB mode, the CBC mode, the CFB mode, and the OFB mode.
Citation Information
Cited By
Neural network model reasoning method for improving intermediate data security
CN121168542A