Large model hierarchical encryption deployment method based on sensitivity perception
Through the sensitivity-aware layered encryption deployment of large language models, sensitive data is routed to the trusted execution environment TEE for encryption calculation, combined with GPU cluster processing low-sensitivity layer, the data privacy and inference security issues of large language models in highly sensitive industries are solved, and data protection and performance optimization are achieved.
Patent Information
- Application Number
- CN202510353963.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
The existing large language models have data privacy and reasoning security issues in the deployment of highly sensitive industries, especially in the fields of medical care, finance, government affairs, etc., and ordinary enterprises find it difficult to train models that meet their needs.
Using a hierarchical encryption deployment method with sensitivity perception, the Transformer Block structure of the large model is divided into high-sensitive layer and low-sensitive layer, and the trusted execution environment TEE is used to perform encryption calculations of sensitive data, combined with GPU cluster processing low-sensitive layer calculations, and ensure data security through fully homomorphic encryption and lightweight symmetric encryption algorithms.
It realizes that while protecting data privacy, it reduces inference performance losses, improves data security and computing efficiency.
Smart Images

Figure CN120337242A_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a hierarchical encryption deployment method for large models based on sensitivity perception, which relates to the field of data security of large language models. Background Art
[0002] Industries in all walks of life are trying to deploy large model applications in this field. Due to the huge scale of large language models, it is difficult for ordinary enterprises to train a large model with performance meeting requirements. Therefore, the common practice is to use open-source large language models and apply them to industrial enterprises through means such as fine-tuning. However, due to the open-source nature of the models, there are bound to be security problems. Especially in the process of accelerating the empowerment of high-sensitivity industries such as healthcare, finance, and government by large language models (LLMs), there are still data privacy and inference security problems. Summary of the Invention
[0003] Aiming at the problems of the prior art, the present invention provides a hierarchical encryption deployment method for large models based on sensitivity perception. Based on the principle of separating sensitive data from computing, by setting up a trusted execution environment security TEE island and adopting a method of dynamically hierarchical encryption of the model, data and model layers with higher sensitivity are calculated separately, so as to achieve the purpose of protecting data privacy and reducing the loss of inference performance.
[0004] The specific solution proposed by the present invention is as follows:
[0005] The present invention provides a hierarchical encryption deployment method for large models based on sensitivity perception, including:
[0006] Step 1: For the Transformer Block structure of the network of the large model, divide the high-sensitivity layer and the low-sensitivity layer, where the Transformer Block structure that is easy to reverse infer the true input is the high-sensitivity layer, and the Transformer Block structure that is not easy to reverse infer the true input is the low-sensitivity layer.
[0007] Step 2: For the request data input by the user, perform sensitive word detection. If there are sensitive words, the part of the request data containing sensitive words is directed to the trusted execution environment TEE of the large model for encrypted calculation of the high-sensitivity layer; the remaining part of the request data is diverted to the GPU cluster for calculation of the low-sensitivity layer.
[0008] When performing encrypted calculations on the highly sensitive layer, first determine whether the memory resources of the trusted execution environment (TEE) are greater than the parameter calculation amount of the highly sensitive layer. If so, directly route the highly sensitive layer to the TEE for encrypted calculation. Otherwise, calculate the maximum tolerable parameter calculation amount based on the memory capacity of the TEE, divide the parameter calculation amount of the highly sensitive layer into several blocks according to the maximum tolerable parameter calculation amount, perform encrypted calculations on each block in the TEE, and encrypt and output the results to the GPU cluster. The GPU cluster calculates the blocks transmitted by the TEE and returns the calculation results to the TEE for decryption to obtain the plaintext results.
[0009] Further, in step 1 of the method for hierarchical encrypted deployment of large models based on sensitivity perception, dividing the highly sensitive layer and the low sensitive layer includes: inputting test data, obtaining the output corresponding to each Transformer Block layer, respectively inputting the output of each layer into the model inversion attack simulator, allowing the attack simulator to reverse-infer the real input, evaluating the error between the input reverse-inferred by the attack simulator and the real input. If the error is lower than the set threshold, the corresponding Transformer Block structure is the highly sensitive layer; if the error is higher than the set threshold, the corresponding Transformer Block structure is the low sensitive layer.
[0010] Further, in step 2 of the method for hierarchical encrypted deployment of large models based on sensitivity perception, the encrypted calculation by the TEE includes:
[0011] Encrypt the parameters and input data of the highly sensitive layer using the pre-generated fully homomorphic encryption public key to ensure that the data is executed in the encrypted state during TEE calculation.
[0012] After the TEE completes the encrypted calculation, generate the intermediate result in ciphertext form.
[0013] Divide the intermediate result into blocks by rows, and attach a message authentication code to each data block.
[0014] Use the lightweight symmetric encryption algorithm AES-GCM to perform secondary encryption on the data blocks. The key is dynamically generated by the TEE and shared with the GPU cluster through a secure channel.
[0015] Further, in step 2 of the method for hierarchical encrypted deployment of large models based on sensitivity perception, calculating the blocks transmitted by the TEE in the GPU cluster and returning the calculation results to the TEE for decryption includes:
[0016] After the GPU cluster receives the encrypted data blocks, first decrypt the AES-GCM layer with the symmetric key, and retain the ciphertext state of the fully homomorphic encryption.
[0017] After the GPU completes the calculation of the encrypted data blocks, re-encrypt the output result with AES-GCM and return it to the TEE through a secure channel.
[0018] After the TEE receives the result, it first decrypts the AES-GCM layer, and then uses the private key of fully homomorphic encryption to finally decrypt the FHE ciphertext to recover the plaintext result.
[0019] Step 2 can use symmetric encryption algorithm AES, asymmetric encryption algorithm RSA, hash algorithm SHA, and fully homomorphic encryption algorithm for encryption calculation.
[0020] The present invention also provides a large model hierarchical encryption deployment device based on sensitivity perception, including a structure division module, a request parsing module, a secure computing routing module, and a distributed execution module.
[0021] The structure division module divides the high-sensitivity layer and the low-sensitivity layer for the Transformer Block structure of the large model network, where the Transformer Block structure that is easily reverse-inferred to the real input is the high-sensitivity layer, and the Transformer Block structure that is not easily reverse-inferred to the real input is the low-sensitivity layer.
[0022] The request parsing module performs sensitive word detection on the request data input by the user. If there are sensitive words, the secure computing routing module will direct the part of the request data containing sensitive words to the trusted execution environment TEE of the large model for high-sensitivity layer encryption calculation; the remaining part of the request data is split to the GPU cluster for low-sensitivity layer calculation.
[0023] When the distributed execution module performs high-sensitivity layer encryption calculation, it first judges whether the memory resource of the trusted execution environment TEE is greater than the parameter calculation amount of the high-sensitivity layer. If so, it directly routes the high-sensitivity layer to the TEE for encryption calculation. Otherwise, according to the memory capacity of the TEE, it calculates the maximum parameter calculation amount that can be accommodated, divides the parameter calculation amount of the high-sensitivity layer into several blocks according to the maximum parameter calculation amount that can be accommodated, performs encryption calculation on each block in the TEE, encrypts and outputs the result to the GPU cluster, calculates the blocks transmitted by the TEE in the GPU cluster, and returns the calculation result to the TEE for decryption to obtain the plaintext result.
[0024] Further, the structure division module of the large model hierarchical encryption deployment device based on sensitivity perception divides the high-sensitivity layer and the low-sensitivity layer, including: inputting test data, obtaining the output corresponding to each layer of the Transformer Block, respectively inputting the output of each layer into the model inversion attack simulator, allowing the attack simulator to reverse-infer the real input, evaluating the error between the input reverse-inferred by the attack simulator and the real input. If the error is lower than the set threshold, the corresponding Transformer Block structure is the high-sensitivity layer. If the error is higher than the set threshold, the corresponding Transformer Block structure is the low-sensitivity layer.
[0025] Furthermore, the distributed execution module of the large model hierarchical encryption deployment device based on sensitivity perception performs encrypted calculations in the TEE, including:
[0026] Encrypt the parameters and input data of the highly sensitive layer using the pre-generated fully homomorphic encryption public key to ensure that the data performs TEE calculations in an encrypted state.
[0027] After the encrypted calculation is completed in the TEE, generate the intermediate result in ciphertext form.
[0028] Divide the intermediate result into blocks by rows, and attach a message authentication code to each data block.
[0029] Use the lightweight symmetric encryption algorithm AES-GCM to perform secondary encryption on the data blocks. The key is dynamically generated by the TEE and shared with the GPU cluster through a secure channel.
[0030] Furthermore, the distributed execution module of the large model hierarchical encryption deployment device based on sensitivity perception calculates the blocks transmitted by the TEE in the GPU cluster and returns the calculation result to the TEE for decryption, including:
[0031] After the GPU cluster receives the encrypted data block, the distributed execution module first decrypts the AES-GCM layer with the symmetric key and retains the ciphertext state of the fully homomorphic encryption.
[0032] After the distributed execution module completes the calculation of the encrypted data block in the GPU, it re-encrypts the output result with AES-GCM and returns it to the TEE through a secure channel.
[0033] After the TEE receives the result, the distributed execution module first decrypts the AES-GCM layer, and then uses the private key of the fully homomorphic encryption to finally decrypt the FHE ciphertext to restore the plaintext result.
[0034] The beneficial effects of the present invention are:
[0035] Based on the principle of separation of sensitive data and calculation, the present invention sets up a trusted execution environment TEE, and uses the method of dynamic hierarchical encryption to calculate the data in layers. Sensitive data is routed to the trusted execution environment TEE for calculation in the highly sensitive layer, and the low sensitive layer is calculated in the GPU cluster, which not only achieves the purpose of protecting data privacy, but also reduces the loss of inference performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic diagram of the method flow of the present invention.
[0037] Figure 2 is a schematic diagram of the interaction between modules in the device of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0038] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited do not limit the present invention.
[0039] Embodiment 1
[0040] The present invention provides a method for hierarchical encryption deployment of a large model based on sensitivity perception, including:
[0041] Step 1: For the Transformer Block structure of the large model network, divide the high-sensitivity layer and the low-sensitivity layer, where the Transformer Block structure that is easily reverse-engineered to obtain the real input is the high-sensitivity layer, and the Transformer Block structure that is not easily reverse-engineered to obtain the real input is the low-sensitivity layer.
[0042] When dividing the high-sensitivity layer and the low-sensitivity layer, it may specifically include: inputting test data, obtaining the output corresponding to each layer of the Transformer Block, respectively inputting the output of each layer into the model inversion attack simulator, allowing the attack simulator to reverse-engineer the real input, evaluating the error between the input reverse-engineered by the attack simulator and the real input. If the error is lower than the set threshold, the corresponding Transformer Block structure is the high-sensitivity layer; if the error is higher than the set threshold, the corresponding Transformer Block structure is the low-sensitivity layer.
[0043] Step 2: For the request data input by the user, perform sensitive word detection. If there are sensitive words, direct the part of the request data containing the sensitive words to the trusted execution environment TEE of the large model for encrypted calculation of the high-sensitivity layer; shunt the remaining part of the request data to the GPU cluster for calculation of the low-sensitivity layer; when inferring the input data without sensitive words, both the high-sensitivity layer and the low-sensitivity layer use plaintext calculation.
[0044] When performing encrypted calculation of the high-sensitivity layer, first determine whether the memory resources of the trusted execution environment TEE are greater than the parameter calculation amount of the high-sensitivity layer. If so, directly route the high-sensitivity layer to the TEE for encrypted calculation. Otherwise, according to the memory capacity of the TEE, calculate the maximum parameter calculation amount that can be accommodated, divide the parameter calculation amount of the high-sensitivity layer into several blocks according to the maximum parameter calculation amount that can be accommodated, perform encrypted calculation on each block in the TEE, and encrypt and output the result to the GPU cluster. Calculate the blocks transmitted by the TEE in the GPU cluster, and return the calculation result to the TEE for decryption to obtain the plaintext result.
[0045] The parameter calculation amount of the highly sensitive layer can be regarded as the overall weight parameter calculation amount. The overall weight parameters are divided into several parts, and each part can be called a weight block. During the inference stage of the large model, the calculation process can be regarded as the multiplication and addition operations of matrices, that is, the operation matrix is sliced into several small matrices for operation. When calculating the maximum allowable parameter calculation amount, that is, the amount of data that can be loaded into the TEE memory capacity, it is equivalent to performing weight block calculations in the TEE. In a computer, it is calculated in bytes. For example, if the weight is a float32 floating point number, each value occupies 4B of storage. Assuming the TEE memory has 2GB, the weight data amount that can be accommodated is 2 * 1024 * 1024 / 4 = 524288, that is, 524288 weight parameter values can be loaded in. When dividing the matrix during operation, it is required that the product of the row and column dimensions of the divided matrix does not exceed this value.
[0046] In step 2, the TEE performs encrypted calculations, which may include:
[0047] Encrypt the parameters of the highly sensitive layer and the input data using the pre-generated fully homomorphic encryption public key to ensure that the TEE calculation is performed in an encrypted state.
[0048] After the TEE completes the encrypted calculation, it generates an intermediate result in ciphertext form.
[0049] Divide the intermediate result into blocks by rows, and attach a message authentication code to each data block.
[0050] Use the lightweight symmetric encryption algorithm AES-GCM to perform secondary encryption on the data blocks. The key is dynamically generated by the TEE and shared with the GPU cluster through a secure channel.
[0051] At the same time, the GPU cluster calculates the blocks transmitted by the TEE and returns the calculation result to the TEE for decryption, including:
[0052] After the GPU cluster receives the encrypted data block, it first decrypts the AES-GCM layer with the symmetric key and retains the ciphertext state of the fully homomorphic encryption.
[0053] After the GPU completes the calculation of the encrypted data block, it re-encrypts the output result with AES-GCM and returns it to the TEE through a secure channel.
[0054] After the TEE receives the result, it first decrypts the AES-GCM layer, and then uses the private key of the fully homomorphic encryption to finally decrypt the FHE ciphertext to restore the plaintext result.
[0055] Step 2 can use symmetric encryption algorithms such as AES, asymmetric encryption algorithms such as RSA, hash algorithms such as SHA, and fully homomorphic encryption algorithms for encrypted calculations.
[0056] Embodiment 2
[0057] The present invention also provides a large model hierarchical encryption deployment device based on sensitivity perception, including a structure division module, a request parsing module, a secure computing routing module, and a distributed execution module.
[0058] The structure division module divides the Transformer Block structure of the large model network into a high-sensitivity layer and a low-sensitivity layer. Among them, the Transformer Block structure that is easily reverse-engineered to obtain the real input is the high-sensitivity layer, and the Transformer Block structure that is not easily reverse-engineered to obtain the real input is the low-sensitivity layer.
[0059] The request parsing module performs sensitive word detection on the request data input by the user. If there are sensitive words, the secure computing routing module directs the part of the request data containing sensitive words to the trusted execution environment TEE of the large model for encrypted calculation of the high-sensitivity layer; the remaining part of the request data is split to the GPU cluster for calculation of the low-sensitivity layer.
[0060] When the distributed execution module performs encrypted calculation of the high-sensitivity layer, it first determines whether the memory resources of the trusted execution environment TEE are greater than the parameter calculation amount of the high-sensitivity layer. If so, it directly routes the high-sensitivity layer to the TEE for encrypted calculation; otherwise, it calculates the maximum parameter calculation amount that can be accommodated according to the memory capacity of the TEE, divides the parameter calculation amount of the high-sensitivity layer into several blocks according to the maximum parameter calculation amount that can be accommodated, performs encrypted calculation on each block in the TEE, encrypts and outputs the results to the GPU cluster, calculates the blocks transmitted by the TEE in the GPU cluster, and returns the calculation results to the TEE for decryption to obtain the plaintext results.
[0061] For the information interaction and execution process memory among the above-mentioned modules in the device, since it is based on the same concept as the method embodiment of the present invention, the specific memory can be referred to the description in the method embodiment of the present invention, and will not be elaborated here.
[0062] Similarly, based on the principle of separation of sensitive data and calculation, the device of the present invention sets up a trusted execution environment TEE, adopts a dynamic hierarchical encryption method to calculate data in layers respectively. Sensitive data is routed to the trusted execution environment TEE for calculation in the high-sensitivity layer, and the low-sensitivity layer is calculated in the GPU cluster, which not only achieves the purpose of protecting data privacy but also reduces the loss of inference performance.
[0063] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as required. The system structures described in the above-mentioned embodiments can be physical structures or logical structures, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities respectively, or they can be jointly implemented by some components in multiple independent devices.
[0064] The above-mentioned embodiments are only preferred embodiments cited to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.
Claims
1. A method for hierarchical encryption deployment of large models based on sensitivity perception, characterized in that Including: Step 1: For the Transformer Block structure of the large model network, divide the highly sensitive layer and the low sensitive layer. Among them, the Transformer Block structure that is easily reverse-engineered to obtain the real input is the highly sensitive layer, and the Transformer Block structure that is not easily reverse-engineered to obtain the real input is the low sensitive layer. Step 2: For the request data input by the user, perform sensitive word detection. If there are sensitive words, direct the part of the request data containing sensitive words to the trusted execution environment TEE of the large model for encrypted calculation of the highly sensitive layer; divert the remaining part of the request data to the GPU cluster for calculation of the low sensitive layer. When performing encrypted calculation of the highly sensitive layer, first determine whether the memory resources of the trusted execution environment TEE are greater than the parameter calculation amount of the highly sensitive layer. If so, directly route the highly sensitive layer to the TEE for encrypted calculation. Otherwise, calculate the maximum parameter calculation amount that can be accommodated according to the memory capacity of the TEE, divide the parameter calculation amount of the highly sensitive layer into several blocks according to the maximum parameter calculation amount that can be accommodated, perform encrypted calculation on each block in the TEE, and encrypt and output the result to the GPU cluster. Calculate the blocks transmitted by the TEE in the GPU cluster and return the calculation result to the TEE for decryption to obtain the plaintext result.
2. The method for hierarchical encryption deployment of a large model based on sensitivity perception according to claim 1, characterized in that In Step 1, dividing the highly sensitive layer and the low sensitive layer includes: inputting test data, obtaining the output corresponding to each layer of the Transformer Block, inputting the output of each layer into the model inversion attack simulator respectively, allowing the attack simulator to reverse-engineer the real input, and evaluating the error between the input reverse-engineered by the attack simulator and the real input. If the error is lower than the set threshold, the corresponding Transformer Block structure is the highly sensitive layer; if the error is higher than the set threshold, the corresponding Transformer Block structure is the low sensitive layer.
3. A method for hierarchical encryption deployment of a large model based on sensitivity perception according to claim 1, characterized in that In Step 2, the encrypted calculation performed by the TEE includes: Encrypting the parameters and input data of the highly sensitive layer using the pre-generated fully homomorphic encryption public key to ensure that the data is executed in the encrypted state in the TEE calculation. After the TEE completes the encrypted calculation, generate the intermediate result in ciphertext form. Divide the intermediate result into blocks by rows, and attach a message authentication code to each data block. Use the lightweight symmetric encryption algorithm AES-GCM to perform secondary encryption on the data blocks. The key is dynamically generated by the TEE and shared with the GPU cluster through a secure channel.
4. The method for hierarchical encryption deployment of a large model based on sensitivity perception according to claim 3, characterized in that In Step 2, calculating the blocks transmitted by the TEE in the GPU cluster and returning the calculation result to the TEE for decryption includes: After receiving the encrypted data block, the GPU cluster first decrypts the AES-GCM layer with the symmetric key and retains the ciphertext state of the fully homomorphic encryption. After the GPU completes the calculation of the encrypted data block, re-encrypt the output result with AES-GCM and return it to the TEE through a secure channel. After receiving the result, the TEE first decrypts the AES-GCM layer, and then uses the private key of the fully homomorphic encryption to finally decrypt the FHE ciphertext to restore the plaintext result.
5. A large model hierarchical encryption deployment device based on sensitivity perception, characterized in that Including a structure division module, a request parsing module, a secure calculation routing module, and a distributed execution module. The structure division module divides the Transformer Block structure of the large model network into highly sensitive layers and less sensitive layers. Among them, the Transformer Block structure that is easily reverse-engineered to obtain the real input is the highly sensitive layer, and the Transformer Block structure that is not easily reverse-engineered to obtain the real input is the less sensitive layer. The request parsing module performs sensitive word detection on the request data input by the user. If there are sensitive words, the secure computing routing module directs the part of the request data containing sensitive words to the trusted execution environment TEE of the large model for encrypted calculation of the highly sensitive layer; the remaining part of the request data is split to the GPU cluster for calculation of the less sensitive layer. When the distributed execution module performs encrypted calculation of the highly sensitive layer, it first determines whether the memory resources of the trusted execution environment TEE are greater than the parameter calculation amount of the highly sensitive layer. If so, it directly routes the highly sensitive layer to the TEE for encrypted calculation. Otherwise, according to the memory capacity of the TEE, it calculates the maximum parameter calculation amount that can be accommodated, divides the parameter calculation amount of the highly sensitive layer into several blocks according to the maximum parameter calculation amount that can be accommodated, performs encrypted calculation on each block in the TEE, and encrypts and outputs the result to the GPU cluster. The GPU cluster calculates the blocks transmitted by the TEE and returns the calculation result to the TEE for decryption to obtain the plaintext result.
6. The apparatus for hierarchical encryption deployment of a large model based on sensitivity perception according to claim 5, characterized in that The structure division module divides the highly sensitive layer and the less sensitive layer, including: inputting test data, obtaining the output corresponding to each layer of TransformerBlock, respectively inputting the output of each layer into the model inversion attack simulator, allowing the attack simulator to reverse-engineer the real input, evaluating the error between the input reverse-engineered by the attack simulator and the real input. If the error is lower than the set threshold, the corresponding Transformer Block structure is the highly sensitive layer; if the error is higher than the set threshold, the corresponding Transformer Block structure is the less sensitive layer.
7. The large model hierarchical encryption deployment device based on sensitivity perception according to claim 5, characterized in that The distributed execution module performs encrypted calculation in the TEE, including: Encrypting the parameters and input data of the highly sensitive layer using the pre-generated fully homomorphic encryption public key to ensure that the data is executed in the encrypted state in the TEE calculation. After the encrypted calculation is completed in the TEE, an intermediate result in ciphertext form is generated. Dividing the intermediate result into blocks by rows, and attaching a message authentication code to each data block. Using the lightweight symmetric encryption algorithm AES-GCM to perform secondary encryption on the data blocks, and the TEE dynamically generates a key and shares it with the GPU cluster through a secure channel.
8. The apparatus for hierarchical encryption deployment of a large model based on sensitivity perception according to claim 7, characterized in that The distributed execution module calculates the blocks transmitted by the TEE in the GPU cluster and returns the calculation result to the TEE for decryption, including: After the GPU cluster receives the encrypted data block, the distributed execution module first decrypts the AES-GCM layer with the symmetric key and retains the ciphertext state of the fully homomorphic encryption. After the distributed execution module completes the calculation of the encrypted data block in the GPU, it re-encrypts the output result with AES-GCM and returns it to the TEE through a secure channel. After the TEE receives the result, the distributed execution module first decrypts the AES-GCM layer and then uses the private key of fully homomorphic encryption to finally decrypt the FHE ciphertext to recover the plaintext result.
Citation Information
Cited By
Model weight parameter protection method, terminal equipment and storage medium
CN121212349A