Large model security reasoning method and system fusing trusted execution environment and differential privacy

By dividing the large model into a distributed deployment architecture that integrates the embedded layer and the core processing layer, combining differential privacy and a trusted execution environment, the contradiction between data privacy protection and computing efficiency of the large model is solved, and efficient security reasoning and privacy protection are achieved.

CN120494087AActive Publication Date: 2025-08-15INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Patent Information

Application Number
CN202510484894.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-15
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

There is a contradiction between data privacy protection and computing efficiency in existing large models. Traditional TEE methods lead to frequent data interactions and trigger delays, and differential privacy methods are difficult to apply to black box service scenarios that protect model intellectual property rights.

Method used

The large model is divided into an embedding layer and a core processing layer. The embedding layer is deployed in a local trusted execution environment. The core processing layer is deployed on a remote GPU server. Differential privacy is used to obfuscate the text, and data is transmitted through the AES-GCM encryption protocol, combining the self-attention mechanism and the feedforward neural network for inference.

Benefits of technology

It significantly reduces the frequency of data interaction between TEE and GPU, reduces the inference delay caused by encryption and decryption operations, and realizes effective protection of user input data. It is suitable for black box service mode, and can adjust the privacy protection intensity and model performance in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494087A_ABST
    Figure CN120494087A_ABST
Patent Text Reader

Abstract

The invention discloses a large model security reasoning method and system fusing a trusted execution environment and differential privacy, and belongs to the technical field of artificial intelligence. The method comprises the steps that a large model is divided into an embedded layer and a core processing layer, the embedded layer is deployed in a local client SGX2 trusted environment, and the core processing layer is deployed in a local client SGX2 trusted environment; the core processing layer is deployed on a remote server equipped with multiple GPU acceleration cards; based on a local differential privacy technology, performing text confusion on original input data of a user by adopting an output noise maximum value algorithm; performing vector mapping on the confused text data through an embedded function; and the embedded vector is encrypted and transmitted to a server GPU, the embedded vector is processed through a self-attention mechanism and a feedforward neural network, and a prediction result is generated and returned to a client. According to the method, the problem of contradiction between data privacy protection and calculation efficiency in existing large model reasoning can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a large-model secure reasoning method and system that integrates a trusted execution environment and differential privacy. Background Art

[0002] Today, artificial intelligence is booming at an unprecedented pace, and the rise of large language models has ushered in a new era of intelligent interaction. At the end of 2022, OpenAI's ChatGPT garnered global attention, surpassing 100 million users in just two months. Since then, various large language models have proliferated. Large language models (LLMs) are language models with tens of billions or more parameters trained on massive amounts of text. These models, hereinafter referred to as large models, have been widely applied in various fields, including healthcare, education, law, and finance, thanks to their powerful human-computer interaction and task reasoning capabilities. These applications have brought significant economic benefits and demonstrated the enormous value potential of large models. However, with the increasing diversity and complexity of application scenarios, large models face significant challenges in security and privacy. Due to their high hardware resource requirements, large models are typically deployed on cloud servers, accessed by users through various interfaces. During interaction, large models must process vast amounts of user data, which contains a significant amount of sensitive personal information and trade secrets. Improper handling can easily lead to data security issues. Large-scale model attacks in recent years have further exacerbated users' concerns about privacy protection.

[0003] To protect user data, researchers have proposed a variety of privacy-preserving technologies, including secure multi-party computation, differential privacy, and confidential computing. Confidential computing, among others, uses hardware-based trusted execution environments (TEEs) to protect data privacy and offers improved computing performance and practicality. Using processor-level security extensions such as ARM TrustZone and Intel SGX, TEEs create an independent execution zone outside the main operating system. This ensures that even if the host system or virtual machine is compromised, the runtime state within the execution zone, including CPU registers, memory, and sensitive I / O, remains confidential and intact. Furthermore, TEEs offer remote attestation capabilities, enabling third-party verification of their trustworthiness. TEEs are typically integrated into CPUs and have resource constraints, making them unsuitable for running large deep learning models. Therefore, researchers have designed a delegation protocol to securely outsource some linear computations during inference to heterogeneous processors, such as GPUs, for acceleration. The TEE converts the linear layer data into an encrypted format, passes it to the GPU for computation, and then decrypts the result back to the original input for the nonlinear layer. To ensure the performance benefits of outsourcing computation, the encryption algorithm used must be minimally complex.

[0004] Differential privacy (DP) is another technical approach used to enhance data privacy in large models during fine-tuning and inference. It is a privacy standard defined in mathematical terms that characterizes the properties of the algorithm rather than the characteristics of the data. Initial differential privacy algorithms inject noise into statistical query results, preventing attackers from accurately inferring information about specific individuals in the dataset, even if they possess the query results. Differential privacy can be categorized as central differential privacy (CDP) or local differential privacy (LDP), depending on the location of data processing. CDP requires uniform noise addition to data on a central server and is suitable for centralized data collection and processing scenarios. LDP, on the other hand, allows data owners to perform localized random perturbations on data before transmitting it to untrusted data managers. In local differential privacy models, the trustworthiness of data managers is no longer a prerequisite, as the data they receive already meets the differential privacy criteria. Compared to central differential privacy, local differential privacy offers a significant advantage: data subjects do not need to trust any entity other than themselves. This advantage has led to the widespread adoption of LDP in real-world systems. Although the local differential privacy method can protect data privacy, it cannot guarantee the security of the noise addition process, and adding noise to all input texts will damage the reasoning performance of large models.

[0005] In summary, trusted execution environments (TEEs) and differential privacy (DP), as two important privacy-preserving technologies, each offer unique advantages in model reasoning and data processing. However, each technology alone still has limitations. Therefore, research on secure reasoning methods for large models that integrate TEEs and DP is of great significance. This not only combines the advantages of both technologies to provide more comprehensive privacy protection, but also effectively addresses increasingly complex security threats, providing new solutions for privacy-preserving computing in the era of large models. Summary of the Invention

[0006] In view of the defects in the existing technology of the trusted execution environment based method that frequent data interaction and encryption and decryption operations lead to a significant increase in inference delay, and the differential privacy based method requires access to model parameters and architecture, which is difficult to apply to black box service scenarios for protecting model intellectual property rights, the present invention proposes a large-model secure inference method and system that integrates trusted execution environment and differential privacy, which can solve the contradiction between data privacy protection and computing efficiency in existing large-model inference.

[0007] To achieve the above objectives, the technical solution of the present invention includes the following contents.

[0008] A large-model secure reasoning method that integrates a trusted execution environment (TEE) and differential privacy, using a local client endpoint, includes:

[0009] The large model is divided into an embedding layer and a core processing layer. The embedding layer is deployed in a trusted execution environment of a local client, and the core processing layer is deployed in a remote server equipped with multiple GPU accelerators. The embedding layer is used to map input text into embedding vectors, and the core processing layer is used to obtain prediction results for the input text based on the embedding vectors.

[0010] After obfuscating the original text based on differential privacy, the embedding vector generated based on the embedding layer is sent to the remote server, so that the remote server obtains and returns the prediction result of the embedding vector using the core processing layer;

[0011] A local vocabulary mapping is performed based on the prediction result of the embedding vector to obtain an inference result of the original text.

[0012] Furthermore, the obfuscating of the original text based on differential privacy includes:

[0013] Divide the words in the original text into sensitive words and non-sensitive words;

[0014] For each word w in the original text, the cosine similarity metric is used to retrieve the k nearest neighboring words w′ of the word in the word embedding space, and the cosine similarity between the word w and the neighboring words w′ is normalized by linear transformation to obtain the similarity cosnorm (w,w′);

[0015] In the similarity cos norm Add Laplace noise of strength ∈ to (w,w′) to generate the perturbation score of the neighboring word w′;

[0016] According to the perturbation score, a neighboring word w′ is selected as a candidate word for the word w;

[0017] If the word w is a sensitive word, the candidate word is used to replace the word w;

[0018] In the case that the word w is a non-sensitive word, the candidate word is used to replace the word w with a probability p, and the word w is kept unchanged with a probability 1-p.

[0019] Furthermore, the similarity Among them, cos(w,w′) represents the cosine similarity between the word w and the adjacent word w′, cos max =maxcos(w,w′), cos min =mincos(w,w′).

[0020] Furthermore, the core processing layer is used to obtain and return the prediction result of the embedding vector, including:

[0021] Processing the embedded vector through a multi-layer decoder network to obtain an output o of the multi-layer decoder network; wherein the decoder in the multi-layer decoder network adopts a self-attention mechanism and a feedforward neural network structure;

[0022] Map the output o to the vocabulary space through the linear layer SoftMax function to calculate the probability distribution of each candidate word;

[0023] The word with the highest probability is selected as the prediction result of the embedding vector.

[0024] Furthermore, the loss function L of the training model is total =L task +λ·L emo ; Among them, the task loss function T represents the number of words in the real corpus, w t represents the embedding vector of the t-th word, and p represents the model prediction w given the input of the first t-1 words. t The probability of w <t represents the first t-1 words of the input, λ is a hyperparameter used to balance the weights of different loss functions, EMO loss function p t represents the probability of generating the next word predicted by the model, et Represents the word embedding vector corresponding to the next word predicted by the large model, e j Represents the word embedding vector corresponding to the next word in the real corpus.

[0025] Furthermore, it is characterized in that the communication between the embedded layer and the core processing layer adopts the AES-GCM encryption protocol to transmit data.

[0026] A large-model secure reasoning system integrating a trusted execution environment and differential privacy, the system comprising:

[0027] A model deployment module, configured to divide the large model into an embedding layer and a core processing layer, deploying the embedding layer in a trusted execution environment on a local client and the core processing layer in a remote server equipped with multiple GPU accelerators; wherein the embedding layer is configured to map input text into embedding vectors, and the core processing layer is configured to obtain prediction results for the input text based on the embedding vectors;

[0028] An embedding vector generation module is configured to, after obfuscating the original text based on differential privacy, send the embedding vector generated based on the embedding layer to a remote server, so that the remote server uses the core processing layer to obtain and return a prediction result of the embedding vector;

[0029] The inference result acquisition module is used to perform local vocabulary mapping based on the prediction result of the embedding vector to obtain the inference result of the original text.

[0030] An electronic device, characterized in that the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements any of the above-mentioned large-model secure reasoning methods that integrate a trusted execution environment and differential privacy.

[0031] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the large-model secure reasoning method that integrates a trusted execution environment and differential privacy as described above is implemented.

[0032] A computer program product, characterized in that when the computer program product is run on a computer device, the computer device executes any of the above-mentioned large-model secure reasoning methods that integrate a trusted execution environment and differential privacy.

[0033] Compared with the prior art, the present invention has at least the following beneficial effects.

[0034] 1. This paper addresses the latency issues associated with frequent encrypted interactions caused by model segmentation in traditional TEE approaches. By proposing a distributed deployment architecture for the embedding and core processing layers, this approach completely confines the sensitive text mapping process to the local SGX trusted environment. Furthermore, it enables efficient communication between the client and the GPU server via PyTorch RPC. Experiments on the Llama series of models demonstrate that this strategy effectively reduces the latency of secure inference on large models.

[0035] 2. To minimize the impact of text obfuscation on model performance, we designed a text obfuscation mechanism based on hierarchical word frequency protection. This mechanism automatically identifies low-frequency sensitive words in the corpus, combines cosine similarity retrieval in the word embedding space to generate a set of semantically similar candidate words, and then adds Laplace noise perturbation to the similarity to achieve differentiated replacement. This replacement strategy ensures the privacy of user input while maintaining semantic coherence before and after the replacement.

[0036] 3. This paper addresses the issue of model performance degradation caused by noise perturbations by utilizing a hybrid loss function to enhance the model's semantic understanding of obfuscated text during downstream task fine-tuning. The feasibility of this approach has been verified in various downstream tasks, including text classification and conversation summarization.

[0037] In summary, the present invention significantly reduces the frequency of data interaction between TEE and GPU by optimizing the model deployment architecture, thereby reducing the inference delay caused by encryption and decryption operations. At the same time, it achieves effective protection of user input data without accessing the complete parameters and architecture of the model, which is suitable for the black box service model adopted by large real-world models. In addition, users can modify the differential privacy budget according to different application scenarios to balance the privacy protection strength and model performance. The present invention has good practicality and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a schematic diagram of the large model reasoning process of the method of the present invention.

[0039] Figure 2 This is a schematic diagram of the distributed deployment of the model proposed by the method of the present invention.

[0040] Figure 3 Schematic diagram of the text obfuscation mechanism proposed by the method of the present invention.

[0041] Figure 4 This is a comparison chart of the reasoning time decomposition of the present invention and the prior art. DETAILED DESCRIPTION

[0042] To better illustrate the large-model secure inference method proposed in the present invention that integrates a trusted execution environment with differential privacy, the present invention is further illustrated below using the application of the Llama2-7B-chat model on the text classification dataset SST-2 as an example, in conjunction with the accompanying drawings and specific implementation methods.

[0043] Figure 1 This is the flow chart of the present invention, which is mainly divided into five parts: model distributed deployment, text obfuscation processing, embedding vector mapping, large model decoding output and model fine-tuning.

[0044] Step 1: Distributed deployment of the model.

[0045] Based on the computational characteristics of large models and the need for user privacy protection, this paper divides large models into two key components: the embedding layer and the core processing layer. The main responsibility of the embedding layer is to map the input text data into a unified vector space for subsequent model processing. Figure 2 As shown, the embedding layer is deployed in the SGX2 trusted environment of the local client, and the privacy and integrity of the user input data are ensured through hardware-level security isolation, while ensuring the confidentiality of the embedding layer on the client. On the other hand, the core processing layer of the large model undertakes more complex computing tasks and is deployed on a remote server equipped with multiple GPU accelerator cards to give full play to the significant advantages of GPU in parallel computing. In order to achieve efficient communication between the embedding layer and the core processing layer, the present invention adopts the Remote Procedure Call (RPC) interface in the PyTorch framework. The layered deployment strategy not only retains the parallel computing advantages of the GPU, but also limits sensitive data processing tasks to a secure local area through a physical isolation mechanism.

[0046] Step 2: Text obfuscation.

[0047] Based on local differential privacy technology, this paper performs text obfuscation on the original input data to protect user sensitive data. Specifically, although the distributed deployment of the model eliminates the need for users to send plaintext input data to the server, the output of the embedding layer may still leak private data. Attackers can implement inversion attacks through the intermediate layer features, namely the embedding vectors. This paper proposes a text obfuscation mechanism based on the output noise maximum algorithm, such as Figure 3 As shown in Figure 1, the mechanism constructs a hierarchical protection strategy through word frequency analysis: words with a frequency lower than a threshold value θ are classified as sensitive words, and the rest are classified as non-sensitive words. Stop words are all non-sensitive words. Assuming that the user dictionary is V, the word frequency distribution frequency(x) in a corpus is counted. x∈V , set the threshold to θ to build a sensitive word set:

[0048] S={x∈V|frequency(x)<θ}.

[0049] The set of non-sensitive words is N = VS, which mainly includes high-frequency words and stop words. Due to their high semantic redundancy and weak privacy relevance, a relaxed handling strategy is adopted. The implementation of the text obfuscation mechanism relies on the following two key functions:

[0050] Candidate word generation function: for the target word w i , this function uses the cosine similarity metric to retrieve the nearest k neighboring words (excluding the target word itself) in the word embedding space and uses normalization to calculate the similarity. First, calculate the standard cosine similarity:

[0051]

[0052] Then, a linear transformation is applied to the similarity values in the candidate set to normalize them:

[0053]

[0054] Among them, cos max =maxcos(w,w j )(1≤j≤k) and cos min =mincos(w,w j )(1≤j≤k) represent the maximum and minimum similarity values in the candidate word set respectively.

[0055] Random mapping function: This function defines the probability function P to implement the differentiated processing strategy. For sensitive words, add Laplace noise with strength ∈ to generate a perturbation score, and select the candidate word with the highest score after noise correction to replace it:

[0056] score(w,w j )=cos norm (w,w j )+z,

[0057]

[0058] For non-sensitive words: execute the above text obfuscation strategy with probability p, and remain unchanged with probability 1-p.

[0059] Step 3: Embedding vector mapping.

[0060] Obfuscated user input data X = <w′1,w′2,…,w′ nIt is first sent to the embedding function deployed in the local trusted environment, which performs the vector mapping operation. The mapping process generates a fixed-dimensional vector representation by multiplying the embedding matrix with the input data. These vector representations retain the semantic information of the original input while hiding the specific content. The calculation formula is:

[0061] h=f emb (X) = W emb X,

[0062] Where h represents the generated embedding vector, W emb represents the embedding matrix, and E(·) represents the embedding function.

[0063] Step 4: Large model decoding output.

[0064] After receiving the encrypted embedding vector, the server performs the following operations to encrypt and transmit the embedding vector to the server's GPU.

[0065] Decoding stage: The embedding vector h is first processed by a multi-layer decoder network. The decoder uses a self-attention mechanism and a feedforward neural network structure, and the formula is expressed as:

[0066] h′=Decoder(h)=MultiHead(LayerNorm(h))+h,

[0067] o=FFN(LayerNorm(h′))+h′,

[0068] Among them, MultiHead represents the multi-head self-attention layer, LayerNorm represents the layer normalization operation, and FFN represents the feedforward neural network.

[0069] Output layer processing: The output o of the decoder is mapped to the word table space through the linear layer SoftMax function, and the probability distribution of each candidate word is calculated:

[0070] p(w n+1 |w′1,w′2,…,w′ n )=SoftMax(W o o+b o ),

[0071] Among them, W o and b o is the weight matrix and bias vector of the output layer. Finally, the word with the highest probability is selected as the prediction result according to the probability distribution:

[0072]

[0073] The server returns the predicted next word ID to the client, which then performs local vocabulary mapping to convert the ID into an actual word.

[0074] Step 5: Model fine-tuning.

[0075] Calculate the hybrid loss function and efficiently fine-tune the parameters of the large model on the downstream task dataset. To reduce the impact of the text obfuscation mechanism on the model performance, the model can be fine-tuned for a small number of cycles on the dataset. Based on the cross-entropy loss function, the EMO loss function is introduced to update the model gradient. The hybrid loss function is calculated as follows:

[0076] L total =L task +λ·L emo ,

[0077] Among them, λ is a hyperparameter used to balance the weights of different loss functions.

[0078] The calculation formula of the task loss function is:

[0079]

[0080] Among them, T represents the number of words in the real corpus, w t represents the embedding vector of the t-th word, and p represents the model prediction w given the input of the first t-1 words. t The probability of w <t represents the first t-1 words of the input.

[0081] The EMO loss function calculation formula is:

[0082]

[0083] Among them, p t is the probability of generating the next word predicted by the model, e t Represents the word embedding vector corresponding to the next word predicted by the model, and both are calculated through the forward propagation process of the model. j The word embedding vector corresponding to the real word is usually obtained by mapping the embedding layer of a large model.

[0084] The gradient update formula during training is:

[0085]

[0086] Where η is the learning rate, is the gradient of the model parameters θ.

[0087] Fine-tuning with a hybrid loss function can enhance the model's generalization ability for obfuscated text. The model can better understand and process text-obfuscated inputs, thereby maintaining high prediction performance while ensuring privacy and security.

[0088] This solution not only retains the parallel computing advantages of GPU, but also protects user sensitive data through the physical isolation mechanism and differential privacy of TEE. While ensuring privacy and security, it maintains high inference performance and effectively balances the relationship between security and efficiency.

[0089] The method of the present invention is further illustrated below by taking the Llama2-7B-chat model and the SST-2 dataset as examples.

[0090] Step 1. Distributed model deployment. The Llama2-7B-chat model is divided into two key components: the embedding layer is deployed on a local client equipped with Intel SGX2, and the core processing layer is deployed on a remote server equipped with two NVIDIA A100 GPUs. The embedding layer contains an embedding matrix of 4,096 dimensions for 32,000 tokens; the core processing layer consists of a 32-layer Transformer decoder, each with 32 attention heads. These two components establish a secure communication channel using PyTorch's RPC interface.

[0091] Step 2. Text obfuscation. Taking the SST-2 dataset as an example, we first calculate the word frequency distribution and set a threshold θ = 0.2 (i.e., words with a frequency of less than 20% are considered sensitive words). For identified sensitive words, we apply Laplace noise with an intensity of ∈ = 0.7 for perturbation; for non-sensitive words, we set a replacement probability p = 0.3. After candidate word screening, we select the k = 20 most similar words, calculate cosine similarity, and perform normalization. For example, the user input text "we never originally feel involved with the movie, as all della its tips remain just that: abstract opinions" is transformed into "we never really feel involved with the story, as all of its ideas remain just that: abstract ideas" after the text obfuscation mechanism. In this text, "originally" is replaced with "really," "movie" is replaced with "story," "della" is replaced with "of," "tips" is replaced with "ideas," and "opinions" is replaced with "ideas." These replacements retain the basic semantic structure of the original text, but reduce the personalized features and sensitive information in the text.

[0092] Step 3: Embedding vector mapping. The obfuscated text is combined with the prompt template "Classify the sentiment of the given movie review text into Positive or Negative. Movie Review:" and mapped into a 4,096-dimensional vector using an embedding function. This calculation is performed within the SGX2 trusted environment. After the embedding calculation is completed, the vector is encrypted using the AES-GCM-256 algorithm, and the key management is securely exchanged through the SGX2 remote attestation mechanism.

[0093] Different templates can be substituted depending on the downstream tasks of the large model. The purpose of concatenating the obfuscated text and the prompt template is to improve the model's performance on specific tasks.

[0094] Step 4. Decode the output. The server GPU receives the encrypted embedding vector, decrypts it, and processes it through the Transformer decoder. For the SST-2 sentiment classification task, the maximum decoding length is set to 10, and the temperature parameter is set to 0.1 to enhance deterministic output. The decoder output passes through a softmax layer, selecting the word sequence with the highest probability. This is returned to the client via a token ID. The client then uses a local vocabulary mapping to restore the actual word. The final output is "Negative," indicating that the obfuscated text has been classified as having negative sentiment.

[0095] Step 5. Model fine-tuning. To reduce the impact of text obfuscation on performance, efficient parameter fine-tuning is performed on the SST-2 training set. A hybrid loss function is used for fine-tuning, with the weight λ set to 0.5, and the cross-entropy loss and EMO loss are fused in a 1:1 ratio. The AdamW optimizer is used, the learning rate is set to 3e-5, the weight decay is 0.01, and fine-tuning is performed for 3 cycles. The LoRA low-rank adaptation technology is used, the rank is set to 8, only the parameters of the model attention module are updated, and the remaining parameters are frozen to save computing resources. Experimental results show that when ε = 2, the accuracy of the proposed method on the SST-2 test set reaches 90%, which is only 7% lower than the original model. At the same time, compared with the method based entirely on TEE, the inference latency is greatly reduced.

[0096] Finally, to verify the effectiveness of the proposed invention, this experiment selected SANTEXT+(Yue X,Du M,Wang T,etal.Differential privacy for text analytics via natural text sanitization[C] / / Zong C,Xia F,Li W,et al.Findings of ACL:ACL / IJCNLP 2021Findings of theAssociation for Computational Linguistics:ACL / IJCNLP 2021,Online Event,August1-6,2021.Association for Computational Linguistics,2021:3853-3866.) and CusText+(Chen S,Mo F,Wang Y,et al.Acustomized text sanitization mechanism withdifferential privacy[C] / / Rogers A,Boyd-Graber JL,Okazaki N.Findings of theAssociation for Computational Linguistics:ACL 2023, Toronto, Canada, July 9-14, 2023. Association for Computational Linguistics, 2023: 5747-5758.) was used as a baseline model for comparison. The Bert-base-uncased model was tested on the text classification task datasets SST-2 and QNLI. The following is a brief introduction to the two baseline models:

[0097] SANTEXT+: This method divides vocabulary into sensitive sets Non-sensitive set We also constructed a utility-optimized local differential privacy (Metric LDP) framework. This framework adaptively allocates privacy budgets based on semantic distance. During pre-training, the model uses perturbed text for mask prediction, enhancing its robustness to noise.

[0098] CusText+: This model uses a mapping function to assign a customizable output set to each input word. The size of this output set can be flexibly adjusted via the hyperparameter k to meet different privacy and practicality requirements. The model uses an exponential sampling function to select a token from the custom output set to replace the original token.

[0099] Table 1 lists the classification accuracy of different privacy-preserving mechanisms on the SST-2 and QNLI datasets, under similar privacy protection levels. k controls the size of the candidate set for each word. As the privacy budget ∈ increases, the accuracy of the BERT model on SST-2 and QNLI improves slightly due to reduced noise perturbations. Across different privacy budgets, the secure inference method proposed in this paper consistently outperforms other baseline methods.

[0100] Table 1 Comparison of the accuracy of different privacy protection mechanisms

[0101]

[0102]

[0103] To evaluate the additional inference overhead introduced by this invention, this experiment uses two lightweight large models, Llama2-7B-chat and Llama3-8B-Instruct, and uses the average token generation time on the SST-2 dataset as the core evaluation metric. The inference efficiency is calculated through multiple measurements. The inference efficiency metric is based on the average forward propagation time and is calculated as follows:

[0104]

[0105] in, is the inference time of the jth sample, and M is the total number of test samples, which defaults to 100. Due to the different preconditions of previous trusted execution environment-based methods, it is difficult to conduct a fair comparison directly. Therefore, this paper establishes two evaluation baseline models:

[0106] CPU+GPU baseline: In this baseline setup, the client executes the embedding layer of the large model in a regular computing environment, while the core layer of the model runs on a GPU server. No privacy protection measures are implemented during the entire inference process.

[0107] GPU baseline: All layers of the large model are deployed on the GPU server for execution. Similarly, no privacy protection mechanism is adopted in the whole process.

[0108] like Figure 4As shown, the method proposed in the present invention introduces additional time overhead of 0.7152 seconds (accounting for 28.78%) and 0.7432 seconds (accounting for 38.12%) on the Llama2-7B and Llama3-8B models respectively compared with the CPU+GPU baseline. The main source of these overheads is the SGX2 encryption mechanism, which uses a hardware-level memory encryption engine to provide a strong guarantee of data confidentiality and integrity for the secure area. The time overhead generated by the text obfuscation mechanism, including the candidate word generation and replacement process, is relatively small, accounting for only 0.6% and 0.74% of the total reasoning time. In addition, since the candidate words of different large models can be pre-calculated and stored, the efficiency of this mechanism in practical applications can be further improved.

[0109] In summary, due to the limited hardware resources within the TEE, researchers have proposed a deep learning model segmentation strategy to improve inference efficiency while ensuring data privacy. This strategy offloads the linear computational portion of the model to the GPU for acceleration. Depending on how the input features are encrypted, these methods can be divided into computation outsourcing methods based on one-time password (OTP) encryption and computation outsourcing methods based on linear transformations. Because nonlinear functions often involve complex operations such as exponential and logarithmic functions, current encryption schemes still struggle to efficiently support them. Large models are typically composed of multiple layers of stacked self-attention mechanisms, with each layer consisting of alternating linear and nonlinear computations. Therefore, when applied to large models, existing model segmentation methods result in frequent data interactions between the TEE and the GPU. Each data exchange requires encryption and decryption, a mechanism that is crucial for protecting data security but also significantly increases inference latency.

[0110] Model security reasoning methods based on differential privacy, based on rigorous mathematical proofs, provide effective defense mechanisms for large models against various privacy threats, including membership inference attacks, attribute inference attacks, and embedding inversion attacks. However, using differential privacy for privacy protection requires access to the model's parameters and architecture, which is difficult to achieve in real-world applications. Model parameters and architecture are often core assets of an enterprise. Therefore, to protect model intellectual property, many deep learning and large models are provided as black-box services, allowing users to obtain output results only through input data.

[0111] Based on the above, this paper proposes a collaborative security framework that integrates a trusted execution environment (TEE) and differential privacy. By splitting large models into a layered deployment architecture consisting of an embedding layer and a core processing layer, embedding vector mapping of sensitive text is performed in the client's local SGX2 trusted environment, leveraging hardware-level isolation to ensure the privacy of input data. Meanwhile, the computationally intensive core processing layer is deployed on a GPU server to maintain efficient parallel computing capabilities, and the two layers achieve collaborative reasoning through encrypted communication.

[0112] Furthermore, to protect the privacy of user input data, this paper also designs a text obfuscation mechanism based on the output noise maximization algorithm. This mechanism applies Laplace noise perturbation to low-frequency sensitive words and generates semantically similar replacement words based on cosine similarity retrieval in the word embedding space. A probabilistic retention strategy is implemented for high-frequency non-sensitive words, thus protecting privacy while maximizing semantic coherence.

[0113] While the specific details, implementation algorithms, and drawings of the present invention are disclosed for illustrative purposes, intended to facilitate understanding and implementation of the present invention, those skilled in the art will appreciate that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. The present invention should not be limited to the preferred embodiments disclosed in this specification and the accompanying drawings; the scope of protection claimed by the present invention shall be determined by the scope defined in the claims.

Claims

1. A large-model secure reasoning method that integrates a trusted execution environment and differential privacy, characterized by: Using a local client endpoint, the method includes: The large model is divided into an embedding layer and a core processing layer. The embedding layer is deployed in a trusted execution environment of a local client, and the core processing layer is deployed in a remote server equipped with multiple GPU accelerators. The embedding layer is used to map input text into embedding vectors, and the core processing layer is used to obtain prediction results for the input text based on the embedding vectors. After obfuscating the original text based on differential privacy, the embedding vector generated based on the embedding layer is sent to the remote server, so that the remote server obtains and returns the prediction result of the embedding vector using the core processing layer; A local vocabulary mapping is performed based on the prediction result of the embedding vector to obtain an inference result of the original text.

2. The method according to claim 1, characterized in that The obfuscation of the original text based on differential privacy includes: Divide the words in the original text into sensitive words and non-sensitive words; For each word w in the original text, the cosine similarity metric is used to retrieve the k nearest neighboring words w′ of the word in the word embedding space, and the cosine similarity between the word w and the neighboring words w′ is normalized by linear transformation to obtain the similarity cos norm (w, w′); In the similarity cos norm Add Laplace noise of strength ∈ to (w, w′) to generate the perturbation score of the neighboring word w′; According to the perturbation score, a neighboring word w′ is selected as a candidate word for the word w; If the word w is a sensitive word, the candidate word is used to replace the word w; In the case that the word w is a non-sensitive word, the candidate word is used to replace the word w with a probability p, and the word w is kept unchanged with a probability 1-p.

3. The method according to claim 2, characterized in that The similarity Among them, cos(w, w′) represents the cosine similarity between the word w and the adjacent word w′, cos max =max cos(w, w′), cos min =min cos(w, w′).

4. The method according to claim 1, wherein The core processing layer is used to obtain and return the prediction result of the embedding vector, including: Processing the embedded vector through a multi-layer decoder network to obtain an output o of the multi-layer decoder network; wherein the decoder in the multi-layer decoder network adopts a self-attention mechanism and a feedforward neural network structure; Map the output o to the vocabulary space through the linear layer SoftMax function to calculate the probability distribution of each candidate word; The word with the highest probability is selected as the prediction result of the embedding vector.

5. The method according to claim 1, wherein The loss function L for training the large model total =L task +λ·L emo ; Among them, the task loss function T represents the number of words in the real corpus, w t represents the embedding vector of the t-th word, and p represents the model prediction w given the input of the first t-1 words. t The probability of w <t represents the first t-1 words of the input, λ is a hyperparameter used to balance the weights of different loss functions, EMO loss function p t represents the probability of generating the next word predicted by the model, e t Represents the word embedding vector corresponding to the next word predicted by the large model, e j Represents the word embedding vector corresponding to the next word in the real corpus.

6. The method according to any one of claims 1 to 5, characterized in that The communication between the embedded layer and the core processing layer uses the AES-GCM encryption protocol to transmit data.

7. A large-model secure reasoning system that integrates a trusted execution environment and differential privacy, characterized by: The system comprises: A model deployment module, configured to divide the large model into an embedding layer and a core processing layer, deploying the embedding layer in a trusted execution environment on a local client and the core processing layer in a remote server equipped with multiple GPU accelerators; wherein the embedding layer is configured to map input text into embedding vectors, and the core processing layer is configured to obtain prediction results for the input text based on the embedding vectors; An embedding vector generation module is configured to, after obfuscating the original text based on differential privacy, send the embedding vector generated based on the embedding layer to a remote server, so that the remote server uses the core processing layer to obtain and return a prediction result of the embedding vector; The inference result acquisition module is used to perform local vocabulary mapping based on the prediction result of the embedding vector to obtain the inference result of the original text.

8. An electronic device, characterized in that: The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the large-model secure reasoning method that integrates a trusted execution environment and differential privacy as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the large-model secure reasoning method that integrates a trusted execution environment and differential privacy as described in any one of claims 1-6.

10. A computer program product, characterized in that When the computer program product runs on a computer device, the computer device executes the large-model secure reasoning method that integrates a trusted execution environment and differential privacy as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Outsourcing deep learning system supporting privacy protection

    CN115587379A

  • Large-model-oriented hybrid reasoning method, device, equipment, medium and product

    CN118798353A

  • Privacy protection method for Transform model in end-side equipment reasoning scene

    CN119026166A

  • Model fine tuning method, text processing method, medium, equipment and program product

    CN119150862A

  • Private data anonymization protection method

    CN119227133A

Cited By

  • Multi-technology fusion intelligent document processing method and system

    CN121682898A

  • Multi-center badminton action evaluation method based on trusted execution environment

    CN121838277A