A trusted execution environment-based transformer model privacy protection inference method

CN122735014APending Publication Date: 2026-09-11XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611024648.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

然而,尽管其表达能力强,纯MPC在实际部署中仍面临巨大的效率挑战

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122735014A_ABST
    Figure CN122735014A_ABST
Patent Text Reader

Abstract

This invention discloses a privacy-preserving inference method for Transformer models based on a trusted execution environment, relating to the field of data security. The method includes: users secretly sharing an additive sequence of token indexes of the data to be inferred, generating two index fragments, which are then sent to both computing servers; both computing servers convert the index fragments into ciphertext embedding vector fragments, using these fragments as the initial hidden state input for the first round of the Transformer layer, performing N rounds of Transformer layer iterative computation to obtain the final ciphertext output; both computing servers perform secure blinding processing on the final ciphertext output to obtain blinded fragments, which are then sent to the user; after receiving the blinded fragments, the user performs a modular addition operation on the blinded fragments locally to reconstruct the plaintext data, thus obtaining the final plaintext inference decision. This invention can reduce communication overhead.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data security, and particularly relates to a Transformer model privacy protection reasoning method based on a trusted execution environment. BACKGROUND

[0002] With the rapid development of machine learning technology, the "model as a service" mode has been widely applied in many fields such as natural language processing, recommendation system and code generation. The Transformer model and its variants have become the cornerstone of this mode due to their excellent feature extraction and reasoning capabilities. In the specific application of many fields such as natural language processing, recommendation system and code generation, more and more users rely on third-party reasoning platforms and tool chains to conveniently access models based on the Transformer architecture. These models are usually developed by one party, deployed on cloud infrastructure, and accessed by end users through remote APIs.

[0003] However, traditional Transformer inference requires users to upload sensitive data to cloud infrastructure, while model providers need to expose model parameters and structures to complete inference, which exacerbates a serious trust crisis: on the one hand, users face the risk of sensitive input leakage during inference; on the other hand, model providers are also exposed to the potential risk of leakage of proprietary model parameters and intellectual property. To solve the above privacy and security challenges, in recent years, academia has proposed various privacy-preserving Transformer inference schemes that use cryptography to achieve secure inference without leaking user input or model parameters. Among them, secure multi-party computation (MPC) is widely used in Transformer inference because it supports general function evaluation and can achieve high accuracy when dealing with nonlinear functions. Pure MPC-based schemes usually use secret sharing technology for secure multiplication. However, despite its strong expressive power, pure MPC still faces significant efficiency challenges in practical deployment. Such schemes have significant drawbacks: on the one hand, linear calculations (such as large-scale matrix multiplication) often rely on a large number of multiplication triples, resulting in high preprocessing and communication overhead; on the other hand, nonlinear calculations often require multiple rounds of interaction, increasing the number of communication rounds and further limiting overall inference efficiency. To overcome the computational and communication efficiency bottlenecks faced by pure MPC schemes, existing cutting-edge research widely introduces trusted execution environments (TEEs) as an effective means to reduce computational and communication overhead. Based on full-trust TEE-assisted privacy-preserving inference schemes, TEE is treated as a completely trusted black-box entity that performs plaintext calculations on decrypted data within the Enclave, significantly improving efficiency. These methods rely on the assumption that the Enclave implementation can provide protection against side-channel leakage. However, such schemes have a fatal flaw: they rely heavily on the strong trust assumption that the Enclave's execution is confidential. However, in the actual inference service scenario, this assumption is often difficult to maintain, and the scheme ignores the realistic threats of code auditability and covert data leakage (intermediate states are stolen). SUMMARY

[0004] The present application aims to at least solve the above technical problems existing in the prior art, and for this purpose, the present application provides a Transformer model privacy-preserving inference method based on a trusted execution environment, comprising: The user performs additive secret sharing on the Token index sequence of the data to be inferred, generates two index shards, and sends them to the two computing servers; Both computing servers execute a secure embedding protocol to convert the index fragments into ciphertext embedding vector fragments. These ciphertext embedding vector fragments are then used as the initial hidden state input for the first round of the Transformer layer. Based on a predefined protocol, N rounds of Transformer layer iterative computation are performed to obtain the final ciphertext output. The predefined protocols include: one-way secure scalar multiplication protocol, secure matrix multiplication protocol, one-way three-input secure multiplication protocol, unintentional DReLU protocol, unintentional ReLU protocol, basic hybrid approximation fitting protocol, hybrid approximation fitting protocol, enhanced bounded function hybrid approximation fitting protocol, enhanced exponential function hybrid approximation fitting protocol, enhanced hybrid approximation fitting protocol, secure nonlinear activation protocol, secure layer normalization protocol, and secure maximum exponential normalization protocol. Both computing servers use local random number masks to perform security blinding processing on the final ciphertext output to obtain blinded fragments, and send the blinded fragments to the user; After receiving the blinded data fragment, the user performs a modulo-addition operation on the blinded data fragment locally to reconstruct the plaintext data and obtain the final plaintext reasoning decision.

[0005] Optionally, the execution of N rounds of Transformer layer iterative computation includes: Based on the initial hidden state input of the current layer, and by invoking the secure matrix multiplication protocol and the secure maximum exponent normalization protocol, the attention output of the current layer is obtained; Based on the attention output of the current layer, the two computing servers use residual connections and call the security layer normalization protocol and the enhanced hybrid approximation fitting protocol to obtain the first intermediate ciphertext feature; the first intermediate ciphertext feature is the normalized intermediate ciphertext feature. For the first intermediate ciphertext feature, the secure matrix multiplication protocol, the secure nonlinear activation protocol, and the secure matrix multiplication protocol are executed sequentially to obtain the second intermediate ciphertext feature; the second intermediate ciphertext feature is an intermediate ciphertext feature processed by a feedforward network and nonlinear activation. Based on the second intermediate ciphertext feature, the two computing servers use a residual connection to execute the security layer normalization protocol and obtain the final ciphertext output of the current layer. Determine if the current layer is the last layer. If not, use the final ciphertext output of the current layer as the initial hidden state input for the next Transformer layer and continue execution. If so, use the final ciphertext output of the current layer as the final ciphertext inference result and output it.

[0006] Optionally, obtaining the attention output of the current layer includes: Based on the initial hidden state input of the current layer, the secure matrix multiplication protocol is invoked to perform ciphertext linear transformations corresponding to the query matrix Q, key matrix K, and value matrix V, respectively, to obtain query features, key features, and value features. The secure matrix multiplication protocol is invoked to calculate the matrix product of the query feature and the key feature to obtain the attention score matrix; The security maximum exponential normalization protocol is invoked to perform security Softmax calculation on the attention score matrix to obtain the attention weight matrix; The step of performing a secure Softmax calculation on the attention score matrix to obtain the attention weight matrix includes: The enhanced exponential function hybrid approximation fitting protocol is invoked to calculate the exponential function; The enhanced hybrid approximation fitting protocol is invoked to calculate the reciprocal, and the one-way secure scalar multiplication protocol is invoked to complete the normalization calculation, thereby obtaining the attention weight matrix; During the execution of Newton's iteration, the enhanced hybrid approximation fitting protocol calls the unidirectional three-input safe multiplication protocol to complete three multiplication calculations. The secure matrix multiplication protocol is invoked to calculate the matrix product of the attention weight matrix and the value feature, thereby obtaining the attention output of the current layer; The secure matrix multiplication protocol includes: The two computing servers first randomly blind the input matrix by partitioning it, and then exchange the blinding results to reconstruct the matrix difference. The matrix product correction calculation is completed based on matrix difference and matrix multiplication auxiliary random numbers generated by TEE, so as to restore the secret shared fragment of the matrix multiplication result locally.

[0007] Optionally, obtaining the first intermediate ciphertext feature includes: The attention output of the current layer is added to the initial hidden state input to obtain the first residual ciphertext feature; The security layer normalization protocol is invoked to calculate the mean, variance, and reciprocal of the standard deviation of the first residual ciphertext feature to obtain the first intermediate ciphertext feature; the first intermediate ciphertext feature is the normalized feature. The step of invoking the security layer normalization protocol to calculate the mean, variance, and reciprocal of the standard deviation of the first residual ciphertext features includes: The enhanced hybrid approximation fitting protocol is invoked to complete the approximate calculation of the reciprocal of the square root; The normalized multiplication calculation is completed by invoking the one-way secure scalar multiplication protocol.

[0008] Optionally, obtaining the second intermediate ciphertext feature includes: The secure matrix multiplication protocol is applied to the first intermediate ciphertext feature and the first layer weight matrix of the feedforward network, and the bias is superimposed to obtain the upgraded feature. Invoking the secure nonlinear activation protocol to perform nonlinear activation calculations and obtain activation output; the invocation of the secure nonlinear activation protocol to perform nonlinear activation calculations includes: The enhanced bounded function hybrid approximation fitting protocol or the enhanced exponential function hybrid approximation fitting protocol is invoked according to the activation function type to complete the approximate calculation of the nonlinear function; For the activated output, the secure matrix multiplication protocol is invoked to complete the calculation of the second linear layer of the feedforward network, thereby obtaining the second intermediate ciphertext feature.

[0009] Optionally, obtaining the final ciphertext output of the current layer includes: Perform ciphertext addition on the first residual ciphertext feature and the second intermediate ciphertext feature to obtain the second residual feature; The security layer normalization protocol is invoked to calculate the mean, variance, and reciprocal of the standard deviation of the second residual feature, thereby obtaining the final ciphertext output of the current Transformer layer. The step of invoking the security layer normalization protocol to calculate the mean, variance, and reciprocal of the standard deviation of the second residual feature includes: The enhanced hybrid approximation fitting protocol is invoked to perform the approximate calculation of the reciprocal of the square root, and the one-way safe scalar multiplication protocol is invoked to perform the normalization calculation.

[0010] Optionally, the enhanced bounded function hybrid approximation fitting protocol includes: Both computing servers invoke the basic hybrid approximation fitting protocol to complete the initial approximation calculation of the bounded function; The input interval is identified based on the initial approximate calculation results of the bounded function and the safe interval determination protocol. The corresponding fitting parameters are selected for local correction according to different intervals, and the secret sharing results of the bounded function are output. The enhanced exponential function hybrid approximation fitting protocol includes: Both computing servers perform interval reduction on the input data and use a basic hybrid approximation fitting protocol to calculate the approximate value of the exponential function within the reduction interval; Based on the approximate value of the exponential function within the reduction interval and the exponential reconstruction relationship, the complete exponential calculation result is recovered, and the secret sharing result of the exponential function is obtained. The enhanced hybrid approximation fitting protocol includes: Both computing servers automatically invoke the corresponding enhanced approximation protocol to complete the local function calculation based on the type of the objective function; The results are fused and corrected based on the calculation results of local functions and Newton's iterative operations, and a secure calculation output secretly shares the approximate result.

[0011] Optionally, the one-way three-input secure multiplication protocol includes: Both computing servers perform random blinding on the three input secret fragments respectively, and exchange the blinding values ​​to restore the corresponding blinding difference; The three-input product correction calculation is completed based on the blinded difference of the recovery and the three-input auxiliary multiplication data output by TEE, and the secret sharing result corresponding to the three-input product is obtained. The one-way secure scalar multiplication protocol includes: Both computing servers use local secret fragments and random mask fragments generated by TEE to calculate the input blinding difference; Both computing servers complete plaintext reconstruction by exchanging blinding differences, and obtain the reconstructed plaintext. The product correction term is calculated based on the reconstructed plaintext and the multiplication auxiliary data output by the TEE to obtain the secret sharing fragment corresponding to the product result; the product result is determined based on the product correction term.

[0012] Optionally, the unintentional DReLU protocol includes: The two computing servers first use a random mask to blind the input data, and then exchange fragments to restore the blinded values; Based on the comparison of the derivatives calculated independently of the public blinding value and the local random mask, the secret sharing result corresponding to the DReLU function is output; The unintentional ReLU protocol includes: Both computing servers first invoke the unintentional DReLU protocol to obtain the input symbol bits; Secure multiplication is performed based on the input symbol bits and the input secret slice, retaining only the data corresponding to the positive interval and masking the data in the negative interval, and outputting the secret shared slice of the ReLU function result.

[0013] Optionally, before the user shares the addition secret of the token index sequence of the data to be inferred, the method further includes: Users and servers initialize the encryption system and negotiate parameters to determine the arithmetic secret sharing mechanism in order to complete the encryption system configuration; The model owner performs a secret addition split on the pre-trained Transformer model parameter set, generating two parameter fragments, which are then sent to two computing servers via a secure channel to achieve secret deployment of the model parameters. The cryptographic service provider deploys TEE code on two computing servers respectively; the computing servers and the corresponding TEE code negotiate and share a random number seed through a secure channel.

[0014] The present invention has at least the following beneficial effects: 1. This invention introduces a dual random masking mechanism to ensure the confidentiality of user input and model parameters, making it very suitable for real-world threat environments of private Transformer inference. This invention divides the overall computation task into two independent components, allowing two computing servers to process independently. On this basis, it makes full use of the computational logic within each component to support parallel execution and cleverly utilizes the idle time during message network transmission to achieve a perfect overlap between local computation and network communication. This not only reduces communication overhead but also maintains extremely high end-to-end performance under high-latency network conditions such as real wide area networks (WANs). 2. This invention introduces a brand-new MPC protocol, which enables one-way communication. Calculations can be completed by transmitting data in only one direction, eliminating the need for multiple rounds of interaction and significantly reducing communication round-trip overhead.

[0015] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0016] Figure 1 This is a flowchart of a privacy-preserving inference method for Transformer models based on a trusted execution environment, provided by an embodiment of the present invention. Figure One ; Figure 2 This is a flowchart of a privacy-preserving inference method for Transformer models based on a trusted execution environment, provided by an embodiment of the present invention. Figure Two . Detailed Implementation

[0017] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0018] This invention provides a privacy-preserving inference method for Transformer models based on a trusted execution environment, such as... Figure 1 and Figure 2 As shown, it includes the following steps: Step 101: The user performs addition and secret sharing on the Token index sequence of the data to be inferred, generates two index fragments, and sends them to the computing servers of both parties.

[0019] In one possible implementation, prior to the user sharing the additive secret of the token index sequence of the data to be inferred, the method further includes: Step 1011: The user and the server initialize the encryption system and negotiate parameters to determine the arithmetic secret sharing mechanism in order to complete the encryption system configuration.

[0020] Specifically, the system initialization phase only needs to be executed once and can be performed when the system is idle. It aims to complete the secret deployment of the encrypted system configuration, model parameters, and negotiation of random number seeds, completely avoiding the network transmission of massive auxiliary data in traditional MPC.

[0021] The user and server negotiate and determine an addition secret sharing mechanism that supports outsourced computation. This embodiment of the invention adopts a mechanism based on... Bit Integer Ring of An arithmetic secret sharing mechanism, exemplarily, can be set up... The encryption system supports the following basic operations and configurations: (1) Fixed-point encoding configuration, system sets scaling factor For example, it can be set This scaling factor will be used to scale floating-point numbers. Mapped to elements on the ring of integers To support fractional operations in the ciphertext field; (2) Encryption primitives, given secret data The system splits it into two random partitions. and ,satisfy These two fragments are not only statistically evenly distributed, but also cannot be leaked individually. Any information; (3) Decrypt (reconstruct) primitives, given two fragments and The system uses modular addition operations. Restore the original data and divide by the scaling factor. Restore to floating-point number; (4) Ciphertext calculation support: The system supports local execution of ciphertext addition. ) and ciphertext constant multiplication ( ), and interactive ciphertext multiplication based on auxiliary data, where, It is a constant.

[0022] Step 1012: The model owner performs a secret addition split on the pre-trained Transformer model parameter set to generate two parameter fragments, which are then sent to two computing servers through a secure channel to achieve secret deployment of the model parameters.

[0023] Specifically, the Model Owner (MO) splits the pre-trained Transformer model parameter set into two statistically uniformly distributed parameter slices, which are then sent to two separate computing servers. and .

[0024] Parameter set Including weight matrix and bias vector The model owner deploys the pre-trained Transformer model parameter set to the server, as follows: (1) The model owner first multiplies all floating-point model parameters by a scaling factor. Round down to the nearest integer and convert to a ring. Integer tensors on; (2) The model owner generates a random tensor with the same shape as the parameter tensor. As the first slice Then calculate the second fragment. ; (3) The model owner sends the first fragment through a secure channel. Send to the first computing server ,Will Send to the second computing server At this point, the model parameters are stored in a distributed manner in the form of secret shards, and no single server can know the details of the model.

[0025] Step 1013: The cryptographic service provider deploys TEE code on two computing servers respectively; the computing servers and the corresponding TEE code negotiate and share a random number seed through a secure channel.

[0026] The cryptographic service provider (SP) deployed TEE code with side-channel protection on two computing servers. , Computing server , And the corresponding TEEs negotiate and share a random number seed through a secure channel (e.g.) , In the subsequent online phase, each party generates the auxiliary random mask required for computation synchronously both locally and within the TEE based on these seeds, without the need for real-time distribution of auxiliary data over the network.

[0027] Deploying TEE code specifically includes the following steps: (1) TEE security hardening and environment initialization: The cryptographic service provider initializes the Trusted Execution Environment (TEE, such as Intel SGX) on the computing server and uses Remote Attestation to verify the integrity of the enclave code. The enclave integrates a side-channel defense mechanism to ensure that, under the assumption of a semi-honest TEE, sensitive data during the computing process will not be leaked to the host machine through the side channel.

[0028] (2) PRF seed and key security negotiation, cryptographic service provider assists in computing server , and its corresponding TEE ( , The pseudo-random function (PRF) seed and key are pre-set or shared through secure channel negotiation. Specifically, this includes: (a) a cross-domain symmetric auxiliary seed; and Negotiating cross-domain seeds ,at the same time and Negotiating cross-domain seeds These two sets of seeds are specifically designed to support two independent directions in a parallel pipeline architecture with decoupled send and receive operations. and One-way basic cryptographic protocols on ) such as secure multiplication Three-input safe multiplication Unintentionally DReLU (a) Synchronously generate core blinding masks, thus perfectly supporting bidirectional computation overlap; (b) Computation server and Negotiating a peer synchronization seed specifically for generating the additional random mask required for the one-way underlying protocol. (c) Nonlinear local sinusoidal fitting protocol for Transformer Each party negotiates its own exclusive key. (for) and )and (for) and ).

[0029] (3) Local mask generation on demand: Using the agreed-upon seed and key, each participant calculates the auxiliary random number sequence required for the current protocol locally on demand using the PRF function before the offline phase or before executing a specific online protocol. For example, during the execution of... Before the agreement, utilize Extract the required set of variables ; (4) The TEE kernel-assisted peer preprocessing mechanism is designed for complex nonlinear local sine fitting protocols. This invention achieves an ultimate fully symmetric and load-balancing mechanism through cross-key injection. Specifically, during the system initialization phase, given the number of truncated Fourier series terms... Fitting interval parameters Cryptographic service providers transmit keys through secure channels and , and Cross-sealing injection and Internally. During preprocessing execution, and Within their respective secure enclaves, all utilize and Synchronous independent generation and , and Within their respective secure enclaves, all utilize and Simultaneously and independently generate corresponding random numbers , , and Thus, without any external communication intervention, each component independently reconstructs the complete plaintext mask. Then, the system will The nonlinear computation task involving a series of terms is performed by a completely symmetric decomposition: for the first half of the terms... Depend on Responsible, Internally, secure nonlinear result calculations and piecewise separation are performed to obtain the corresponding sinusoidal fundamental parameters. Sum and cosine fundamental parameters Its precise correction formula is: ,as well as The second half of the item directly corresponds to the generation of random numbers. and After the calculation is completed, Only to the local host machine Release its corrected slice vector and Regarding the second half of the item Depend on Responsible, Internally, safe nonlinear result calculations and piecewise separation are performed to obtain the corresponding sinusoidal fundamental parameters. Sum and cosine fundamental parameters Its precise correction formula is: ,as well as The first half of the item directly corresponds to the randomly generated number to be filled in. , After the calculation is completed, Only to the local host machine Release its corrected slice vector and This mechanism of "cross-injection dual-end reconstruction + task half-symmetric stripping" not only makes... and Random number injection and alignment can be completed through zero cross-network communication, which completely solves the serious bandwidth bottleneck caused by the online distribution of massive auxiliary data in the traditional MPC framework. It also perfectly distributes the heavy trigonometric function computing power consumption inside the TEE, realizing a truly fully symmetrical parallel architecture.

[0030] The data to be inferred is text or image data. The user preprocesses the text or image data to be inferred to a length of [length missing]. Token index sequence in integer form Users generate locally and Random integer tensors of the same dimension As the first slice And calculate the second fragment. Subsequently, the user will Send to the first computing server ,Will Send to the second computing server .

[0031] Step 102: Both computing servers execute a secure embedding protocol to convert the index fragment into a ciphertext embedding vector fragment. This ciphertext embedding vector fragment is then used as the initial hidden state input for the first round of the Transformer layer. Based on a predefined protocol, N rounds of Transformer layer iterative computation are performed to obtain the final ciphertext output. The predefined protocols include: one-way secure scalar multiplication protocol, secure matrix multiplication protocol, one-way three-input secure multiplication protocol, unintentional DReLU protocol, unintentional ReLU protocol, basic hybrid approximation fitting protocol, hybrid approximation fitting protocol, enhanced bounded function hybrid approximation fitting protocol, enhanced exponential function hybrid approximation fitting protocol, enhanced hybrid approximation fitting protocol, secure nonlinear activation protocol, secure layer normalization protocol, and secure maximum exponential normalization protocol.

[0032] Transformer model evaluation is performed when a user initiates an inference request. During this phase, the system follows the semi-honest TEE assumption, meaning that any data entering the TEE must undergo random number blinding. Both computation servers collaborate to perform N layers of Transformer module computations on the input data without recovering the plaintext data. Specifically, this is done to avoid leaking the index. Under the premise of extracting vectors, both parties calculate the server from the cryptographic service provider ( Obtain (or pre-generate during the offline phase) a set of auxiliary data pairs. .in It is a mask of random integers uniformly distributed within the vocabulary size. The position index is One-hot encoded vector (i.e., the first) (Bit 1 is used, the rest are 0); both parties' servers collaborate to calculate the blinded offset. The specific calculation formula is as follows: ,in, This is the vocabulary size. After calculation, the two computing servers exchange and reconstruct the plaintext offset. .because Randomness, public The real index will not be leaked. Information; both parties calculate the server's use of plaintext offset Slice the random One-Hot vectors held Perform a circular left shift operation. The shifted vector In fact, it corresponds to the real index. One-Hot vector (i.e. Both sides computed the server's use of the shifted ciphertext One-Hot vector. And the encrypted embedding weight matrix stored locally on the server Performing safe matrix multiplication Use this as the initial hidden state ciphertext input for the first round of the Transformer layer. .

[0033] In one possible implementation, performing N rounds of Transformer layer iterative computation includes: Step 201: Based on the initial hidden state input of the current layer, and by invoking the secure matrix multiplication protocol and the secure maximum exponent normalization protocol, obtain the attention output of the current layer.

[0034] In one possible implementation, obtaining the attention output of the current layer includes: Step 2011: Based on the initial hidden state input of the current layer, call the secure matrix multiplication protocol to perform the ciphertext linear transformation corresponding to the query matrix Q, key matrix K and value matrix V respectively, and obtain the query features, key features and value features.

[0035] Step 2012: Invoke the secure matrix multiplication protocol to calculate the matrix product of the query feature and the key feature to obtain the attention score matrix.

[0036] Step 2013: Invoke the secure maximum exponent normalization protocol to perform secure Softmax calculation on the attention score matrix to obtain the attention weight matrix.

[0037] The step of performing a secure Softmax calculation on the attention score matrix to obtain the attention weight matrix includes: The enhanced exponential function hybrid approximation fitting protocol is invoked to calculate the exponential function; The enhanced hybrid approximation fitting protocol is invoked to calculate the reciprocal, and the one-way secure scalar multiplication protocol is invoked to complete the normalization calculation, thereby obtaining the attention weight matrix; During the execution of Newton iteration, the enhanced hybrid approximation fitting protocol calls the unidirectional three-input safe multiplication protocol to complete three multiplication calculations.

[0038] Step 2014: Invoke the secure matrix multiplication protocol to calculate the matrix product of the attention weight matrix and the value feature to obtain the attention output of the current layer.

[0039] The secure matrix multiplication protocol includes: The two computing servers first randomly blind the input matrix by partitioning it, and then exchange the blinding results to reconstruct the matrix difference. The matrix product correction calculation is completed based on matrix difference and matrix multiplication auxiliary random numbers generated by TEE, so as to restore the secret shared fragment of the matrix multiplication result locally.

[0040] Obtain projection slices for each attention head. Both computing servers utilize offline-obtained Beaver multiplication triples to perform a safe linear transformation, converting the input... Secure matrix multiplication is performed with the query, key, and value weight matrices of this layer respectively to obtain the encrypted projection fragments. , , The expression is as follows: ; ; ; in, It is a one-way secure matrix multiplication protocol. For attention head Query weight matrix sharding, For attention head The key weight matrix is ​​partitioned. For attention head The value weight matrix is ​​segmented.

[0041] Attention score calculation, both parties' calculation servers will With the transposed key matrix Collaboratively perform secure matrix multiplication and multiply by a public constant of the scaling factor. ( (For feature dimensions), obtain the unnormalized attention score matrix slices. .

[0042] Security attention weighting assessment, both parties calculate the attention score matrix on the server. Collaborative invocation of a predefined security maximum exponent normalization protocol In the ciphertext domain, a softmax function is fitted to prevent overflow, and the normalized attention weight ciphertext matrix is ​​obtained. .

[0043] Attention output calculation involves both parties using the attention weight ciphertext matrix. and Perform secure matrix multiplication on the projected ciphertext to obtain the output of a single header. Subsequently, The results of the size are concatenated according to the feature dimensions and then compared with the output projection matrix. Perform secure matrix multiplication to obtain the multi-head attention ciphertext output of this layer. .

[0044] Step 202: Based on the attention output of the current layer, the two computing servers adopt residual connection and call the security layer normalization protocol and the enhanced hybrid approximation fitting protocol to obtain the first intermediate ciphertext feature; the first intermediate ciphertext feature is the normalized intermediate ciphertext feature.

[0045] In one possible implementation, obtaining the first intermediate ciphertext feature includes: Step 2021: Add the attention output of the current layer to the initial hidden state input to obtain the first residual ciphertext feature.

[0046] Specifically, both computing servers input the initial hidden state of this layer locally. The attention-ciphertext output obtained in the preceding steps By directly adding them together, the first residual ciphertext feature can be obtained. This process is performed entirely locally, without requiring network communication.

[0047] Step 2022: Invoke the security layer normalization protocol to calculate the mean, variance, and reciprocal of the standard deviation of the first residual ciphertext feature to obtain the first intermediate ciphertext feature; the first intermediate ciphertext feature is the normalized feature.

[0048] The step of invoking the security layer normalization protocol to calculate the mean, variance, and reciprocal of the standard deviation of the first residual ciphertext features includes: The enhanced hybrid approximation fitting protocol is invoked to complete the approximate calculation of the reciprocal of the square root; The normalized multiplication calculation is completed by invoking the one-way secure scalar multiplication protocol.

[0049] The first layer of normalization calculation, specifically, involves both computing servers calculating the characteristics of the summed residuals, i.e., the first residual ciphertext characteristics. Collaborative invocation of a predefined security layer normalization protocol The feature dimension variance is calculated using an enhanced hybrid approximation fitting protocol, combined with a learnable affine parameter scaling factor. With translation factor Finally, the first-stage normalized intermediate ciphertext feature is output, which is the first intermediate ciphertext feature. .

[0050] Step 203: For the first intermediate ciphertext feature, execute the secure matrix multiplication protocol, the secure nonlinear activation protocol, and the secure matrix multiplication protocol in sequence to obtain the second intermediate ciphertext feature; the second intermediate ciphertext feature is an intermediate ciphertext feature processed by a feedforward network and nonlinear activation.

[0051] In one possible implementation, obtaining the second intermediate ciphertext feature includes: Step 2031: Perform the secure matrix multiplication protocol on the first intermediate ciphertext feature and the first layer weight matrix of the feedforward network, and superimpose the bias to obtain the upgraded feature.

[0052] This step is the first layer of secure linear transformation. Specifically, both computing servers use Beaver triples to transform the first intermediate ciphertext features. With the first layer weight matrix of the feedforward network Perform secure matrix multiplication and add bias locally. Obtain the upgraded feature pieces .

[0053] Step 2032: Invoke the secure nonlinear activation protocol to perform nonlinear activation calculation and obtain activation output; the invocation of the secure nonlinear activation protocol to perform nonlinear activation calculation includes: The nonlinear function approximation calculation is completed by calling either the enhanced bounded function hybrid approximation fitting protocol or the enhanced exponential function hybrid approximation fitting protocol according to the activation function type.

[0054] Specifically, the computing servers of both parties target the dimensionality-upgrading feature. Collaborative invocation of a predefined secure nonlinear activation protocol ( or By utilizing interval determination and polynomial / sine hybrid fitting techniques, the ciphertext features after nonlinear activation are calculated with high precision. Here we adopt Let's take an example to illustrate. The protocol incorporates one-way... One-way unintentional DReLU ( ) and TEE-assisted local sine function ( ).

[0055] Step 2033: For the activated output, call the secure matrix multiplication protocol to complete the calculation of the second linear layer of the feedforward network and obtain the second intermediate ciphertext feature.

[0056] This step is a second-level secure linear transformation; specifically, both computing servers will activate the feature. With the second layer weight matrix Perform secure matrix multiplication and add biases. This is then mapped back to the original feature dimension to obtain the final ciphertext output calculated by the feedforward network, which is the second intermediate ciphertext feature. .

[0057] Step 204: Based on the second intermediate ciphertext feature, the two computing servers use a residual connection to execute the security layer normalization protocol and obtain the final ciphertext output of the current layer.

[0058] In one possible implementation, obtaining the final ciphertext output of the current layer includes: Step 2041: Perform ciphertext addition on the first residual ciphertext feature and the second intermediate ciphertext feature to obtain the second residual feature.

[0059] Specifically, both computing servers locally compute the first residual ciphertext feature. The second intermediate ciphertext feature output by the feedforward network Directly performing ciphertext addition yields the second residual superposition feature, i.e., the second residual feature. .

[0060] Step 2042: Invoke the security layer normalization protocol to calculate the mean, variance, and reciprocal of the standard deviation of the second residual feature, and obtain the final ciphertext output of the current Transformer layer; The step of invoking the security layer normalization protocol to calculate the mean, variance, and reciprocal of the standard deviation of the second residual feature includes: The enhanced hybrid approximation fitting protocol is invoked to perform the approximate calculation of the reciprocal of the square root, and the one-way safe scalar multiplication protocol is invoked to perform the normalization calculation.

[0061] The second-level normalization calculation, specifically, involves both computing servers targeting the second residual feature. Combined with the second set of learnable affine parameters of this layer and The security layer normalization protocol is invoked again in collaboration. Get the current number The final ciphertext output of the layer .

[0062] Step 205: Determine if the current layer is the last layer. If not, use the final ciphertext output of the current layer as the initial hidden state input for the next Transformer layer and continue execution. If so, use the final ciphertext output of the current layer as the final ciphertext inference result and output it.

[0063] Specifically, determine the Transformer layer number of the current iteration. Compared with the preset total number of model layers If the relationship, Then the current output features Automatically used as the next round (i.e., the first round) The input from the layer is returned to continue the evaluation of the safe multi-head attention mechanism; if Then the loop iteration calculation of the Transformer Block ends, and the finally obtained ciphertext output features are... Proceed to step 103 to prepare for the final result blinding and client response.

[0064] Before evaluation using the Transformer model, this embodiment of the invention designs five core privacy protection protocols, six types of function-based hybrid fitting protocols, six types of function-wide-range input hybrid fitting protocols, and three nonlinear operators as the basic components for constructing the upper-level Transformer operators.

[0065] I. Five Core Privacy Protection Agreements One-way secure scalar multiplication protocol ( ) In one possible implementation, the one-way secure scalar multiplication protocol includes: Both computing servers use local secret fragments and random mask fragments generated by TEE to calculate the input blinding difference; Both computing servers complete plaintext reconstruction by exchanging blinding differences, and obtain the reconstructed plaintext. The product correction term is calculated based on the reconstructed plaintext and the multiplication auxiliary data output by the TEE to obtain the secret sharing fragment corresponding to the product result; the product result is determined based on the product correction term.

[0066] This protocol is a core component for element-wise multiplication in nonlinear modules of the processing system (such as GeLU activation, Softmax, etc.), and is used for unidirectional computation in encrypted state. Its execution process is as follows: (a) and Using cross-domain seeds Generate scalar mask ; and Using synchronization seeds Generate scalar mask .

[0067] (b) Calculate the blinding value and And transmit to a secure enclave ; Combine these values ​​with the mask to calculate And return it unidirectionally to .

[0068] (c) Locally calculate the blinding value and And send unidirectionally . Multiply the received plaintext blinding values ​​to obtain and combined The return value is used to calculate its own product result locally. .at the same time, Calculate directly on the local machine using the held mask. The whole process No need to Return any intermediate state One-way secure matrix multiplication protocol ( ) Specifically, to address the dense linear mapping requirements of the Multi-Head Attention (MHSA) mechanism and Feedforward Network (FFN) in the Transformer architecture, we designed a dedicated matrix multiplication protocol for computation. Its execution process is as follows: (a) and Using cross-domain seeds Generate the first scalar mask ; and Using synchronization seeds Generate the second scalar mask .

[0069] (b) Calculate the blinding value and And transmit to a secure enclave ; Combine these values ​​with the mask to calculate And return it unidirectionally to .

[0070] (c) Locally calculate the blinding value and And send unidirectionally . Multiply the received plaintext blinding values ​​to obtain and combined The return value is used to calculate its own product result locally. .at the same time, Calculate directly on the local machine using the held mask. The whole process No need to Return any intermediate state.

[0071] One-way three-input secure multiplication protocol ( ) In one possible implementation, the one-way three-input secure multiplication protocol includes: Both computing servers perform random blinding on the three input secret fragments respectively, and exchange the blinding values ​​to restore the corresponding blinding difference; The three-input product correction calculation is completed based on the blinded difference of the recovery and the three-input auxiliary multiplication data output by TEE, and the secret sharing result corresponding to the three-input product is obtained.

[0072] Specifically, this protocol aims to complete the multiplication of three ciphertext variables within a single round of communication, enabling unidirectional computation in ciphertext mode. Its execution process is as follows: (a) and Using cross-domain seeds Generate a third scalar mask , and Using synchronization seeds Generate the fourth scalar mask .

[0073] (b) Second computing server Calculate the blinding value , , And transmit to a secure enclave ; Combine these values ​​with the mask to calculate the second combined value. .

[0074] (c) Locally calculate the blinding value And send unidirectionally . Calculated using the received plaintext blinding value , Transmit three plaintext blind values ​​to the secure enclave. , Combine these values ​​with the mask to calculate the third combined value. and mix items Its one-way return to .

[0075] (d) Combination The return value calculates its own product result fragment locally, which is the fourth product result fragment. Meanwhile, computing servers Calculate directly locally using the held mask and Therefore, we can conclude that .

[0076] One-way unintentional DReLU protocol ( ) In one possible implementation, the unintentional DReLU protocol includes: The two computing servers first use a random mask to blind the input data, and then exchange fragments to restore the blinded values; Based on the comparison of the derivatives calculated independently of the public blinding value and the local random mask, the secret sharing result corresponding to the DReLU function is output.

[0077] This protocol is used to prevent the source data from being leaked. Under the premise of safely extracting its sign bit (i.e., determining) Based on a predetermined range constraint (assuming the data length is...). , ,and Its execution process is as follows: (a) Generate random numbers , and Shared random numbers and based on Positive and negative setting flip flag . and Shared mask .

[0078] (b) Calculate the blinding value Transmit to a secure enclave ; Combine these values ​​with the mask to calculate and return it unidirectionally to .

[0079] (c) Locally calculate the blinding value One-way sending . Combine the two parts to restore ,because If the signs are the same and there are range constraints, determine... It can be obtained without damage The judgment result .

[0080] (d) Will One-way transmission to , According to the set flag position Calculate the initial fragmentation and return it to :like ,return ;like If so, perform a reverse correction and return. . After receiving, subtract the mask locally. To obtain the final sharding results: .at the same time, The other half of the result fragment is calculated directly using the locally held mask. The exact corresponding formula is: If ,but ;like ,but .

[0081] One-way unintentional ReLU protocol ( ) In one possible implementation, the unintentional ReLU protocol includes: Both computing servers first invoke the unintentional DReLU protocol to obtain the input symbol bits; Secure multiplication is performed based on the input symbol bits and the input secret slice, retaining only the data corresponding to the positive interval and masking the data in the negative interval, and outputting the secret shared slice of the ReLU function result.

[0082] This protocol is used to compute the standard modified linear unit (ReLU) activation function, i.e., the output within the safety region. Due to the present invention The result of the protocol output It can be used directly as an arithmetic result (without the need for expensive Boolean-to-arithmetic (B2A) conversion), therefore The execution process is extremely simple, and the specific process is as follows: (a) Both computing servers first make a direct call. Protocol, obtaining judgment Is it greater than Arithmetic partitioning .

[0083] (b) The servers of both parties then invoke the basic secure multiplication protocol once. Calculate the product of the judgment result and the original feature input. The final output This is a safe ReLU result partitioning.

[0084] II. Six Types of Functions Basic Mixed Fitting Protocol Basic Hybrid Approximation Fitting Protocol ( ) This protocol employs a linear combination of polynomials and truncated Fourier series (sine functions) to achieve high-precision fitting of nonlinear functions (such as sigmoid, tanh, erf, exp) within a certain range within a single round of communication, given the number of truncated Fourier series terms. Fitting interval parameters and coefficient Its execution process is as follows: (a) Both parties compute the alignment mask fragment generated by the TEE directly from the server. extract ;at the same time extract .

[0085] (b) Online phase, and Each input fragment is obfuscated using its own blinding mask, and the results are calculated. Both parties send the blinded fragment to each other and reconstruct the plaintext blinded value by performing a modulo-add operation locally. .

[0086] (c) Both computing servers use random mask fragmentation and the pre-calculated optimal fitting coefficients ,for The calculation is as follows: ; In the formula, , , These represent the safe computation result slices corresponding to the linear, quadratic, and cubic terms of the polynomial, respectively, which are used for subsequent linear combination with the pre-calculated polynomial coefficients.

[0087] (d) For each Both parties based on plain text Extract the sinusoidal basis terms independently locally. Sum of cosine basis terms Finally, combining the modified mask fragments held by each party. Fit coefficients The final fitting result is obtained locally through multiply-accumulate operations. .

[0088] Basic Mixed Approximation Fit protocol( ) This protocol employs a linear combination of polynomials and truncated Fourier series (sine function), followed by multiple rounds of Newton iteration to improve fitting accuracy. The protocol has two implementation modes: "fitting the reciprocal of the square root" and "fitting the reciprocal." The execution flow for "fitting the reciprocal of the square root" is as follows: (a) Both computing servers compute a hybrid approximation fitting protocol ,in To indicate the participating parties The input data is secretly shared in fragments. The secret shared slice represents the initial approximation of the objective function, which serves as the initial value for subsequent Newton iterations.

[0089] (b) Both parties' servers linearly scale the input secret fragment according to a preset scaling factor to obtain the scaled input secret shared fragment. .

[0090] (c) To construct the iterative update term, both computing servers first invoke the three-party secure multiplication protocol. Calculate the current approximate value Secret Sharing of the Third Power .

[0091] (d) Both parties calculate the server's calculation .

[0092] (e) Both computing servers begin the second round of Newton iterations to calculate the current approximation. Secret Sharing of the Third Power .

[0093] (f) Both parties calculate the final result obtained by the server. .

[0094] The execution flow for fitting the reciprocal is as follows: (a) Both computing servers compute a hybrid approximation fitting protocol .

[0095] (b) Both computing servers begin the first round of Newton iterations and obtain the first result. .

[0096] (c) Both computing servers begin the second round of Newton iterations and obtain the second result. .

[0097] (d) Both computing servers begin the third round of Newton iterations and obtain the final result. .

[0098] III. Wide-range input fusion fitting protocol for six types of functions Enhanced bounded function hybrid approximation fitting protocol ( , , ) In one possible implementation, the enhanced bounded function hybrid approximation fitting protocol includes: Both computing servers invoke the basic hybrid approximation fitting protocol to complete the initial approximation calculation of the bounded function; Based on the initial approximate calculation results of the bounded function and the safe interval determination protocol, the input interval is identified, and the corresponding fitting parameters are selected for local correction according to different intervals, and the secret sharing result of the bounded function is output.

[0099] Specifically, this protocol targets nonlinear activation functions with horizontal asymptotes (such as sigmoid, tanh, and erf). Leveraging their property of tending to a constant beyond a specific interval, it uses boundary comparison with a constant multiplexing to eliminate wide-area approximation errors while reducing communication overhead. Its execution flow is as follows: (a) Both parties' calculation servers calculate based on pre-set thresholds. Calculate in parallel with pre-calculated optimal fitting coefficients .

[0100] (b) The results are obtained by parallel computation of the two computing servers. in and is the boundary constant of the function.

[0101] Enhanced exponential function hybrid approximation fitting protocol ( ) In one possible implementation, the enhanced exponential function hybrid approximation fitting protocol includes: Both computing servers perform interval reduction on the input data and use a basic hybrid approximation fitting protocol to calculate the approximate value of the exponential function within the reduction interval; By combining the approximate value of the exponential function within the reduction interval with the exponential reconstruction relationship, the complete exponential calculation result is recovered, and the secret sharing result of the exponential function is obtained.

[0102] This protocol targets unbounded and exponentially growing functions (such as exp). Due to their extremely steep curves and the inability to utilize asymptotic constants, it employs a piecewise hybrid fitting strategy in the encrypted domain to eliminate approximation errors and maintain high accuracy over a very wide dynamic range. Its execution flow is as follows: (a) Both parties' calculation servers calculate based on pre-set thresholds. Parallel calculation with pre-calculated optimal fitting coefficients for different ranges .

[0103] (b) The results are obtained by parallel computation of the two computing servers. .

[0104] Enhanced hybrid approximation protocol( / ) In one possible implementation, the enhanced hybrid approximation fitting protocol includes: Both computing servers automatically invoke the corresponding enhanced approximation protocol to complete the local function calculation based on the type of the objective function; The results are fused and corrected based on the calculation results of local functions and Newton's iterative operations, and a secure calculation output secretly shares the approximate result.

[0105] This protocol innovatively integrates encrypted input scaling with a piecewise hybrid fitting strategy (embedded Newton iteration) to support accurate approximation over an ultra-wide dynamic range. It addresses nonlinear functions (such as reciprocals and square root reciprocals) that are extremely steep near zero (possessing singularities) and whose initial estimations are prone to numerical inflation over a large dynamic range. The protocol has two implementation modes: "fitting the square root reciprocal" and "fitting the reciprocal." The execution flow for "fitting the square root reciprocal" is as follows: (a) Both parties calculate the scaling factor of the stable region based on the pre-set numerical values ​​on the server. With interval threshold ( , After linearly mapping the input ciphertext, the optimal fit results for candidate segments with high and low identifiers are calculated in parallel: .

[0106] (b) The results are obtained by parallel computation of the two computing servers. .

[0107] (c) Both computing servers calculate the final result locally. .

[0108] The execution flow for fitting the reciprocal is as follows: (a) Both parties calculate the scaling factor of the stable region based on the pre-set numerical values ​​on the server. With interval threshold ( , After linearly mapping the input ciphertext, the optimal fit results for candidate segments with high and low identifiers are calculated in parallel: .

[0109] (b) The results are obtained by parallel computation of the two computing servers. .

[0110] (c) Both computing servers calculate the final result locally. .

[0111] IV. Three Major Nonlinear Operators Secure nonlinear activation protocol ( ) This protocol integrates interval determination and piecewise approximation strategies, and is specifically designed to handle the GeLU activation function. Its specific execution flow is as follows: (a) Both parties' calculation servers calculate based on pre-set thresholds. Calculate in parallel with pre-calculated optimal fitting coefficients .

[0112] (b) The results are obtained by parallel computation of the two computing servers. .

[0113] Security layer normalization protocol ( ) This protocol is specifically designed for feature normalization operations with extremely large dynamic ranges. To ensure numerical stability and computational efficiency under a wide range of inputs, this protocol combines an optimized comparison protocol with a piecewise enhanced reciprocal square root protocol. The specific execution flow includes the following sub-steps: (a) Both computing servers compute the input fragments Perform addition locally to obtain the sum of the feature dimensions, and use a truncation protocol to obtain the mean slice. Variance partitioning is calculated using mean partitioning. ; (b) Both parties' servers calculate based on pre-set thresholds. The reciprocal of the square root of variance piecewise variance is obtained by parallel calculation of the pre-calculated optimal fitting coefficients for different ranges. ; (c) Both parties calculate the standardized intermediate value on the server. ; (d) Both computing servers perform affine transformations to obtain normalization. .

[0114] Maximum Security Index Normalization Protocol (MPI) ) This protocol employs a clipping-based alternative for the Softmax function, avoiding the overhead of global maximum search. The specific execution flow includes the following sub-steps: (a) Both parties' servers calculate based on pre-set thresholds. , parallel computing ; (b) Both parties' calculation servers will input Cut off and translate to obtain ; (c) Both parties calculate the value of the fitted exp function on the server. ; (d) Both parties calculate the sum of the fitted exp functions calculated by the server. ; (e) Both parties' calculation servers calculate based on pre-set thresholds. and multiple factors The reciprocal slices are obtained by parallel computation of the pre-calculated optimal fitting coefficients for different ranges. ; (f) Both parties calculate the final result value calculated by the server. ).

[0115] To clarify the concurrent topology of the basic privacy protection protocol of this invention and the generality of constructing upper-layer operators, the following system-level architecture and extension declarations are hereby made: Bidirectional peer-to-peer concurrent execution topology of the underlying protocol All of the above underlying one-way privacy protection protocols (including , , , and All of these (etc.) possess complete symmetry and role interchangeability in their physical architecture. In actual model evaluation and scheduling, the system data flow is not limited to a single "..." send, Instead of a single, opposite direction, it fully supports bidirectional concurrent execution. Specifically, the first computing server... Second computing server It can simultaneously act as both the "sender" and the "receiver" of a one-way protocol. The system can achieve this by splitting feature dimensions or attention heads. and At the same time, a unidirectional protocol data stream is initiated in the opposite direction to the other end. This bidirectional concurrency mechanism fully utilizes the uplink and downlink bandwidth of the full-duplex network, enabling a perfect bidirectional overlap between computation time and network transmission time, thereby macroscopically constructing a bidirectional parallel pipeline architecture that supports the overlap of computation and communication.

[0116] Arbitrariness and Generality of Implementation Methods for Upper-Level Composite Operators It should be noted that this system utilizes the aforementioned underlying protocols to construct the core composite operators in the Transformer architecture (such as...). , , The process (etc.) is not limited to specific mathematical fitting methods or iterative logic. The six basic privacy protection protocols mentioned above serve as general "secure computing atoms," supporting flexible scheduling and free recombination. In practical implementation, the system can use any mathematical approximation method, piecewise logic, or iterative algorithm (e.g., but not limited to: Taylor series expansion of arbitrary order, Chebyshev polynomial fitting, Newton-Raphson iteration with different bases, or other numerical approximation strategies for activation functions) to call the aforementioned underlying basic protocols, thereby achieving the required upper-level nonlinear composite operator evaluation. Any upper-level model evaluation process constructed based on the efficient combination of underlying protocols provided by this invention is equivalent to an equivalent substitution of this invention and falls within the patent protection scope of this invention.

[0117] Step 103: Both computing servers use a local random number mask to perform security blinding processing on the final ciphertext output to obtain blinded fragments, and send the blinded fragments to the user.

[0118] Specifically, both sides calculate the server using a locally random mask vector. The final evaluation result is fragmented, i.e., the encrypted output features. Perform safety blinding treatment The blinded results were then split into pieces. Send to the user client that initiated the request.

[0119] Step 104: After receiving the blinded fragment, the user performs a modulo addition operation on the blinded fragment locally to reconstruct the plaintext data and obtain the final plaintext inference decision.

[0120] After receiving the two fragments, the user performs a modulo-add operation locally to reconstruct the plaintext data, and then uses the scaling factor negotiated during the system initialization phase. Perform division decoding to map the data back to the original floating-point probability space and obtain the final plaintext inference decision.

[0121] Specifically, the user performs a modulo addition operation on the two received fragments to recover the obfuscated result. Then, by using the mask information acquired synchronously to remove interference terms, the integer ring representation of the true result is reconstructed. The user utilizes the scaling factor negotiated during the system initialization phase. Perform calculation This maps the data back from the ring of integers to the original floating-point probability space. Numerical data is used to obtain the final plaintext reasoning decision.

[0122] This invention also provides a privacy-preserving inference system for Transformer models based on a trusted execution environment, the system comprising: The model owner, the holder of the pre-trained Transformer model, wants to provide services to the outside world without disclosing the model parameters (weights and biases), and participate in the system initialization and model deployment process.

[0123] Cryptographic service provider: A trusted or semi-honest third-party auxiliary organization responsible for generating and distributing the offline auxiliary data required by the system, and responsible for the initialization of the Trusted Execution Environment (TEE) and key sealing injection, participating only in the offline preprocessing process.

[0124] Users, the holders of the data to be inferred, wish to obtain the inference results without disclosing the privacy of their input data, and participate in the model evaluation request and result response process.

[0125] The computing servers, including a first computing server and a second computing server, serve as the core computing entities of the system. They possess computing resources but are assumed to be semi-honest and non-colluding. Trusted execution environments are deployed on both servers. , As a security coprocessor, it is responsible for collaboratively performing Transformer inference tasks in encrypted form.

[0126] The functions of the model owner include: Model training and quantization: The model owner trains the Transformer model locally and maps the model parameters (weights and biases) from the floating-point domain to the integer ring domain to generate a set of fixed-point parameters; Parameter Secret Splitting: The model owner uses an additive secret sharing mechanism to split the fixed-point parameter set into two parameter fragments, ensuring that no single fragment contains valid information from the original model. Secure distribution: The model owner sends two parameter fragments to the first computing server and the second computing server respectively through a secure channel, thus completing the privacy protection deployment of the model.

[0127] The functions of a cryptography service provider include: Random Number and Key Generation: Cryptographic service providers use true random number generators to generate high-quality random seeds and proprietary encryption keys (such as...). ); Auxiliary data construction: Based on a random seed, the cryptographic service provider pre-generates auxiliary data fragments required to support multi-party secure computation in the offline stage, including Beaver triples for matrix multiplication, random tuples for optimized three-input multiplication, mask pairs for comparison protocols, and truncation parameters for hybrid approximation fitting, etc. Environment Deployment and Cross-Distribution: The cryptographic service provider pre-distributes the generated auxiliary data fragments to both parties' computing servers. More importantly, the cryptographic service provider is responsible for initializing the Trusted Execution Environment (TEE) on both servers and cross-sealing the exclusive key into the peer's TEE via a secure channel (i.e.,...). injection , injection This provides an underlying trust foundation for "local one-way algebra stripping" and "zero cross-network alignment" in the online inference stage.

[0128] User functions include: Data preprocessing and secret sharing: Users convert plaintext data to be reasoned into integer indices and split the input data into two secret fragments based on an additive secret sharing mechanism; In the fragmented data transmission, the user sends two fragments of input data to the first computing server and the second computing server respectively through a secure channel to initiate the inference request; The results are reconstructed and decoded. The user receives the result fragments returned from two computing servers, performs a modulo addition operation to eliminate random masks and reconstruct the plaintext data, and then divides by the scaling factor to decode and obtain the final Transformer inference result.

[0129] The specific functions of the computing server include: System initialization and deployment: The server negotiates and establishes integer ring parameters and secret sharing mechanism, receives model parameter fragments from the model owner and auxiliary data fragments from the cryptographic service provider, and completes inference preparation; Input encryption and secure embedding computation: The server collaboratively executes the secure embedding protocol, utilizes offline auxiliary masks to perform index blinding and shift reconstruction, and invokes a one-way secure matrix multiplication protocol ( ), extract the corresponding initial ciphertext embedding vector without restoring the plaintext index; Evaluation of secure multi-head attention mechanism: Server collaborative invocation of basic security matrix multiplication protocol ( Obtain the projection matrix (Q / K / V) of the multi-head attention and call the secure maximum exponential normalization protocol ( ). ), to complete the ciphertext calculation and output of the attention mechanism weights; Security layer normalization: The server, in conjunction with the encrypted residual connection, collaboratively invokes the security layer normalization protocol. The feature variance is calculated using the inverse square root protocol, and the intermediate ciphertext feature map is standardized. Secure feedforward networks and nonlinear activation: server cooperative invocation Protocol, and in conjunction with a secure activation function protocol (such as...) The GeLU activation function is evaluated with high accuracy in the ciphertext state using a hybrid approximation fitting technique, thus completing the ciphertext computation of the feedforward network. Results Blinding and Backhaul: The server generates or uses a pre-generated random mask to perform additive blinding on the final inference result fragments, and then sends the obfuscated result fragments to the user.

[0130] In summary, this invention proposes a framework based on a semi-honest TEE model. This method significantly reduces communication overhead and improves inference efficiency through a unidirectional communication protocol suite, while substantially weakening the trust assumptions of the TEE and fully protecting the privacy of user data and the server model. This provides a solution that balances security and ultimate performance for the practical implementation of large-scale privacy-preserving Transformer inference.

[0131] Specifically, in the inference service scenario, the TEE is formally modeled as a semi-honest participant. Under this model, it is acknowledged that the Enclave code provider may steal sensitive information through internal intermediate computation states. Based on this, this invention designs a general MPC protocol that can be integrated with the semi-honest TEE. By introducing a double random masking mechanism, the confidentiality of user input and model parameters is ensured, making it highly suitable for the real-world threat environment of private Transformer inference. A highly efficient one-way communication protocol is designed (breaking through the interaction round bottleneck): Addressing the network latency bottleneck caused by multiple rounds of bidirectional interaction in pure MPC schemes, this invention utilizes TEE assistance to design a novel MPC protocol for core arithmetic and nonlinear primitives (such as safe multiplication, three-input multiplication, and unintentional DReLU). These protocols innovatively achieve one-way communication, meaning that computation can be completed by transmitting data in only one direction, without multiple rounds of interaction. Compared to previous pure MPC and traditional TEE-assisted schemes, this concept fundamentally and significantly reduces communication round-trip overhead.

[0132] This invention innovatively integrates a Trusted Execution Environment (TEE) with multi-channel secure multi-party computation, constructing a bidirectional concurrent architecture with extremely low communication rounds and extreme performance. This solution achieves efficient and secure inference for large-scale Transformer deep neural networks (such as pre-trained models like BERT and GPT) while ensuring absolute privacy and security of user data and server model intellectual property. Its unique "cross-key sealing injection" and "local peer-to-peer reconstruction" mechanisms completely break through the bandwidth bottleneck of traditional MPC's online auxiliary data distribution, perfectly adapting to cloud-based outsourced computing scenarios involving "model owner - cryptographic service provider - computing server - user," and is particularly suitable for large-scale commercial deployment in low-bandwidth, high-latency real wide area network (WAN) environments.

[0133] This invention proposes an asynchronous overlapping optimization strategy. Based on the characteristics of unidirectional communication, this invention divides the overall computation task into multiple independent components and reorganizes the computation flow to support parallel execution. This design enables the computing server to synchronously perform local pre-computation during idle time in network data transmission, effectively hiding communication latency, reducing the actually observed communication overhead by nearly half, and significantly improving end-to-end efficiency.

[0134] This invention implements a privacy-preserving architecture that supports outsourced computation and multi-scenario deployment. The invention constructs a complete ecosystem logic including model owners, users, computation servers, and cryptographic service providers, supporting end-to-end encrypted computation. This architecture is highly flexible, adapting to low-bandwidth, high-latency scenarios such as wide area networks (WANs) through a one-way communication protocol, and also supporting load balancing by alternating computation roles. This effectively solves the privacy and performance imbalance problem in large-scale Transformer inference in cloud outsourcing scenarios.

[0135] This invention proposes a high-precision, low-communication-rounds nonlinear function evaluation method based on hybrid mathematical approximation. Addressing the issues of high communication overhead or insufficient accuracy in traditional MPC schemes using polynomial approximation, this invention innovatively designs a hybrid approximation fitting protocol. This combines low-order polynomials with truncated Fourier series (sine functions) to achieve [something] within a single round of communication. High-precision approximation of nonlinear functions.

[0136] This invention designs a series of enhanced hybrid fitting protocols for wide-range inputs: to overcome the approximation error and numerical inflation caused by the wide dynamic range characteristics in model inference, this invention deeply integrates ciphertext segmentation, dynamic scaling, and basic one-way operators. Specifically, this includes: designing protocols for bounded functions (such as Sigmoid). Protocol, utilizing Perform interval determination to achieve multiplexing of boundary constants and basic fitting terms; design for unbounded exponential functions (such as Exp). The protocol employs piecewise parallel fitting in the encrypted domain to suppress accuracy degradation; it is designed for steep functions with singularities. / The protocol is the first to integrate ciphertext value stable scaling with segmentation strategies.

[0137] This invention proposes a low-round-specific privacy computation protocol adapted to the Transformer architecture: For the GeLU / SiLU activation function, a computation strategy integrating interval determination and piecewise approximation is proposed, solving the high latency problem caused by the serial execution of comparison and computation in traditional schemes; For the Softmax mechanism, an alternative scheme based on numerical truncation is proposed, utilizing a security exponent and an enhanced reciprocal protocol to avoid the high-overhead global maximum search operation, effectively solving the numerical overflow problem and improving computational efficiency; For LayerNorm normalization, an enhanced reciprocal square root approximation protocol supporting a large dynamic range is proposed, ensuring the numerical stability of the model when dealing with different distribution characteristics.

[0138] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0139] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A privacy-preserving inference method for Transformer models based on a trusted execution environment, characterized in that, include: Users share a secret by adding the token index sequence of the data to be inferred, generating two index fragments, and sending them to the computing servers of both parties; Both computing servers execute a secure embedding protocol to convert the index fragments into ciphertext embedding vector fragments. These ciphertext embedding vector fragments are then used as the initial hidden state input for the first round of the Transformer layer. Based on a predefined protocol, N rounds of Transformer layer iterative computation are performed to obtain the final ciphertext output. The predefined protocols include: one-way secure scalar multiplication protocol, secure matrix multiplication protocol, one-way three-input secure multiplication protocol, unintentional DReLU protocol, unintentional ReLU protocol, basic hybrid approximation fitting protocol, hybrid approximation fitting protocol, enhanced bounded function hybrid approximation fitting protocol, enhanced exponential function hybrid approximation fitting protocol, enhanced hybrid approximation fitting protocol, secure nonlinear activation protocol, secure layer normalization protocol, and secure maximum exponential normalization protocol. Both computing servers use local random number masks to perform security blinding processing on the final ciphertext output to obtain blinded fragments, and send the blinded fragments to the user; After receiving the blinded data fragment, the user performs a modulo-addition operation on the blinded data fragment locally to reconstruct the plaintext data and obtain the final plaintext reasoning decision.

2. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 1, characterized in that, The execution of N rounds of Transformer layer iterative computation includes: Based on the initial hidden state input of the current layer, and by invoking the secure matrix multiplication protocol and the secure maximum exponent normalization protocol, the attention output of the current layer is obtained; Based on the attention output of the current layer, the two computing servers use residual connections and call the security layer normalization protocol and the enhanced hybrid approximation fitting protocol to obtain the first intermediate ciphertext feature; the first intermediate ciphertext feature is the normalized intermediate ciphertext feature. For the first intermediate ciphertext feature, the secure matrix multiplication protocol, the secure nonlinear activation protocol, and the secure matrix multiplication protocol are executed sequentially to obtain the second intermediate ciphertext feature; the second intermediate ciphertext feature is an intermediate ciphertext feature processed by a feedforward network and nonlinear activation. Based on the second intermediate ciphertext feature, the two computing servers use a residual connection to execute the security layer normalization protocol and obtain the final ciphertext output of the current layer. Determine if the current layer is the last layer. If not, use the final ciphertext output of the current layer as the initial hidden state input for the next Transformer layer and continue execution. If so, use the final ciphertext output of the current layer as the final ciphertext inference result and output it.

3. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 2, characterized in that, Obtaining the attention output of the current layer includes: Based on the initial hidden state input of the current layer, the secure matrix multiplication protocol is invoked to perform ciphertext linear transformations corresponding to the query matrix Q, key matrix K, and value matrix V, respectively, to obtain query features, key features, and value features. The secure matrix multiplication protocol is invoked to calculate the matrix product of the query feature and the key feature to obtain the attention score matrix; The security maximum exponential normalization protocol is invoked to perform security Softmax calculation on the attention score matrix to obtain the attention weight matrix; The step of performing a secure Softmax calculation on the attention score matrix to obtain the attention weight matrix includes: The enhanced exponential function hybrid approximation fitting protocol is invoked to calculate the exponential function; The enhanced hybrid approximation fitting protocol is invoked to calculate the reciprocal, and the one-way secure scalar multiplication protocol is invoked to complete the normalization calculation, thereby obtaining the attention weight matrix; During the execution of Newton's iteration, the enhanced hybrid approximation fitting protocol calls the unidirectional three-input safe multiplication protocol to complete three multiplication calculations. The secure matrix multiplication protocol is invoked to calculate the matrix product of the attention weight matrix and the value feature, thereby obtaining the attention output of the current layer; The secure matrix multiplication protocol includes: The two computing servers first randomly blind the input matrix by partitioning it, and then exchange the blinding results to reconstruct the matrix difference. The matrix product correction calculation is completed based on matrix difference and matrix multiplication auxiliary random numbers generated by TEE, so as to restore the secret shared fragment of the matrix multiplication result locally.

4. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 2, characterized in that, Obtaining the first intermediate ciphertext feature includes: The attention output of the current layer is added to the initial hidden state input to obtain the first residual ciphertext feature; The security layer normalization protocol is invoked to calculate the mean, variance, and reciprocal of the standard deviation of the first residual ciphertext feature to obtain the first intermediate ciphertext feature; the first intermediate ciphertext feature is the normalized feature. The step of invoking the security layer normalization protocol to calculate the mean, variance, and reciprocal of the standard deviation of the first residual ciphertext features includes: The enhanced hybrid approximation fitting protocol is invoked to complete the approximate calculation of the reciprocal of the square root; The normalized multiplication calculation is completed by invoking the one-way secure scalar multiplication protocol.

5. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 2, characterized in that, The process of obtaining the second intermediate ciphertext feature includes: The secure matrix multiplication protocol is applied to the first intermediate ciphertext feature and the first layer weight matrix of the feedforward network, and the bias is superimposed to obtain the upgraded feature. Invoking the secure nonlinear activation protocol to perform nonlinear activation calculations and obtain activation output; the invocation of the secure nonlinear activation protocol to perform nonlinear activation calculations includes: The enhanced bounded function hybrid approximation fitting protocol or the enhanced exponential function hybrid approximation fitting protocol is invoked according to the activation function type to complete the approximate calculation of the nonlinear function; For the activated output, the secure matrix multiplication protocol is invoked to complete the calculation of the second linear layer of the feedforward network, thereby obtaining the second intermediate ciphertext feature.

6. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 4, characterized in that, Obtaining the final ciphertext output of the current layer includes: Perform ciphertext addition on the first residual ciphertext feature and the second intermediate ciphertext feature to obtain the second residual feature; The security layer normalization protocol is invoked to calculate the mean, variance, and reciprocal of the standard deviation of the second residual feature, thereby obtaining the final ciphertext output of the current Transformer layer. The step of invoking the security layer normalization protocol to calculate the mean, variance, and reciprocal of the standard deviation of the second residual feature includes: The enhanced hybrid approximation fitting protocol is invoked to perform the approximate calculation of the reciprocal of the square root, and the one-way safe scalar multiplication protocol is invoked to perform the normalization calculation.

7. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 1, characterized in that, The enhanced bounded function hybrid approximation fitting protocol includes: Both computing servers invoke the basic hybrid approximation fitting protocol to complete the initial approximation calculation of the bounded function; The input interval is identified based on the initial approximate calculation results of the bounded function and the safe interval determination protocol. The corresponding fitting parameters are selected for local correction according to different intervals, and the secret sharing results of the bounded function are output. The enhanced exponential function hybrid approximation fitting protocol includes: Both computing servers perform interval reduction on the input data and use a basic hybrid approximation fitting protocol to calculate the approximate value of the exponential function within the reduction interval; Based on the approximate value of the exponential function within the reduction interval and the exponential reconstruction relationship, the complete exponential calculation result is recovered, and the secret sharing result of the exponential function is obtained. The enhanced hybrid approximation fitting protocol includes: Both computing servers automatically invoke the corresponding enhanced approximation protocol to complete the local function calculation based on the type of the objective function; The results are fused and corrected based on the calculation results of local functions and Newton's iterative operations, and the secure calculation output secretly shares the approximate result.

8. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 1, characterized in that, The one-way three-input secure multiplication protocol includes: Both computing servers perform random blinding on the three input secret fragments respectively, and exchange the blinding values ​​to restore the corresponding blinding difference; The three-input product correction calculation is completed based on the blinded difference of the recovery and the three-input auxiliary multiplication data output by TEE, and the secret sharing result corresponding to the three-input product is obtained. The one-way secure scalar multiplication protocol includes: Both computing servers use their local secret fragments and the random masked fragments generated by the TEE to calculate the input blinding difference; Both computing servers complete plaintext reconstruction by exchanging blinding differences, and obtain the reconstructed plaintext. The product correction term is calculated based on the reconstructed plaintext and the multiplication auxiliary data output by the TEE to obtain the secret sharing fragment corresponding to the product result; the product result is determined based on the product correction term.

9. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 1, characterized in that, The unintentional DReLU protocol includes: The two computing servers first use a random mask to blind the input data, and then exchange fragments to restore the blinded values; Based on the comparison of the derivatives calculated independently of the public blinding value and the local random mask, the secret sharing result corresponding to the DReLU function is output; The unintentional ReLU protocol includes: Both computing servers first invoke the unintentional DReLU protocol to obtain the input symbol bits; Secure multiplication is performed based on the input symbol bits and the input secret slice, retaining only the data corresponding to the positive interval and masking the data in the negative interval, and outputting the secret shared slice of the ReLU function result.

10. The privacy-preserving inference method for Transformer models based on a trusted execution environment according to claim 1, characterized in that, Before the user shares the addition secret of the token index sequence of the inference data, the following is also included: Users and servers initialize the encryption system and negotiate parameters to determine the arithmetic secret sharing mechanism in order to complete the encryption system configuration; The model owner performs a secret addition split on the pre-trained Transformer model parameter set, generating two parameter fragments, which are then sent to two computing servers via a secure channel to achieve secret deployment of the model parameters. The cryptographic service provider deploys TEE code on two computing servers respectively; the computing servers and the corresponding TEE code negotiate and share a random number seed through a secure channel.