Method, device and equipment for model trusted reasoning and storage medium
By using encryption and decryption mechanisms between multiple trusted execution environments in a computing cluster, the problem of communication security in the inference phase of machine learning models in cloud environments is solved, secure and reliable data transmission is achieved, and the security boundary of the cloud environment is expanded.
Patent Information
- Application Number
- CN202510901314.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
When deploying machine learning models in a cloud environment, how can we ensure that the communication process between different inference stages of the machine learning model is secure and reliable? In particular, how can we prevent data from being exposed outside the trusted execution environment and breaking the trust boundary during communication between multiple servers or server clusters?
By providing multiple trusted execution environments in a computing cluster and assigning corresponding first and second keys to them, intermediate processing results are encrypted in the first trusted execution environment and decrypted in the second trusted execution environment, ensuring that encrypted data is transmitted between trusted execution environments, thereby achieving secure data transmission.
It ensures that the communication process between different stages of the machine learning model's reasoning process is secure and reliable, expands the security boundary of the cloud environment, and improves the overall security of the cloud environment.
Smart Images

Figure CN120806147A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to methods, apparatuses, devices, computer-readable storage media, and computer program products for model trusted inference. BACKGROUND
[0002] With the integration of cloud computing technology and machine learning technology, more and more providers of machine learning models choose to deploy machine learning models in a cloud environment. The cloud environment can provide sufficient computing resources to perform inference services of machine learning models, which is conducive to improving the inference efficiency and response speed of machine learning models. However, the cloud environment faces certain data security problems. SUMMARY
[0003] In a first aspect of the present disclosure, a method for model trusted inference is provided. The method comprises: performing, in a first trusted execution environment of a computing cluster for model inference, a first stage processing of a machine learning model on a model input to generate an intermediate processing result corresponding to the model input, the first trusted execution environment being one of a plurality of trusted execution environments of the computing cluster, each trusted execution environment of the plurality of trusted execution environments holding at least one of: a first key of the computing cluster or a second key corresponding to the first key; encrypting, in the first trusted execution environment, the intermediate processing result using the first key to obtain encrypted data; providing the encrypted data from the first trusted execution environment to a second trusted execution environment of the plurality of trusted execution environments; decrypting, in the second trusted execution environment, the encrypted data using the second key to obtain the intermediate processing result; and performing, in the second trusted execution environment, a second stage processing of the machine learning model on the intermediate processing result to generate a model output corresponding to the model input.
[0004] In a second aspect of the disclosure, an apparatus for model trusted inference is provided. The apparatus comprises: a first processing module configured to perform, in a first trusted execution environment of a computing cluster for model inference, a first stage processing of a machine learning model on a model input to generate an intermediate processing result corresponding to the model input, the first trusted execution environment being one of a plurality of trusted execution environments of the computing cluster, each of the plurality of trusted execution environments holding at least one of: a first key of the computing cluster or a second key corresponding to the first key; an encryption module configured to encrypt, in the first trusted execution environment, the intermediate processing result with the first key to obtain encrypted data; a providing module configured to provide the encrypted data from the first trusted execution environment to a second trusted execution environment of the plurality of trusted execution environments; a decryption module configured to decrypt, in the second trusted execution environment, the encrypted data with the second key to obtain the intermediate processing result; and a second processing module configured to perform, in the second trusted execution environment, a second stage processing of the machine learning model on the intermediate processing result to generate a model output corresponding to the model input.
[0005] In a third aspect of the disclosure, an electronic device is provided. The device comprises at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the disclosure, a computer program product is provided, comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to the first aspect of the disclosure.
[0008] It is to be understood that the particulars shown herein are by way of example and for purposes of illustrative discussion of the embodiments of the present disclosure only and are not intended to limit the scope of the present disclosure to the particular embodiment illustrated. Other objectives, features and aspects of the disclosure will become apparent from the following description taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, aspects and advantages of embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters designate like elements in which:
[0010] Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure can be implemented is shown;
[0011] Figure 2 A schematic diagram showing an example architecture for model trustworthy inference is shown in accordance with some embodiments of the present disclosure;
[0012] Figure 3 A flowchart showing a process for model trustworthy inference is shown in accordance with some embodiments of the present disclosure;
[0013] Figure 4 A schematic block diagram showing an example apparatus for model trustworthy inference is shown in accordance with some embodiments of the present disclosure; and
[0014] Figure 5 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It will be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0016] In the description of embodiments of the present disclosure, the term "includes" and its derivatives, such as "including," should be understood in an open, inclusive sense, that is, "including, but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicit or implicit definitions can also be included below.
[0017] In this document, unless explicitly stated otherwise, performing a step "in response to" an action means performing the step at least partially in response to the action. It should be understood that the steps of the various embodiments of the present disclosure can be implemented in hardware, software, firmware or any combination thereof.
[0018] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0019] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to the relevant laws and regulations.
[0020] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user, so that the user can autonomously select whether to provide the personal information to the software or hardware, such as an electronic device, an application program, a server or a storage medium, performing the operation of the technical solution of the present disclosure according to the prompt information.
[0021] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information may be presented in a text manner. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0022] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0023] As used herein, the term "model" can learn an association between respective inputs and outputs from training data, such that a corresponding output can be generated for a given input after training is completed. The generation of a model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is one example of a model based on deep learning. In this document, a "model" can also be referred to as a "machine learning model", a "learning model", a "machine learning network", or a "learning network", which terms are used interchangeably herein.
[0024] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, which typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence, such that the output of a previous layer is provided as input to a subsequent layer, with the input layer receiving the input to the neural network and the output of the output layer as the final output of the neural network. Each layer of a neural network includes one or more nodes (also referred to as processing nodes or neurons), each of which processes input from the previous layer.
[0025] Generally, machine learning can include three stages, i.e., a training stage, a testing stage, and an application stage (also referred to as an inference stage). In the training stage, a given model can be trained using a large amount of training data, iteratively updating parameter values until the model is able to obtain consistent inferences from the training data that satisfy an expected target. Through training, the model can be considered to have learned an association (also referred to as a mapping) between input and output from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to determine whether the model is able to provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs based on the parameter values obtained through training to determine corresponding outputs.
[0026] Figure 1 A schematic diagram illustrating an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As shown, the example environment 100 can include a cloud environment 110 in which one or more computing clusters 120 are deployed. The computing clusters 120 can provide computing resources required for inference tasks of a machine learning model 160. That is, the cloud environment 110 can utilize the computing clusters 120 to perform inference tasks of the machine learning model 160. Figure 1
[0027] In some embodiments, the computing clusters 120 can include a plurality of instances (also referred to as computing instances). Each computing instance can include a physical instance or a virtual instance. Examples of the virtual instance can include, but are not limited to, a container, a virtual machine, and the like. The computing clusters 120 can provide a trusted execution environment (TEE). The computing clusters 120 can perform inference tasks of the machine learning model 160 in the trusted execution environment to improve security. The trusted execution environment is a hardware-based security technology that builds a secure computing environment isolated from the outside by dividing a secure part and a non-secure part. The secure computing environment can guarantee the confidentiality and integrity of data and code loaded inside the trusted execution environment. The trusted execution environment is isolated from the ordinary environment and has a higher security level, which is suitable for performing processing on sensitive data therein. Private cloud computing (PCC) can be run in the trusted execution environment. The private cloud computing is a secure computing framework on a cloud based on the TEE, which aims to build a set of secure computing services trusted by users to provide a secure and reliable cloud running environment for users and guarantee the security of the entire link of end-to-cloud collaboration.
[0028] In some embodiments, the terminal device 130 can be communicatively connected with the cloud environment 110. The terminal device 130 can send an inference request to the cloud environment 110 based on the communicative connection, to request the machine learning model 160 to perform an inference task. The cloud environment 110 can perform the inference task using the machine learning model 160 in response to the inference request, and obtain an inference result. The cloud environment 110 can feed back the inference result to the terminal device 130.
[0029] In some examples, the terminal device 130 can be an electronic device of the user 140. An application or a digital assistant, etc. can be deployed in the terminal device 130, which can utilize the machine learning model 160 to support the interaction with the user 140. As an example, the application or the digital assistant can utilize the machine learning model 160 to provide a question-answering service to the user 140. The terminal device 130 can present a user interface 150 of the application, e.g. a conversation page. The terminal device 130 can present conversation content in the user interface 150, e.g. a question from the user 140 or an answer to the question, etc. Of course, the terminal device 130 is not limited to providing a question-answering service to the user 140 using the machine learning model 160, but can also provide any appropriate service to the user 140 using the machine learning model 160, and embodiments of the present disclosure are not limited in this regard.
[0030] In some examples, the terminal device 130 can be an electronic device (e.g. a server) of a service provider of the application or the digital assistant. The provider of the machine learning model 160 can deploy the machine learning model 160 in the cloud environment 110, and the service provider of the application can utilize the terminal device 130 to request the inference service of the machine learning model 160 from the cloud environment 110, to support the function of the application.
[0031] The machine learning model 160 can be different types of models. In some embodiments, the machine learning model 160 can be constructed based on a language model (LM). The machine learning model used is a content generative model, which is capable of generating a corresponding output based on a model input. In some embodiments, the machine learning model based on the language model is capable of processing a model input in a text modality (e.g. natural language and / or machine language) and / or a model input in a non-text modality (e.g. image, voice, video, etc.), and is capable of generating a desired output according to the model input and a prompt word. The prompt word here is used to guide the machine learning model to generate an output that can solve the user demand indicated by the model input. In the application scenario for supporting user conversation, the input of the user 140 can be provided to the machine learning model 160 as at least part of the model input (other parts can include the prompt word). The user input is regarded as a question. Based on the model output, a corresponding answer can be generated to be provided to the user 140.
[0032] In some embodiments, the machine learning model 160 can be a large model. For example, the machine learning model 160 can be a large language model (LLM). As another example, the machine learning model 160 can be a large model capable of processing multi-modal inputs, such as a visual language model (VLM).
[0033] In some embodiments, the terminal device 130 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a smartbook, a media tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combinations of the aforementioned and the like, including wearables, accessories, peripherals, and the like of such devices and / or any other suitable device.
[0034] In some embodiments, the cloud environment 110 can include a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content distribution network, and big data and artificial intelligence platform. Embodiments of the present disclosure are not limited in this respect.
[0035] It should be understood that the structure and function of the environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure.
[0036] As mentioned above, with the integration of cloud computing technology and machine learning technology, more and more providers of machine learning models choose to deploy machine learning models in a cloud environment. The cloud environment can provide sufficient computing resources to perform the inference service of the machine learning model, which is conducive to improving the inference efficiency and response speed of the machine learning model.
[0037] However, the cloud environment faces certain data security problems. Specifically, the inference task of the machine learning model can require the coordination of multiple servers or multiple server clusters in the cloud environment to complete. By deploying a trusted execution environment for a single server or server cluster, the security of the processing process within the trusted execution environment can be ensured to a certain extent. However, the communication process between multiple servers or multiple server clusters can cause data to be exposed outside the trusted execution environment, resulting in the breaking of the trust boundary of the trusted execution environment.
[0038] For example, inference tasks of a machine learning model can be divided into pre- padding tasks and decoding tasks according to different types of requirements for computing resources. The pre-padding tasks are computationally intensive tasks, and a computationally intensive computing instance can be selected to execute the pre-padding tasks. The decoding tasks are memory intensive tasks, and a memory intensive computing instance can be selected to execute the decoding tasks. For example, a computationally intensive computing instance and a memory intensive computing instance can be selected to construct a pre-padding cluster and a decoding cluster, respectively. In this case, how to ensure that the communication process between the pre-padding cluster and the decoding cluster is trustworthy becomes a problem to be solved urgently.
[0039] Therefore, embodiments of the present disclosure propose an improved solution for model trustworthy inference. In the improved solution, a computing cluster for model inference can provide a plurality of trusted execution environments, the computing cluster can have a corresponding first key and a second key, and each of the plurality of trusted execution environments can hold at least one of the first key or the second key. A first stage processing of a machine learning model is performed on a model input in a first trusted execution environment of the plurality of trusted execution environments to generate an intermediate processing result corresponding to the model input. The intermediate processing result is encrypted in the first trusted execution environment using the first key to obtain encrypted data. The encrypted data is provided from the first trusted execution environment to a second trusted execution environment of the plurality of trusted execution environments. The encrypted data is decrypted in the second trusted execution environment using the second key to obtain the intermediate processing result. Then, a second stage processing of the machine learning model is performed on the intermediate processing result in the second trusted execution environment to generate a model output corresponding to the model input.
[0040] In embodiments of the present disclosure, the intermediate processing result is encrypted using a key before being passed. The encrypted data is passed between trusted execution environments. After obtaining the encrypted data, the trusted execution environment can decrypt the intermediate processing result from the encrypted data using the key, so that the inference process of the machine learning model is securely and efficiently performed. In this way, the communication process between different inference stages of the machine learning model can be ensured to be secure and reliable, so that the security boundary of the cloud environment can be expanded and the security of the cloud environment can be improved.
[0041] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings. Figure 2 A schematic diagram of an example architecture 200 for model trustworthy inference according to some embodiments of the present disclosure is shown. Embodiments of the present disclosure will be described below in conjunction with the environment 100 and the example architecture 200 in Figure 1 The improved solution of embodiments of the present disclosure will be described below in conjunction with the environment 100 and the example architecture 200 in FIG. 1 and FIG. 2, but this is only exemplary. It should be understood that the improved solution of embodiments of the present disclosure is not only limited to being applied to the environment 100 or the example architecture 200, but can also be applied to any other appropriate environment or architecture, and embodiments of the present disclosure do not limit this.
[0042] In some embodiments of the present disclosure, the computing cluster 120 is used for model inference, and can provide multiple trusted execution environments. The computing cluster 120 has a corresponding first key and a second key. The first key and the second key here can be referred to as “shared keys” or “common keys” of the multiple trusted execution environments in the computing cluster 120. Each trusted execution environment in the multiple trusted execution environments can hold at least one of the first key or the second key. The first key and the second key are used to perform encryption or decryption on data transmitted between the trusted execution environments to establish an encrypted channel between the trusted execution environments, thereby extending the security boundary of the cloud environment 110.
[0043] The first key and the second key can be the same or different. In some examples, the first key and the second key can be symmetric keys. In this case, the first key and the second key are the same. Alternatively, the first key and the second key can be asymmetric keys. In this case, the first key and the second key are different. Of course, the above-mentioned first key and second key are only exemplary, and any appropriate key can be selected according to actual needs. Embodiments of the present disclosure are not limited in this regard.
[0044] In some examples, if the first key and the second key are the same, each trusted execution environment can hold the key. If the first key and the second key are different, each trusted execution environment can hold both the first key and the second key, or each trusted execution environment can also hold one of them. For example, if the data transmission direction between the trusted execution environments is unidirectional, the trusted execution environment that outputs data can hold the first key, and the trusted execution environment that receives data can hold the second key. Of course, the above-mentioned key holding manner is only exemplary, and embodiments of the present disclosure are not limited in this regard.
[0045] In order to facilitate understanding of the improvement scheme of the embodiments of the present disclosure, the computing cluster 120 will be introduced first in the following. Then the creation and acquisition process of the key of the computing cluster 120 will be introduced. Finally, the process of encrypted communication between the trusted execution environments using the key will be introduced.
[0046] In some embodiments of the present disclosure, the inference process of the machine learning model 160 can include multiple stages, and the computing cluster 120 can provide multiple trusted execution environments corresponding to the multiple stages. In each trusted execution environment in the multiple trusted execution environments, the corresponding stage processing of the machine learning model 160 is performed. In some embodiments, as shown in FIG. 1, the machine learning model 160 can include a first stage 161, a second stage 162, and a third stage 163. The computing cluster 120 can provide a first trusted execution environment 121, a second trusted execution environment 122, and a third trusted execution environment 123. The first trusted execution environment 121 can execute the first stage 161 of the machine learning model 160, the second trusted execution environment 122 can execute the second stage 162 of the machine learning model 160, and the third trusted execution environment 123 can execute the third stage 163 of the machine learning model 160. Figure 2As shown, the computing cluster 120 can deploy multiple trusted execution environment (TEE) instances 212, 214, each of which can provide a corresponding trusted execution environment 216, 218. In some embodiments, the computing cluster 120 can include multiple subclusters corresponding to multiple stages of the machine learning model, and each subcluster can include multiple TEE instances 212, 214.
[0047] The TEE instances 212 and 214 in the computing cluster 120 can select any appropriate computing instance based on the different types of computing resources required at each stage. In some examples, examples of TEE instances 212 and 214 may include, but are not limited to, physical instances or virtual instances, etc., and virtual instances may include, but are not limited to, containers, virtual machines, etc. In some examples, TEE instances 212 and 214 may include compute-intensive computing instances or memory-intensive computing instances. Of course, the above TEE instances are only exemplary, and any appropriate instance can be selected to construct a TEE instance according to actual needs, and the embodiments of the present disclosure are not limited to this.
[0048] In some embodiments, the machine learning model 160 may include a first-stage process and a second-stage process. The plurality of trusted execution environments 216, 216 may include a plurality of trusted execution environments 216 configured to execute the first-stage process (sometimes also referred to herein as "first trusted execution environments") and a plurality of trusted execution environments 218 configured to execute the second-stage process (sometimes also referred to herein as "second trusted execution environments").
[0049] In some examples, such as Figure 2 As shown, the computing cluster 120 may include at least one first sub-cluster and at least one second sub-cluster. The first sub-cluster may include multiple TEE instances 212-1, 212-2, ..., 212-M, and the multiple TEE instances 212 respectively form multiple trusted execution environments 216-1, 216-2, ..., 216-M, where M is a positive integer. The second sub-cluster may include multiple TEE instances 214-1, 214-2, ..., 214-N, and the multiple TEE instances 214 respectively form multiple trusted execution environments 218-1, 218-2, ..., 218-N, where N is a positive integer.
[0050] The first-stage processing and the second-stage processing can be any two stages of a plurality of stages in an inference process of the machine learning model 160. The first-stage processing and the second-stage processing can be different in cases where the model architecture and the inference process of the machine learning model 160 are different. The plurality of computing instances of the computing cluster 120 can include at least one first type of instance configured to perform the first-stage processing and at least one second type of instance configured to perform the second-stage processing. Different types of instances have different consumption patterns of computing resources, and thus different types of instances are deployed or included in different TEE instances. For example, one TEE instance 212 can include one or more first type of instances, while one TEE instance 214 can include one or more second type of instances.
[0051] In some embodiments, the first-stage processing can include a prefill (also referred to as “full inference”) of the machine learning model 160, and the second-stage processing can include a decode (also referred to as “incremental inference”) of the machine learning model 160. The prefill is a compute-intensive task, and the plurality of TEE instances 212 can include, for example, a plurality of compute-intensive computing instances. The decode is a memory-intensive task, and the plurality of TEE instances 214 can include, for example, a plurality of memory-intensive computing instances. As such, efficient execution of the first-stage processing and the second-stage processing can be ensured.
[0052] In some examples, each TEE instance 212 in the first sub-cluster can be deployed with or include one or more prefill instances (PrefillInstances). The prefill instances can be utilized to perform a prefill task of the machine learning model 160. Each TEE instance 214 in the second sub-cluster can be deployed with or include one or more decode instances (DecodeInstances). The decode instances can be utilized to perform a decode task of the machine learning model 160. The prefill instances and the decode instances can include physical instances or virtual instances, such as containers, virtual machines, and the like.
[0053] In some embodiments, the computing cluster 120 can further include a network node 220. The plurality of TEE instances 212, 214 can be communicatively connected through the network node 220 to enable data transmission between the plurality of TEE instances 212, 214. The network node 220 can be configured to schedule the plurality of TEE instances 212, 214 to perform inference tasks of the machine learning model 160. For example, the network node 220 can be configured to schedule the plurality of TEE instances 212 to perform the first stage tasks and the plurality of TEE instances 214 to perform the second stage tasks based on a load balancing policy. For example, if the network node 220 receives an encrypted intermediate processing result from a certain TEE instance 212, the network node 220 can select a TEE instance 214 from the plurality of TEE instances 214 that currently has computing power resources. It should be appreciated that the computing power resources of a TEE instance depend on the computing power resources of the computing instance running therein. In some examples, the plurality of TEE instances 212 in the first sub-cluster can be communicatively connected with the plurality of TEE instances 214 in the second sub-cluster through the network node 220. The computing cluster 120 can utilize the network node 220 to schedule the plurality of TEE instances 212 and the plurality of TEE instances 214 to perform the pre-filling tasks and the decoding tasks of the machine learning model 160, respectively. In some examples, the network node 220 can be a gateway, e.g., a Prefill-Decode Gateway. The Prefill-Decode Gateway can schedule the TEE instances 212 and the TEE instances 214 to perform the pre-filling tasks and the decoding tasks based on a load balancing policy. Of course, the above network node 220 is only exemplary, and any appropriate network device(s) can be selected to form the network node 220 according to actual needs.
[0054] In some embodiments of the present disclosure, the cloud environment 110 can include a Trusted Key Service (TKS) 230. Alternatively, the TKS can be deployed outside of the cloud environment 110, e.g., in other cloud environments. The TKS is a secure key service running in a PCC, aiming to provide users with hardware-protected key management and proxy services. In some embodiments, the plurality of computing instances of the computing cluster 120 can respectively obtain at least one of the first key or the second key from the Trusted Key Service 230 based on identification information of the computing cluster. For example, the first type of instances obtain at least the first key, and the second type of instances obtain at least the second key.
[0055] The identification information can include any appropriate information capable of identifying the computing cluster 120, such as a cluster ID, a cluster name, a cluster address, and the like. The identification information can be determined in any one of the plurality of trusted execution environments 216, 218. The identification information is shared or common to the plurality of computing instances of the computing cluster 120.
[0056] In some embodiments, the identification information indicative of the compute cluster 120 is determined in one of the plurality of trusted execution environments 216, 218. Thereafter, a key generation request is sent to the trusted key service 230 in the trusted execution environment to request the trusted key service 230 to generate at least one of the first key or the second key based on the identification information. It can be appreciated that the operations performed in the trusted execution environment can be performed by the corresponding TEE instance in the trusted execution environment. For example, performing an operation in the trusted execution environment 216-1 can be understood as performing the operation by the TEE instance 212-1 in the trusted execution environment 216-1. For example, a compute instance of the plurality of compute instances can send a key generation request to the trusted key service 230 to request the trusted key service 230 to generate at least one of the first key or the second key based on the identification information. In some examples, the TEE instance 212-1 in the first sub-cluster can be the first instance of the compute cluster 120. The TEE instance 212 can determine a cluster number of the compute cluster 120, and the TEE instance 212 can send a key generation request to the trusted key service 230 based on the cluster number. The trusted key service 230 can generate the first key and / or the second key in response to the key generation request.
[0057] Example embodiments regarding the identification information are described below. In some embodiments, the identification information of the compute cluster is generated by one of the plurality of compute instances. The instance performs authentication of the remaining compute instances of the plurality of compute instances except for the compute instance by a remote attestation service. The identification information is provided by the compute instance to the remaining instances in response to the remaining instances passing the authentication.
[0058] Continuing with reference to Figure 2Example embodiments are described. In some embodiments, the cloud environment 110 can include a remote attestation service 240. Alternatively, the remote attestation service 240 can be deployed outside of the cloud environment 110, e.g., in other cloud environments. The remote attestation service 240 can be configured to perform verification on the security of the trusted execution environments 216, 218. In one of the plurality of trusted execution environments 216, 218, based on a plurality of pieces of verification information (also referred to as “second verification information” herein) respectively corresponding to the plurality of trusted execution environments 216, 218, an identity verification request for the plurality of trusted execution environments 216, 218 is sent to the remote attestation service 240 to request the remote attestation service 240 to verify the security of the corresponding trusted execution environment. The remote attestation service 240 can verify the security of the plurality of trusted execution environments 216, 218 based on the plurality of pieces of verification information to obtain a verification result in response to the identity verification request. The remote attestation service 240 can feed back the verification result to the corresponding trusted execution environment. If the verification result indicates that at least one remaining trusted execution environment of the plurality of trusted execution environments 216, 218 passes the verification, the trusted execution environment can provide identification information to the at least one remaining trusted execution environment. In this way, the TEE instance that passes the verification can join the computing cluster 120 using the identification information, which can ensure that the TEE instance that joins the computing cluster 120 is secure and trusted.
[0059] The verification information herein can include information related to the TEE instance 212, 214 and information related to the trusted execution environment 216, 218. The information related to the TEE instance can include, for example, the model number, version number of the TEE instance 212, 214, and identity attestation indicating whether the TEE instance is deployed with a TEE, etc. The information related to the trusted execution environment 216, 218 can include a measurement value indicating whether the TEE is tampered, or a configuration parameter indicating the security mode and permission settings of the TEE, etc. Of course, the above verification information is only exemplary, and any appropriate verification information can be selected to verify the security of the trusted execution environment 216, 218 according to actual needs. Embodiments of the present disclosure are not limited in this regard.
[0060] In some examples, as Figure 2As shown, after determining the cluster number, the TEE instance 212-1 can send the verification information corresponding to the plurality of TEE instances 212 and the plurality of TEE instances 214 to the remote attestation service 240, to request the remote attestation service 240 to verify the security of the trusted execution environments 216, 218 corresponding to the plurality of TEE instances 212 and the plurality of TEE instances 214. If the verification result indicates that the plurality of TEE instances 212 and the plurality of TEE instances 214 are both verified, the TEE instance 212-1 can provide the cluster number to the TEE instances 212-2, …, 212-M, and the TEE instances 214-1, 214-2, …, 214-N.
[0061] In some embodiments, in one of the plurality of trusted execution environments 216, 218, based on verification information (also referred to as “first verification information” herein sometimes) indicating the security of the trusted execution environment and the identification information of the computing cluster 120, a key acquisition request is sent to the trusted key service 230. The trusted key service 230 can respond to the key acquisition request, based on the verification information, to perform verification on the security of the trusted execution environment requesting the key. If the verification passes, meaning that the trusted execution environment is trusted, the trusted key service 230 can feed back at least one of the first key or the second key to the trusted execution environment. If the verification fails, meaning that the trusted execution environment is not trusted, the trusted key service 230 can prohibit feeding back the key to the trusted execution environment. In this way, it can be ensured that the key is provided to the trusted TEE instance, and the untrusted TEE instance is prevented from obtaining the key.
[0062] The first verification information can be the same as or different from the second verification information described above. For example, the first verification information can include all or part of the content in the second verification information. In some examples, the trusted key service 230 can assist the remote attestation service 240 to verify the security of the trusted execution environment requesting the key. For example, the trusted key service 230 can forward the verification information in the key acquisition request to the remote attestation service to request the remote attestation service 240 to assist in verifying the security of the trusted execution environment. Of course, if the trusted key service 230 has the ability to verify the security of the trusted execution environment, the trusted key service 230 can also perform the verification on the security of the trusted execution environment. Embodiments of the present disclosure do not limit this.
[0063] In some examples, if the first key and the second key are the same, trusted key service 230 may provide the first key or the second key back to the trusted execution environment. If the first key and the second key are different, trusted key service 230 may provide both the first key and the second key back to the trusted execution environment. Of course, if the data transfer relationship between the trusted execution environments is one-way, trusted key service 230 may provide the first key for encryption to the trusted execution environment that outputs data, and may provide the second key for decryption to the trusted execution environment that receives data.
[0064] In some embodiments, the key acquisition process described above can be performed by a computing instance. A computing instance among the multiple computing instances sends a key acquisition request to the trusted key service 230 based on verification information indicating the trusted execution environment in which the computing instance is located and identification information of the computing cluster 120. The computing instance receives at least one of the first key or the second key fed back by the trusted key service 230 in response to the key acquisition request. For example, a first-type instance can request a first key from the trusted key service 230, and a second-type instance can request a second key from the trusted key service 230.
[0065] As an example, Figure 2 As shown, the trusted key service 230 generates a first key and / or a second key in response to a key generation request sent by the TEE instance 212-1. The trusted key service 230 may feedback notification information to the TEE instance 212-1 to notify the TEE instance 212-1 that key generation is complete. The TEE instance 212-1 may send a key acquisition request to the trusted key service 230 based on the verification information corresponding to the cluster number. If the security verification of the TEE instance 212-1 passes, the trusted key service 230 may send the first key and / or the second key to the TEE instance 212-1.
[0066] As another example, Figure 2As shown, if TEE instance 214-1 receives the cluster number of the computing cluster 120 from TEE instance 212-1, TEE instance 214-1 can send a key acquisition request to the trusted key service 230 based on the cluster number and the corresponding verification information. The trusted key service 230 can verify the security of the TEE instance 214-1, and if the verification is passed, meaning that the TEE instance 214-1 is trusted, the trusted key service 230 can feed back the first key and / or the second key corresponding to the computing cluster 120 to the TEE instance 214-1. If the verification fails, meaning that the TEE instance 214-1 is not trusted, the trusted key service 230 can prohibit feeding back the key to the TEE instance 214-1, for example, the trusted key service 230 can feed back a notification to the TEE instance 214-1 that the verification fails. It can be understood that each TEE instance in the computing cluster 120 can acquire the key from the trusted key service 230 in a similar manner.
[0067] In some embodiments of the present disclosure, the first stage processing of the machine learning model 160 is performed on the model input in the trusted execution environment 216 to generate an intermediate processing result corresponding to the model input. For example, the first type instance can perform the first stage processing of the machine learning model 160 on the model input to generate the intermediate processing result corresponding to the model input. It can be known from the foregoing analysis that the first stage processing can be different in the case of different model architectures or inference processes of the machine learning model 160. The first stage processing can be any appropriate stage in the inference process of the machine learning model 160. It can be understood that the trusted execution environment 216 can be any one of the plurality of trusted execution environments 216.
[0068] In some examples, the first stage processing can include pre-padding of the machine learning model 160. The pre-padding of the machine learning model 160 is performed on the model input in the first trusted execution environment to generate at least one of the key cache data, the value cache data, or the first output token of the attention mechanism as the intermediate processing result. As an example, in combination with the foregoing analysis, the first stage processing can include the pre-padding of the machine learning model 160. Figure 1 and Figure 2 As shown, if the cloud environment 110 receives the inference request from the terminal device 130, the cloud environment 110 can determine the model input based on the inference request (for example, which can include the question of the user 140). The TEE instance 212-1 can perform the pre-padding task of the machine learning model 160 based on the model input to generate the key cache data, the value cache data, and the first output token.
[0069] In some embodiments of the present disclosure, the intermediate processing result is encrypted by the first key pair in the trusted execution environment 216 to obtain encrypted data. In some embodiments, the first type instance can encrypt the intermediate processing result by the first key pair to obtain encrypted data. As an example, the TEE instance 212-1 can encrypt the key cache data, the value cache data, and the first output token by the first key pair to obtain an encrypted data package (i.e., encrypted data). It should be understood that the encryption process can be different in the case of different types of the first key. Embodiments of the present disclosure do not limit the type of the first key and the encryption process.
[0070] In some embodiments of the present disclosure, the TEE instance 212 provides the encrypted data from the trusted execution environment 216 to the trusted execution environment 218. In some embodiments, the TEE instance 212 can provide the encrypted data from the trusted execution environment 216 to the network node 220. For example, the first type instance included in the TEE instance 212 can send the encrypted data to the network node 220. The network node 220 can select at least one trusted execution environment 218 from a plurality of trusted execution environments 218. Then, the network node 220 can provide the encrypted data to the at least one trusted execution environment 218. As an example, the TEE instance 212-M can provide the encrypted data to the network node 220, which can select a TEE instance with a relatively small load, such as the TEE instance 214-2, from a plurality of TEE instances 214 based on a load balancing strategy. Then, the network node 220 can provide the encrypted data to the TEE instance 214-2. Of course, the network node 220 can also use other any appropriate scheduling strategy to perform the scheduling task of the encrypted data. For example, the network node 220 can also randomly select a feasible execution environment 218 from a plurality of trusted execution environments 218. Embodiments of the present disclosure do not limit this.
[0071] In some embodiments of the present disclosure, the encrypted data is decrypted by the second key pair in the trusted execution environment 218 to obtain the intermediate processing result. For example, the second type instance can decrypt the encrypted data by the second key pair to obtain the intermediate processing result. As an example, if the TEE instance 214-2 receives the encrypted data package, the TEE instance 214-2 can decrypt the encrypted data package by the second key pair in the trusted execution environment 218-2 to obtain the key cache data, the value cache data, and the first output token.
[0072] In some embodiments of the present disclosure, a second stage processing of the machine learning model 160 is performed on the intermediate processing result in the trusted execution environment 218 to generate a model output corresponding to the model input. For example, the second type instance can perform the second stage processing of the machine learning model 160 on the intermediate processing result to generate the model output corresponding to the model input. The second stage processing can be any appropriate stage in an inference process of the machine learning model 160. In some embodiments, the second stage processing can include a decoding of the machine learning model 160. The TEE instance 214 can perform the decoding in the respective trusted execution environment 218 based on the key cache data, the value cache data, and the first output token to generate at least one remaining output token after the first output token. As an example, the TEE instance 218-2 can determine a response to the question of the user 140 based on the at least one remaining output token. Thereafter, the cloud environment 110 can feed back the response to the terminal device 130.
[0073] It should be appreciated that the actions described above with reference to the TEE instances can be performed by the computing instances of the computing cluster. For example, the actions described with reference to the TEE instances 212 can be performed by the first type instances, and the actions described with reference to the TEE instances 214 can be performed by the second type instances.
[0074] In this way, the communication process between different inference stages of the machine learning model can be ensured to be secure and reliable. As a result, the security boundary of the cloud environment can be expanded, and the security of the cloud environment can be improved.
[0075] Figure 3 A flowchart of a process 300 for model trusted inference is shown in accordance with some embodiments of the present disclosure. The process 300 can be applied to the computing cluster 120 for model inference. The process 300 will be described below with reference to the environment 100 in Figure 1 It should be appreciated that the actions described with reference to the computing cluster 120 can be implemented by the instances (e.g., computing instances) or nodes in the computing cluster 120.
[0076] At block 310, the computing cluster 120 performs a first stage processing of a machine learning model on a model input in a first trusted execution environment to generate an intermediate processing result corresponding to the model input, the first trusted execution environment being one of a plurality of trusted execution environments of the computing cluster, each trusted execution environment of the plurality of trusted execution environments holding at least one of: a first key of the computing cluster or a second key corresponding to the first key.
[0077] At block 320, the computing cluster 120 encrypts the intermediate processing result with the first key in the first trusted execution environment to obtain encrypted data.
[0078] At block 330, the computing cluster 120 provides the encrypted data from the first trusted execution environment to a second trusted execution environment of the plurality of trusted execution environments.
[0079] At block 340, the computing cluster 120 decrypts, in the second trusted execution environment, the encrypted data with the second key to obtain an intermediate processing result.
[0080] At block 350, the computing cluster 120 performs, in the second trusted execution environment, a second stage processing of the machine learning model on the intermediate processing result to generate a model output corresponding to the model input.
[0081] In some embodiments, providing the encrypted data from the first trusted execution environment to the second trusted execution environment comprises: providing the encrypted data from the first trusted execution environment to a network node; and providing, by the network node, the encrypted data to the second trusted execution environment of the plurality of second trusted execution environments, the plurality of second trusted execution environments being configured for the second stage processing.
[0082] In some embodiments, performing the first stage processing on the model input comprises: generating, as the intermediate processing result, at least one of key cache data, value cache data, or a first output token of an attention mechanism by performing, in the first trusted execution environment, a prefill of the machine learning model on the model input.
[0083] In some embodiments, performing the second stage processing on the intermediate processing result comprises: performing, in the second trusted execution environment, a decoding based on the key cache data, the value cache data, and the first output token to generate at least one remaining output token after the first output token.
[0084] In some embodiments, the first stage processing on the model input is performed by a first type instance of the computing cluster for model inference, and the intermediate processing result is encrypted by the first type instance with the first key, and the encrypted data is decrypted by a second type instance of the computing cluster for model inference with the second key, and the second stage processing on the intermediate processing result is performed by the second type instance.
[0085] In some embodiments, the computing cluster for model inference comprises a plurality of instances including at least one first type instance configured to perform the first stage processing and at least one second type instance configured to perform the second stage processing, and the process 300 further comprises: obtaining, by the plurality of instances, at least one of the first key or the second key from a trusted key service based on identification information of the computing cluster, respectively.
[0086] In some embodiments, multiple instances obtain identification information in the following manner: one instance among the multiple instances generates identification information of the computing cluster; the instance performs authentication on the remaining instances among the multiple instances except the instance through a remote attestation service; and in response to the remaining instances passing the authentication, the instance provides the identification information to the remaining instances.
[0087] In some embodiments, obtaining at least one of the first key or the second key from a trusted key service includes: sending a key acquisition request to the trusted key service by an instance among multiple instances based on verification information indicating the trusted execution environment in which the instance is located and identification information of the computing cluster; and receiving, by the instance, at least one of the first key or the second key fed back by the trusted key service in response to the key acquisition request.
[0088] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 1 shows a schematic structural block diagram of an example apparatus 400 for model trustworthy reasoning according to certain embodiments of the present disclosure. Apparatus 400 may be implemented as or included in cloud environment 110. Each module / component in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0089] like Figure 4 As shown, the device 400 includes: a first processing module 410, configured to perform a first stage processing of a machine learning model on a model input in a first trusted execution environment in a computing cluster to generate an intermediate processing result corresponding to the model input, the first trusted execution environment being one of a plurality of trusted execution environments of the computing cluster, and each of the plurality of trusted execution environments holding at least one of the following: a first key of the computing cluster or a second key corresponding to the first key; an encryption module 420, configured to encrypt the intermediate processing result using the first key in the first trusted execution environment to obtain encrypted data; a providing module 430, configured to provide the encrypted data from the first trusted execution environment to a second trusted execution environment among the plurality of trusted execution environments; a decryption module 440, configured to decrypt the encrypted data using the second key in the second trusted execution environment to obtain the intermediate processing result; and a second processing module 450, configured to perform a second stage processing of the machine learning model on the intermediate processing result in the second trusted execution environment to generate a model output corresponding to the model input.
[0090] In some embodiments, the providing module 430 is further configured to: provide the encrypted data from the first trusted execution environment to the network node; and provide the encrypted data to a second trusted execution environment among a plurality of second trusted execution environments through the network node, the plurality of second trusted execution environments being configured for second stage processing.
[0091] In some embodiments, the first processing module 410 is further configured to generate, as the intermediate processing result, at least one of the key cache data, the value cache data, or the first output token of the attention mechanism by performing, in the first trusted execution environment, the pre-padding of the machine learning model on the model input.
[0092] In some embodiments, the second processing module 450 is further configured to perform, in the second trusted execution environment, the decoding based on the key cache data, the value cache data, and the first output token to generate at least one remaining output token after the first output token.
[0093] In some embodiments, the first processing module 410 and the encryption module 420 are included in a first type of instance in a computing cluster for model inference, and the decryption module 440 and the second processing module 450 are included in a second type of instance in the computing cluster for model inference.
[0094] In some embodiments, the computing cluster for model inference includes a plurality of instances including at least one first type of instance configured to perform the first stage processing and at least one second type of instance configured to perform the second stage processing, and the apparatus 400 is further configured to obtain, by the plurality of instances, at least one of the first key or the second key from the trusted key service based on the identification information of the computing cluster respectively.
[0095] In some embodiments, the plurality of instances obtain the identification information by: generating, by one of the plurality of instances, the identification information of the computing cluster; performing, by the instance, authentication of remaining instances of the plurality of instances other than the instance through a remote attestation service; and in response to the remaining instances passing the authentication, providing, by the instance, the identification information to the remaining instances.
[0096] In some embodiments, obtaining, from the trusted key service, at least one of the first key or the second key includes: sending, by an instance of the plurality of instances, a key obtaining request to the trusted key service based on verification information indicating a trusted execution environment in which the instance is located and the identification information of the computing cluster; and receiving, by the instance, at least one of the first key or the second key that is fed back by the trusted key service in response to the key obtaining request.
[0097] The units and / or modules included in the apparatus 400 can be implemented utilizing various means, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine-executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 400 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include Field- programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-Chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0098] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the disclosure can be implemented is shown. It should be understood that Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the scope of the embodiments described herein. Figure 5 The electronic device 500 shown can include or be implemented as Figure 1 the computing cluster 120 of FIG. 1, or Figure 4 the apparatus 400 of FIG. 4.
[0099] As shown in Figure 5 The electronic device 500 is in the form of a general-purpose electronic device. Components of the electronic device 500 can include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processor 510 can be a real or virtual processor and is capable of executing various processing in accordance with computer-executable instructions stored in the memory 520. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve parallel processing capabilities of the electronic device 500.
[0100] The electronic device 500 typically includes a number of computer storage media. Such media can be any available media that is accessible by the electronic device 500 and includes both volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable media and can include machine-readable media, such as a flash drive, a magnetic disk, or any other media that can be used to store information and / or data and that can be accessed by the electronic device 500.
[0101] The electronic device 500 can further include additional detachable / non-detachable, volatile / non-volatile storage media. Although not shown in FIG. 5, a disk drive for reading from or writing to a detachable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a detachable, non-volatile optical disk (e.g., a CD-ROM) can be provided. In these cases, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more computer-executable instruction modules configured to perform various methods or actions of various embodiments of the present disclosure. Figure 5
[0102] The communication unit 540 enables communication with other electronic devices through communication media. Additionally, the functionality of the components of the electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating with one another through a communication connection. As such, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0103] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc., one or more devices that enable a user to interact with the electronic device 500, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 500 to communicate with one or more other electronic devices, through the communication unit 540, as needed. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0104] According to an example implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, there is also provided a computer program product tangibly stored on a non-transitory computer-readable storage medium and comprising computer-executable instructions, where the computer-executable instructions are executed by a processor to implement the method described above.
[0105] The computer executable instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer executable instructions can also be stored in a computer readable storage medium that can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the computer executable instructions such that the instruction retrieval mechanisms of a general purpose computer, special purpose computer, or other programmable data processing apparatus can access the computer executable instructions to implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0106] The computer executable instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0107] The computer executable instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0108] The computer executable instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0109] Having described various implementations of the disclosure above, the descriptions are not exhaustive and do not limit the disclosure to the disclosed implementations. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The scope of the disclosure includes all the implementations of which an equivalent would be apparent to those skilled in the art from the disclosure given and the associated drawings. The selection of the terms to be used in the written description is not intended to limit the scope of the present disclosure, but rather to best describe the principles of various implementations in preference to a detailed inventory of all technical terms.
Claims
1. A method for model-trusted reasoning, comprising: Performing a first-stage processing of a machine learning model on a model input in a first trusted execution environment of a computing cluster used for model inference to generate an intermediate processing result corresponding to the model input, wherein the first trusted execution environment is one of multiple trusted execution environments of the computing cluster, and each of the multiple trusted execution environments holds at least one of the following: a first key of the computing cluster or a second key corresponding to the first key; Encrypting the intermediate processing result using the first key in the first trusted execution environment to obtain encrypted data; providing the encrypted data from the first trusted execution environment to a second trusted execution environment among the plurality of trusted execution environments; decrypting the encrypted data using the second key in the second trusted execution environment to obtain the intermediate processing result; as well as Perform second-stage processing of the machine learning model on the intermediate processing result in the second trusted execution environment to generate a model output corresponding to the model input.
2. The method of claim 1 , wherein providing the encrypted data from the first trusted execution environment to the second trusted execution environment comprises: providing the encrypted data from the first trusted execution environment to a network node; as well as The encrypted data is provided, via the network node, to a second trusted execution environment of a plurality of second trusted execution environments, the plurality of second trusted execution environments being configured for the second stage processing.
3. The method of claim 1 , wherein performing the first stage of processing on the model input comprises: By performing pre-filling of the machine learning model on the model input in the first trusted execution environment, at least a portion of the key cache data, value cache data or the first output word of the attention mechanism is generated as the intermediate processing result.
4. The method according to claim 3, wherein performing the second stage processing on the intermediate processing result comprises: In the second trusted execution environment, decoding is performed based on the key cache data, the value cache data, and the first output word to generate at least one remaining output word following the first output word.
5. The method according to claim 1, wherein the first stage of processing is performed on the model input by a first type instance in the computing cluster for model inference, and the first type instance encrypts the intermediate processing result using the first key, and The encrypted data is decrypted by a second type instance in the computing cluster used for model inference using the second key, and the second type instance performs the second stage processing on the intermediate processing result.
6. The method according to claim 1, wherein the computing cluster for model inference includes multiple instances, the multiple instances including at least one first type instance configured to perform the first stage processing and at least one second type instance configured to perform the second stage processing, and the method further comprises: The multiple instances respectively obtain at least one of the first key or the second key from a trusted key service based on the identification information of the computing cluster.
7. The method according to claim 6, wherein the identification information of the plurality of instances is obtained by: generating, by one of the plurality of instances, the identification information of the computing cluster; The instance performs authentication on the remaining instances of the plurality of instances except the instance through a remote attestation service; and In response to the other instances passing the authentication, the instance provides the identification information to the other instances.
8. The method of claim 6, wherein obtaining at least one of the first key or the second key from the trusted key service comprises: An instance among the multiple instances sends a key acquisition request to the trusted key service based on verification information indicating the trusted execution environment in which the instance is located and identification information of the computing cluster; as well as The instance receives at least one of the first key or the second key fed back by the trusted key service in response to the key acquisition request.
9. A device for model-trusted reasoning, comprising: a first processing module configured to perform a first stage of processing of a machine learning model on a model input in a first trusted execution environment of a computing cluster used for model inference to generate an intermediate processing result corresponding to the model input, wherein the first trusted execution environment is one of multiple trusted execution environments of the computing cluster, and each of the multiple trusted execution environments holds at least one of the following: a first key of the computing cluster or a second key corresponding to the first key; An encryption module is configured to encrypt the intermediate processing result using the first key in the first trusted execution environment to obtain encrypted data; A providing module configured to provide the encrypted data from the first trusted execution environment to a second trusted execution environment among the plurality of trusted execution environments; a decryption module, configured to decrypt the encrypted data using the second key in the second trusted execution environment to obtain the intermediate processing result; as well as The second processing module is configured to perform a second stage of processing of the machine learning model on the intermediate processing result in the second trusted execution environment to generate a model output corresponding to the model input.
10. An electronic device comprising: at least one processor; as well as At least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 8 when executed by the at least one processor.
11. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 8.
12. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Trusted computing program calling method and device, electronic equipment and storage medium
CN112948810A
Verifiability for execution in trusted execution environment
CN114270778A
Model reasoning method and device
CN116232562A
Model training method, model using method, system, trusted node and device
US20220350898A1
Cross-subnet calling
WO2024001022A1
Cited By
Large model safety protection system, large model safety protection method and medium
CN121711199A