Method, apparatus and program product for reasoning
By dividing the neural network model into two parts deployed in and outside the secure space, and using encrypted inference data, the problems of data security and privacy protection in machine learning tasks are solved, and secure deployment under the balance of resources and performance is achieved.
Patent Information
- Application Number
- CN202311873051.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
In machine learning tasks, prior art is difficult to protect data security and privacy while ensuring resource balance and model performance, especially when data or parameters are exposed in the inference phase, preventing further leakage.
By determining whether the weight matrix of the layer in the neural network model is an irreversible target matrix, the model is divided into two parts deployed in the secure space and outside, and the results are obtained using encrypted inference data to ensure that some models operate in a trusted and secure area.
It realizes that data security and privacy can still be protected when intermediate data or parameters are exposed, providing a safe and reliable deployment environment, avoiding further leakage, while maintaining resource and performance balance.
Smart Images

Figure CN120235239A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to the field of computers, and more particularly to methods, devices, and program products for inference. Background Art
[0002] With the rapid development of artificial intelligence and / or machine learning (AI / ML) technologies, their applications in numerous fields are becoming increasingly widespread. However, despite the significant improvements brought about by the popularization of these technologies in many fields (such as preference recommendation, smart cities, assisted / autonomous driving, etc.), data security and privacy protection issues have become increasingly prominent.
[0003] In all aspects of machine learning tasks, it is particularly important to ensure the security and reliability of operations and data flows. Appropriate and effective security measures and technical means are expected to be taken to ensure the secure execution of machine learning operations, so that data can be used reasonably and legally while remaining complete and confidential. Summary of the Invention
[0004] Embodiments of the present disclosure provide a solution for inference, through which secure deployment can be provided for the execution of a model, thereby ensuring the security and reliability of operations and data flows.
[0005] In a first aspect of the present disclosure, a method for inference is provided. The method includes receiving an input / output (I / O) request based on a protocol related to a first type of storage medium. The method includes determining, for a layer in a neural network model, whether the weight matrix of the layer is an irreversible target matrix. The method further includes, in response to determining that the weight matrix of the layer is the target matrix, dividing the neural network model into a first part and a second part, where the first part includes the layer and the previous layers before the layer and is deployed within a secure space, and the second part includes the subsequent layers after the layer and is deployed outside the secure space. The method further includes using the divided neural network model to obtain an inference result based on encrypted inference data.
[0006] In another aspect of the present disclosure, an electronic device for inference is provided. The electronic device includes a processor and a memory. The memory is coupled to the processor and stores instructions thereon. When executed by the processor, the instructions cause the device to perform actions. The actions include determining, for a layer in a neural network model, whether the weight matrix of the layer is an irreversible target matrix. The actions further include, in response to determining that the weight matrix of the layer is the target matrix, dividing the neural network model into a first part and a second part, where the first part includes the layer and the previous layers before the layer and is deployed within a secure space, and the second part includes the subsequent layers after the layer and is deployed outside the secure space. The actions further include using the divided neural network model to obtain an inference result based on encrypted inference data.
[0007] In another aspect of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer-readable storage medium and includes computer-executable instructions that, when executed, cause a computer to perform the method or process according to an embodiment of the present disclosure.
[0008] According to the inference solution of the embodiment of the present disclosure, reasonably partially deploying the model in a trusted security area can provide a secure and reliable deployment environment for machine learning tasks while ensuring resource balance and model performance. Even in the case of exposure of some intermediate parameters or data, it is impossible to deduce further leakage forward, thereby protecting data security and privacy.
[0009] Note that the Summary of the Invention section is provided to introduce a series of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify the key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] By describing the embodiments of the present disclosure in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present disclosure will become more clearly understood. In the drawings:
[0011] Figure 1 A schematic diagram of an example environment in which the method and / or process according to an embodiment of the present disclosure can be implemented is illustrated;
[0012] Figure 2 A flowchart of a method for inference according to an embodiment of the present disclosure is illustrated;
[0013] Figure 3 A schematic diagram of a model input-output stream according to an embodiment of the present disclosure is illustrated;
[0014] Figure 4 A flowchart of a security space verification process according to an embodiment of the present disclosure is illustrated;
[0015] Figure 5 A schematic diagram of an operator mapping process according to an embodiment of the present disclosure is illustrated;
[0016] Figure 6 A schematic diagram of an inference application for route planning according to an embodiment of the present disclosure is illustrated; and
[0017] Figure 7 A schematic block diagram of an example device that can be used to implement an embodiment of the present disclosure is illustrated.
[0018] In all the figures, the same or similar reference numerals generally denote the same or similar elements. Specific embodiments
[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0020] In the description of the embodiments of the present disclosure, the term "comprising" and its variants should be understood as open-ended inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects, unless clearly indicated otherwise.
[0021] As described above, with the wide application of AI / ML technologies, data security and privacy protection issues have become increasingly prominent. In all aspects of machine learning tasks, it is crucial to ensure the security and reliability of operations and data flows. To ensure the secure execution of machine learning operations and to ensure that data is used reasonably and legally while being complete and confidential, appropriate and effective security measures and technical means are required. In practical applications, it is necessary to always pay attention to the integrity and confidentiality of data and to ensure that machine learning operations are not affected by security degradations such as illegal attacks.
[0022] By way of example and not limitation, in the development process of an assisted / autonomous driving system, geographical data may be utilized, for example, for training a model or for inference through a trained model. However, such geographical data may not be directly accessible to the system developer or service provider or may be invisible to the developer, as such data is sensitive to the administrator or user. In addition, during the use of such data, it is also necessary to prevent adverse attacks from learning the original input through the final result or intermediate results (such as data or parameters, etc.). Therefore, the importance of protecting data security in such or similar scenarios is self-evident.
[0023] For example, a large number of Internet of Things (IoT) devices are now deployed for corresponding functionality. For machine learning applications, these IoT devices are typically used for inference so that computations can be performed at or near the data generation site. At the same time, such devices are often placed in physically insecure locations where hackers may steal the generated data.
[0024] In this case, data encryption (e.g., confidentiality, authentication, etc.) would be the solution to protect data from direct access. Ideally, we also hope that there are differences between the training phase and the inference phase. For example, the weight matrix and bias terms during the inference phase are visible or accessible. Once some intermediate data or parameters are exposed through hacking, the consequences would be unthinkable. All of these would make it more difficult to protect data security during the inference phase.
[0025] Traditional methods may deploy the entire model in a secure space to obtain a completely secure environment. However, in reality, the size of the secure space that a processor can usually provide is limited. Even if the size of the secure space of some chips is relatively large, their usage cost is very high. In addition, some secure spaces only support CPU computing, which may reduce performance compared to GPU implementation.
[0026] To address at least some of the above and other potential problems, embodiments of the present disclosure propose a solution for inference. The solution includes determining, for a layer in a neural network model, whether the weight matrix of the layer is an irreversible target matrix. The solution further includes, in response to determining that the weight matrix of the layer is the target matrix, dividing the neural network model into a first part and a second part, where the first part includes the layer and the previous layers before the layer and is deployed in a secure space, and the second part includes the subsequent layers after the layer and is deployed outside the secure space. The solution also includes using the divided neural network model to obtain an inference result based on encrypted inference data. In this way, by partially deploying the model in a trusted secure area, it is possible to provide a secure deployment environment for machine learning tasks while ensuring resource balance and model performance. Even if some intermediate data or parameters are exposed, it is impossible to deduce further leakage forward, thereby protecting data security and privacy.
[0027] The following refers to Figures 1 to 7 to illustrate the basic principles and several example implementations of the present disclosure. It should be understood that these exemplary embodiments are only provided to enable those skilled in the art to better understand and then implement the embodiments of the present disclosure, rather than limiting the scope of the present disclosure in any way.
[0028] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which the method and / or process according to an embodiment of the present disclosure can be implemented. The example environment 100 exemplarily shows a hierarchical storage system including multiple nodes. As Figure 1As shown, the example environment 100 may include an input 110, a computing device 120, and a storage device 130. It should be understood that only a limited number of components or units are shown here for the purpose of facilitating understanding and easy illustration, but the embodiments of the present disclosure are not limited thereto and may further include other devices or systems. For example, the example environment 100 may further include a display (not shown), which is configured to display the inference results of the computing device 120.
[0029] According to an embodiment of the present disclosure, the input 110 includes encrypted inference data (hereinafter, also simply referred to as encrypted inference data). The inference data can be encrypted using a key at the data source, where the data source can be the origin where the inference data is generated. In some embodiments, in response to its own generation, the inference data can be encrypted at its generation location. The type of inference data may include, but is not limited to, text, image, audio, video, etc. It should be understood that the present disclosure does not impose limitations and constraints on the size and type of the inference data, etc., as long as it corresponds to the model it is fed into.
[0030] As described above, the inference data is encrypted at the data source. According to an embodiment of the present disclosure, the encrypted inference data can be decrypted only in a trusted security area, for example, decrypted by performing confidential computing in this area based on the key at the above data source. In this way, the integrity and confidentiality of the inference data from its generation location to the security area can be ensured, preventing security degradation such as data leakage from occurring during this process. Hereinafter, the trusted security area according to the embodiment of the present disclosure will be further described in detail.
[0031] According to an embodiment of the present disclosure, a neural network model for predetermined functionality (such as classification and prediction, etc.), for example, a deep learning model, can be run on the computing device 120. The computing device 120 may have computing capabilities corresponding to running this model and may be arranged locally or distributedly in the cloud, or a combination thereof. The computing device 120 can load the neural network model from the storage device 130, and during the running of the neural network model, it can access the storage device 130 and use the data stored in the storage device 130 to perform corresponding calculations. It should be understood that the computing device 120 is Figure 1 schematically shown as one computing device here, but this is only for the purpose of facilitating illustration and easy understanding. In the example environment 100, more computing devices can be arranged according to actual needs.
[0032] The computing device 120 can be configured with a secure space, i.e., the above-mentioned trusted secure area. The secure space according to an embodiment of the present disclosure can be an independent and isolated secure execution environment implemented, for example, on-chip (such as through a secure processor). Such an environment provides a confidential space for sensitive data and computations to enable secure operations, integrity protection, etc., protecting it from external threats, thereby ensuring that the data and computations are protected by security, integrity, and privacy throughout the processing within it. Examples of the secure space according to an embodiment of the present disclosure can include, but are not limited to, a trusted execution environment (TEE), etc. In some embodiments, the secure space can be configured to execute secure computations including decryption, authentication, etc., and the secure space can be configured to have a size corresponding to the computing resources (the computing power of the chip or processor). Below, the corresponding operations on the computing device 120 will be further described in detail.
[0033] By way of example and not limitation, the computing device 120 can include, but is not limited to, a personal computer, a laptop computer, a server computer, a mobile device (such as a smartphone, a tablet computer, etc.), a wearable electronic device, a multimedia player, a personal digital assistant (PDA), a smart home device, a consumer electronic product, or a distributed computing environment including any one or more of the above devices, etc. In some embodiments, a part of the computing device 120 can be arranged locally, while another part is arranged in the cloud.
[0034] According to an embodiment of the present disclosure, the storage device 130 can be configured to store the model to be run on the computing device 120 and its various parameters, etc. In addition, the storage device 130 can be configured to store the input and output of the model (i.e., inference data and inference results), etc. The situation in training is similar, so it will not be elaborated here. It should be understood that the storage device 130 is Figure 1 schematically shown as a single storage device for the purpose of facilitating illustration and easy understanding, but in the example environment 100, more storage devices can also be arranged according to actual needs.
[0035] By way of example and not limitation, the storage device 130 can include, but is not limited to, local storage devices, remote storage devices, and their combinations. In some embodiments, multiple storage devices in the storage device 130 can include, but are not limited to, a mechanical hard disk drive (HDD), a solid-state drive (SSD), etc., and some of the multiple storage devices can be arranged locally, while others can be arranged remotely and are coupled together, for example, via a line or a network, etc.
[0036] The above combines Figure 1 described the example environment 100 in which the methods and / or processes according to the embodiments of the present disclosure can be implemented. Below will be combined withFigure 2 The flowchart of method 200 for inference according to an embodiment of the present disclosure is described. Through this method 200, the model can be reasonably and securely deployed, with at least a part of it being deployed in a trusted secure area. In this way, machine learning tasks can be securely and confidentially executed while ensuring resource balance and model performance. Even if some intermediate results are unexpectedly exposed, the possibility of deducing the original data forward is eliminated, thereby protecting the security and privacy of data and computing.
[0037] At block 210, for a layer in the neural network model, it is determined whether the weight matrix of this layer is a non-invertible target matrix. As described above, the computing device 120 loads the neural network model from the storage device 130 and runs the model thereon for specific functionality. The model architecture of the neural network includes one or more layers, and each layer has its corresponding weight matrix. The weight matrix is a key parameter connecting different layers and can have different shapes or patterns. According to an embodiment of the present disclosure, for these neural network layers, it can be determined whether their corresponding weight matrices are non-invertible. If the weight matrix of a certain layer is non-invertible, it means that it is impossible to deduce the original input or information about the layers before this layer from this layer forward. The process of determining the target matrix according to an embodiment of the present disclosure will be further described in detail below.
[0038] At block 202, in response to determining that the weight matrix of this layer is a target matrix, the neural network model is divided into a first part and a second part, where the first part includes this layer and the previous layers before this layer and is deployed within a secure space, and the second part includes the subsequent layers after this layer and is deployed outside the secure space. The secure space according to an embodiment of the present disclosure is an isolated and independent execution environment where data and computing are protected by security, integrity, and privacy. As discussed above, deploying the entire model within the secure space can ensure operational security, but it will result in high deployment costs and a risk of performance degradation. Through the operation at block 202 described herein, an effective deployment division of the model can be achieved.
[0039] According to an embodiment of the present disclosure, the layer determined to include a non-invertible target matrix and the layers before this layer (hereinafter also referred to as previous layers) are deployed within the secure space. In other words, the weight matrix of at least one neural network layer among the neural network layers deployed within the secure space is non-invertible, such that even if the output of any layer among the layers after this layer (hereinafter also referred to as subsequent layers) is known, it is impossible to deduce the original input. The process of securely deploying the model according to an embodiment of the present disclosure will be further described in detail below.
[0040] At block 203, an inference result is obtained based on encrypted inference data by using the partitioned neural network model. According to an embodiment of the present disclosure, the model is reasonably partitioned, with a part of it deployed within a secure space and another part deployed outside the secure space. In this way, feeding the encrypted inference data into the reasonably partitioned neural network model can infer the desired result with high performance while ensuring data and computing security, for example, performing inference using a GPU. Hereinafter, the inference process according to an embodiment of the present disclosure will be described in further detail.
[0041] Therefore, the method 200 for inference according to an embodiment of the present disclosure reasonably deploys part of the model in a trusted secure area, which can provide a secure and reliable deployment environment for machine learning tasks while ensuring resource balance and model performance. Even if some intermediate parameters or data are exposed, it is impossible to deduce further leakage forward, thus protecting data security and privacy.
[0042] Figure 3 FIG. is a schematic diagram illustrating a model input-output stream 300 according to an embodiment of the present disclosure. It should be understood that hereinafter, a neural network model will be used for non-limiting description, and other different models can be adopted according to specific usage needs, such as a deep learning model or a backpropagation neural network model, etc.
[0043] As Figure 3 shown, the input 110 including encrypted inference data will be fed into a neural network model separately arranged in the secure space 310 and the space outside the secure space (hereinafter, for convenience of reference, it will be referred to as the general space 320). In some embodiments, the neural network model is a trained model and has a parameter set corresponding to the input 110 including encrypted inference data, that is, a parameter set adjusted for this data. The complete inference process based on encrypted inference data according to an embodiment of the present disclosure will be introduced below.
[0044] To further ensure the reliability of the secure space 310, a secure space verification process 400 according to an embodiment of the present disclosure can be adopted. Figure 4The flowchart of the security space verification process 400 according to an embodiment of the present disclosure is illustrated. At block 410, the security space 310 can be authenticated, and at block 420, in response to the security space 310 passing the authentication, a secret key 311 can be sent from the data source of the encrypted inference data to the security space 310. The secret key 311 encrypts the inference data at the data source and decrypts the encrypted inference data within the security space 310. The decrypted encrypted inference data 312 is only used by the neural network layer within the security space 310 for corresponding calculations and is invisible outside the security space 310. That is, the outside cannot identify the decrypted encrypted inference data 312. Return reference Figure 3 , the decrypted encrypted inference data 312 can be fed into the first part 313 of the neural network model deployed within the security space 310.
[0045] According to an embodiment of the present disclosure, starting from the first layer of the neural network model, the corresponding weight matrix of each layer is sequentially determined layer by layer to be an irreversible target matrix, where the target matrix is a non-square matrix. In the case of forward propagation, starting from the first neural network layer, the weight matrix of each neural network layer can be sequentially determined layer by layer to be an irreversible matrix. It should be understood that a non-square matrix is one of many examples of the target matrix according to an embodiment of the present disclosure, and the target matrix can also include, but is not limited to, singular matrices, diagonal matrices, upper triangular matrices, lower triangular matrices, atomic matrices, pseudo-inverse matrices, etc.
[0046] As Figure 3 shown, it is determined from left to right whether the weight matrix of each neural network layer is irreversible. It should be understood that sequentially determining layer by layer is only one implementation of determining whether the weight matrix of a layer is irreversible, and the embodiments of the present disclosure are not limited thereto, and may also include determining layer by layer in reverse order (in the case of backpropagation), or randomly selecting a predetermined number of layers (which may not be continuous) and determining whether their weight matrices are irreversible.
[0047] According to an embodiment of the present disclosure, during the layer-by-layer determination, in response to first determining that the corresponding weight matrix of a layer is an irreversible target matrix, a reasonable boundary for model partial deployment can be determined. As Figure 3 exemplarily shown, starting from the first layer, it is sequentially determined layer by layer, and in response to a neural network layer (by way of example and not limitation, such as Figure 3 the layer 314 shown) being a non-square matrix, the layer 314 and the previous layers before it can be determined as the first part of the neural network and can be deployed within the security space 310. In addition, the layer 314 and the subsequent layers after it can be determined as the second part of the neural network and can be deployed within the general space 320 outside the security space 310. As Figure 3As shown, the output of the previous layer will be used as the input of the subsequent layer for corresponding calculations. For example, the output O of layer 314 at the boundary i can be used as the input of the layer after it. Through the calculations of the layers of the neural network model within the secure space 310 and the general space 320, the inference result 330 is finally output.
[0048] In some embodiments, taking the forward propagation of a neural network as an example, the output of a neural network layer can be as shown in the following formula (1):
[0049] O i = σ(M i *O i-1 + b i ) (1)
[0050] where O i is the output of the i-th layer, which is based on the weight matrix M of this layer i , the input of this layer (i.e., the output of the previous layer) O i-1 , the bias term b i , and other weights and parameters σ.
[0051] There are some differences from the training stage. The weight matrix M of each layer of the trained neural network model i and the bias term b i may be known or obtainable. Once the intermediate result O i-1 is exposed, it is very likely to reverse-engineer the results of the previous layer or even the original input, resulting in deteriorated security. According to the embodiments of the present disclosure, at least one neural network layer with an irreversible weight matrix M i is included in the secure space 310, completely eliminating the occurrence of the above reverse-engineering situation, thereby significantly improving security. It should be understood that the above-described model deployment division can also be adopted in the retraining stage, and the process is similar to the above content and will not be elaborated here.
[0052] Figure 5 Illustrates a schematic diagram of the operator mapping process 500 according to an embodiment of the present disclosure. As described above, the neural network model is separately deployed within the secure space 310 and the general space 320 outside the secure space. Hereinafter, the interaction between operators across spaces of the model will be described. According to the operator mapping process 500 of the embodiments of the present disclosure, the cost of the interaction between operators across spaces of the model can be effectively reduced (for example, saving computing and communication resources).
[0053] According to an embodiment of the present disclosure, within the secure space 510, identify at least one layer of secure space operators for the first part of the neural network model deployed within the secure space 510, and within the neural network framework 520, identify at least one layer of secure space outside operators for the second part of the neural network model deployed outside the secure space 510. As Figure 5 Exemplarily shown, the secure space 510 may include secure space operators 511, 512, 513, etc., and the neural network framework 520 may include secure space outside operators 523, 524, etc.
[0054] In addition, according to an embodiment of the present disclosure, in the neural network framework 520, shadow operators corresponding to each secure space operator in the secure space can be determined, for example Figure 5 the shadow operators 521, 522, and 523 shown. In this way, a corresponding relationship is formed between each secure space operator and the corresponding shadow operator. As Figure 5 Schematically shown, the shadow operators 521, 522, and 523 and the secure space outside operators 524, 525 form a continuous operator sequence in the neural network framework 520, and the last shadow operator among the shadow operators 521, 522, and 523 (i.e., Figure 5 the shadow operator 523) is coupled to the first secure space outside operator among the secure space outside operators 524, 525 (i.e., the secure space outside operator 524). Here, the shadow operator can be a dummy operator and can be configured to send and receive instructions to the secure space operator and forward the secure space internal calculation result to the secure space outside operator. The inference process according to the embodiment of the present disclosure will be further described in detail below with the help of Figure 5 Further details.
[0055] According to an embodiment of the present disclosure, in response to the encrypted inference data being decrypted, the first secure space operator 511 among the secure space operators 511, 512, and 513 performs a calculation corresponding to the operator within the secure space 510 based on the decrypted encrypted inference data (hereinafter, simply referred to as the first secure space internal calculation), and sends the first calculation result to the second secure space operator 512 after the first secure space operator 511 and sends the first virtual calculation result to the first shadow operator 521 among the shadow operators 521, 522, and 523 corresponding to the first secure space operator 511. The first virtual calculation result does not mean the true settlement result of the first secure space operator 511, but can indicate the completion of the calculation corresponding to the operator. Additionally or alternatively, the first secure space operator 511 can start the calculation in response to an instruction from the first shadow operator 521.
[0056] Next, in response to the second in - safe - space operator 512 among the in - safe - space operators 511, 512, and 513 receiving the first calculation result from the first in - safe - space operator 511 and the indication from the second shadow operator 522 among the shadow operators 521, 522, and 523, the second in - safe - space operator 512 can perform a second calculation corresponding to this operator, and send the second calculation result to the third in - safe - space operator 513 after the second in - safe - space operator 512 and send a second virtual calculation result to the second shadow operator 522 among the shadow operators 521, 522, and 523 corresponding to the second in - safe - space operator 521. According to an embodiment of the present disclosure, this process can be repeated until the last in - safe - space operator.
[0057] According to an embodiment of the present disclosure, in response to the last in - safe - space operator (e.g., Figure 5 the in - safe - space operator 513 in Figure 5 receiving the penultimate calculation result (i.e., the above - mentioned second result) of its penultimate in - safe - space operator before it (e.g., Figure 5 the in - safe - space operator 512 in
[0058] and the indication from the last shadow operator among the shadow operators corresponding to the in - safe - space operator 513 (e.g.,
[0059] the shadow operator 523 in
[0060] Figure 6 the in - safe - space operator 513 performs the last in - safe - space calculation and sends the in - safe - space calculation result (i.e., the final calculation result in the safe space 510) to the shadow operator 523. Figure 6As shown in the figure, the input 100 including encrypted inference data is fed into the neural network model 610 that is reasonably deployed and partitioned, with a part of it in the secure space to ensure security and another part outside the secure space to ensure performance.
[0061] According to an embodiment of the present disclosure, the encrypted inference data may include an encrypted current driving scene image of the vehicle. Such geographical data is sensitive to the management side or the user side and thus needs to be protected through alignment. The encrypted current driving scene image can be fed into the first part of the neural network model deployed in the secure space to obtain the first output of the first part. The encrypted current driving scene image can be decrypted in the secure space, that is, the decrypted current driving scene image 611.
[0062] According to an embodiment of the present disclosure, the first output of the first part can be fed into the second part of the neural network model deployed outside the secure space to obtain the second output of the second part, and the second output of the second part can be output, which indicates the predicted future travel route for the vehicle, that is, the inference result 620.
[0063] Figure 7 The figure shows a schematic block diagram of an example device 700 that can be used to implement some embodiments according to the present disclosure. As Figure 7 shown in the figure, the device 700 includes a central processing unit (CPU) 701, which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) 702 or the computer program instructions loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0064] Multiple components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0065] The various processes and treatments described above, such as method 200, may be executed by processing unit 701. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of method 200 described above may be performed.
[0066] The present disclosure may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present disclosure.
[0067] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., optical pulses through a fiber optic cable), or electrical signals transmitted through a wire.
[0068] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0069] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0070] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0071] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is created that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0072] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0073] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0074] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.
Claims
1. A method for inference, comprising: For a layer in a neural network model, determining whether the weight matrix of the layer is a target matrix that is non-invertible; In response to determining that the weight matrix of the layer is the target matrix, dividing the neural network model into a first part and a second part, where the first part includes the layer and the previous layers before the layer and is deployed within a secure space, and the second part includes the subsequent layers after the layer and is deployed outside the secure space; And Using the divided neural network model to obtain an inference result based on encrypted inference data.
2. The method according to claim 1, wherein for the layer in the neural network model, determining whether the weight matrix of the layer is the target matrix includes: Starting from the first layer of the neural network model, sequentially determining layer by layer whether the corresponding weight matrix of each layer is the target matrix, the target matrix being a non-square matrix; And Wherein determining that the weight matrix of the layer is the target matrix includes: During the layer-by-layer determination, first determining that the corresponding weight matrix of a layer is the non-square matrix.
3. The method according to claim 1, wherein: The encrypted inference data is encrypted using a key at a data source outside the secure space, and Decryption of the encrypted inference data is provided within the secure space, and the decrypted encrypted inference data cannot be identified outside the secure space.
4. The method according to claim 3, further comprising: Authenticating the secure space; And In response to the secure space passing the authentication, transmitting the key to the secure space for decrypting the encrypted inference data, Wherein the decrypted encrypted inference data is fed into the first part of the neural network model that is deployed within the secure space.
5. The method according to claim 1, wherein the encrypted inference data includes an encrypted current driving scene image of a vehicle, and the method further comprises: Feeding the encrypted current driving scene image into the first part of the neural network model that is deployed within the secure space to obtain a first output of the first part; Feeding the first output of the first part into the second part of the neural network model that is deployed outside the secure space to obtain a second output of the second part; And Outputting the second output of the second part, the second output indicating a predicted future travel route for the vehicle.
6. The method according to claim 1, further comprising: Identifying, within the secure space, operators within the secure space for at least one layer of the first part of the neural network model that is deployed within the secure space; Identifying, in a neural network framework, operators outside the secure space for at least one layer of the second part of the neural network model that is deployed outside the secure space; Determining, in the neural network framework, shadow operators corresponding to each operator within the secure space among the operators within the secure space, Wherein the shadow operator and the out-of-safe-space operator form a continuous operator sequence in the neural network framework, and the last shadow operator in the shadow operators is coupled to the first out-of-safe-space operator in the out-of-safe-space operators.
7. The method according to claim 6, wherein obtaining the inference result based on the encrypted inference data includes: In response to the decryption of the encrypted inference data, the first in-safe-space operator in the in-safe-space operators performs a first in-safe-space calculation in the safe space based on the decrypted encrypted inference data, and sends a first calculation result to a second in-safe-space operator after the first in-safe-space operator and sends a first virtual calculation result to a first shadow operator in the shadow operators corresponding to the first in-safe-space operator; And In response to the last in-safe-space operator in the in-safe-space operators receiving a penultimate calculation result from a penultimate in-safe-space operator before the last in-safe-space operator and an indication from a last shadow operator in the shadow operators corresponding to the last in-safe-space operator, the last in-safe-space operator performs a last in-safe-space calculation and sends the in-safe-space calculation result to the last shadow operator; and The last shadow operator sends the in-safe-space calculation result to the first out-of-safe-space operator in the out-of-safe-space operators.
8. The method according to claim 7, wherein obtaining the inference result based on the encrypted inference data further includes: In response to the first out-of-safe-space operator in the out-of-safe-space operators receiving the in-safe-space calculation result from the last shadow operator in the shadow operators, each out-of-safe-space operator in the out-of-safe-space operators sequentially performs corresponding calculations operator by operator until an out-of-safe-space calculation result is generated.
9. The method according to claim 1, wherein: The safe space is configured to perform secure calculations including decryption and authentication therein, and The safe space is configured to have a size corresponding to the computing resources.
10. The method according to claim 1, wherein: The neural network model is a trained model and has an adjusted parameter set corresponding to input data including the encrypted inference data.
11. An electronic device, comprising: A processor; And A memory, the memory being coupled to the processor and storing instructions that, when executed by the processor, cause the device to perform actions, the actions including: For a layer in a neural network model, determining whether the weight matrix of the layer is an irreversible target matrix; In response to determining that the weight matrix of the layer is the target matrix, dividing the neural network model into a first part and a second part, wherein the first part includes the layer and previous layers before the layer and is deployed in a safe space, and the second part includes subsequent layers after the layer and is deployed outside the safe space; and Using the partitioned neural network model, an inference result is obtained based on encrypted inference data.
12. The apparatus according to claim 11, wherein determining whether the weight matrix of the layer in the neural network model is the target matrix includes: starting from the first layer of the neural network model, sequentially determining whether the corresponding weight matrix of each layer is the target matrix layer by layer, the target matrix being a non-square matrix; and wherein determining that the weight matrix of the layer is the target matrix includes: during the layer-by-layer determination, it is first determined that the corresponding weight matrix of a layer is the non-square matrix.
13. The apparatus according to claim 11, wherein: the encrypted inference data is encrypted using a key at a data source outside the secure space, and decryption of the encrypted inference data is provided within the secure space, and the decrypted encrypted inference data cannot be identified outside the secure space.
14. The apparatus according to claim 13, the action further includes: authenticating the secure space; and in response to the secure space passing the authentication, transmitting the key to the secure space for decrypting the encrypted inference data, wherein the decrypted encrypted inference data is fed into the first part of the neural network model deployed within the secure space.
15. The apparatus according to claim 11, wherein the encrypted inference data includes an encrypted current driving scene image of a vehicle, the action further includes: feeding the encrypted current driving scene image into the first part of the neural network model deployed within the secure space to obtain a first output of the first part; feeding the first output of the first part into the second part of the neural network model deployed outside the secure space to obtain a second output of the second part; and outputting the second output of the second part, the second output indicating a predicted future travel route for the vehicle.
16. The apparatus according to claim 11, the action further includes: identifying in the secure space operators within the secure space for at least one layer of the first part of the neural network model deployed within the secure space; identifying in the neural network framework operators outside the secure space for at least one layer of the second part of the neural network model deployed outside the secure space; determining in the neural network framework shadow operators corresponding to each operator within the secure space, wherein the shadow operators and the operators outside the secure space form a continuous sequence of operators in the neural network framework, and the last shadow operator in the shadow operators is coupled to the first operator outside the secure space.
17. The apparatus according to claim 16, wherein obtaining the inference result based on the encrypted inference data includes: In response to the decryption of the encrypted inference data, the first in-safe-space operator among the in-safe-space operators in the safe space performs a first in-safe-space calculation within the safe space based on the decrypted encrypted inference data, and sends a first calculation result to a second in-safe-space operator after the first in-safe-space operator and sends a first virtual calculation result to a first shadow operator among the shadow operators corresponding to the first in-safe-space operator; and in response to the last in-safe-space operator among the in-safe-space operators in the safe space receiving a penultimate calculation result from a second-to-last in-safe-space operator before the last in-safe-space operator and an indication from a last shadow operator among the shadow operators corresponding to the last in-safe-space operator, the last in-safe-space operator performs a last in-safe-space calculation and sends the in-safe-space calculation result to the last shadow operator; and the last shadow operator sends the in-safe-space calculation result to a first out-safe-space operator among the out-safe-space operators outside the safe space.
18. The apparatus according to claim 17, wherein obtaining the inference result based on the encrypted inference data further comprises: in response to the first out-safe-space operator among the out-safe-space operators outside the safe space receiving the in-safe-space calculation result from the last shadow operator among the shadow operators, each out-safe-space operator among the out-safe-space operators sequentially performs a corresponding calculation operator by operator until an out-safe-space calculation result is generated.
19. The apparatus according to claim 11, wherein: the neural network model is a trained model and has an adjusted set of parameters corresponding to input data including the encrypted inference data.
20. A computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and comprising computer-executable instructions that, when executed by a processor of a computer, cause the computer to: determine, for a layer in a neural network model, whether a weight matrix of the layer is an irreversible target matrix; in response to determining that the weight matrix of the layer is the target matrix, divide the neural network model into a first part and a second part, wherein the first part includes the layer and previous layers before the layer and is deployed within a safe space, and the second part includes subsequent layers after the layer and is deployed outside the safe space; and utilize the divided neural network model to obtain an inference result based on encrypted inference data.