Security Inference Method, Device, Electronic Device and Storage Medium of a Model

By using randomly generated permutation matrix to perform matrix permutation processing when the model is deployed to the cloud platform, the problem of model parameters and inference data leakage is solved, and efficient privacy protection and rapid inference process are achieved.

CN119990336BActive Publication Date: 2025-06-24PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510466565.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-24
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Breach of model parameters and inference data can lead to serious privacy risks when a model is deployed to a cloud platform.

Method used

By randomly generating a permutation matrix between the model developer side and the cloud platform, matrix permutation processing is performed to generate hidden parameter sets and inference data, and recovery and reconstruction processing is performed on the user side to achieve privacy protection of model parameters and inference data.

Benefits of technology

It effectively reduces the risk of leakage of model parameters and inference data, improves the security of private data, and reduces the calculation and communication overhead caused by ciphertext calculations, and improves processing speed and communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990336B_ABST
    Figure CN119990336B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device and storage medium for secure inference of a model, relating to the technical field of model privacy inference. In the secure inference method, the cloud platform cannot obtain the parameter set of the model based on each hidden parameter set; the cloud platform cannot obtain the original inference data based on the second hidden inference; the user side performs recovery and reconstruction processing based on the first permutation matrix, the first inference result and the second inference result to obtain the target inference result. The user side cannot obtain the parameter set of the model, so the user side cannot obtain the privacy of the model developer side. Moreover, the present application does not need to convert the original inference data and the parameter set of the model into ciphertext by using a traditional protocol, but performs matrix multiplication processing on the parameter set of the model and a randomly generated permutation matrix, saving the computational and communication overhead caused by ciphertext calculation, with a faster processing speed and a smaller communication overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model privacy inference, and particularly to a secure inference method, device, electronic device, and storage medium for a model. Background Art

[0002] In related technologies, more and more models are deployed on cloud platforms to provide high-quality services for customers, such as chatting, virtual assistants, and code generation. However, this service model requires the model developer side to upload model parameters to the cloud platform and the user side to upload inference data to the cloud platform. This process may lead to the leakage of model parameters and inference data, resulting in a serious risk of privacy leakage. Summary of the Invention

[0003] This application aims to at least solve one of the technical problems existing in the prior art. For this purpose, this application proposes a secure inference method, device, electronic device, and storage medium for a model, which can protect the privacy of the parameter set and inference data of the model, reduce the risk of leakage of private data, and thus improve the security of private data.

[0004] To achieve the above object, a first aspect embodiment of this application provides a secure inference method for a model, which is applied to the model developer side, and the model is provided on the model developer side;

[0005] The method includes:

[0006] Randomly generate a first permutation matrix based on a preset input sequence length, randomly generate a second permutation matrix based on the dimension of the model, and randomly generate a third permutation matrix based on the dimension of the linear layer of the feed-forward neural network of the model;

[0007] Perform matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set, perform matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set, and perform matrix permutation processing on the second permutation matrix, the third permutation matrix, and the linear layer parameter set in the feed-forward neural network to obtain a third hidden parameter set;

[0008] Receive the first hidden inference data sent by the user side, and perform first privacy protection inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a first inference result;

[0009] Send the first set of hidden parameters, the second set of hidden parameters, and the third set of hidden parameters to the cloud platform, so that the cloud platform performs second privacy-preserving inference based on the second hidden inference data, the first set of hidden parameters, the second set of hidden parameters, and the third set of hidden parameters, obtains a second inference result, and sends the second inference result to the client. The first hidden inference data and the second hidden inference data are obtained by the client processing the original inference data based on the secret sharing algorithm;

[0010] Send the first permutation matrix and the first inference result to the client, so that the client performs restoration and reconstruction processing based on the first permutation matrix, the first inference result, and the second inference result to obtain a target inference result.

[0011] To achieve the above object, a second aspect embodiment of the present application provides a secure inference method for a model, which is applied to a cloud platform. The method includes:

[0012] Receive the first set of hidden parameters, the second set of hidden parameters, and the third set of hidden parameters sent by the model developer side; wherein, the first set of hidden parameters is obtained by the model developer side based on the attention mechanism parameter set of the linear layer of the model and a randomly generated first permutation matrix; the second set of hidden parameters is obtained by the model developer side based on the embedding layer parameter set of the model and a randomly generated second permutation matrix; the third set of hidden parameters is obtained by the model developer side based on the linear layer parameter set of the pre-feed neural network of the model and a randomly generated third permutation matrix;

[0013] Receive the second hidden inference data sent by the client;

[0014] Randomly generate a first permutation matrix based on a preset input sequence length; randomly generate a second permutation matrix based on the dimension of the model; randomly generate a third permutation matrix based on the dimension of the linear layer of the feed-forward neural network of the model;

[0015] Perform matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first set of hidden parameters; perform matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second set of hidden parameters; perform matrix permutation processing on the second permutation matrix, the third permutation matrix, and the linear layer parameter set in the feed-forward neural network to obtain a third set of hidden parameters;

[0016] Receive the first hidden inference data sent by the client, and perform second privacy-preserving inference based on the first hidden inference data, the first set of hidden parameters, the second set of hidden parameters, and the third set of hidden parameters to obtain a second inference result;

[0017] Send the second inference result to the client, so that the client performs restoration and reconstruction processing based on the first permutation matrix, the first inference result, and the second inference result to obtain the target inference result; wherein, the first inference result is obtained by the model developer side performing first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set.

[0018] To achieve the above object, the third aspect embodiment of the present application provides a secure inference method for a model, which is applied to a secure inference system. The secure inference system includes a model developer side, a client, and a cloud platform;

[0019] The method includes:

[0020] The model developer side randomly generates a first permutation matrix based on a preset input sequence length; randomly generates a second permutation matrix based on the dimension of the model; randomly generates a third permutation matrix based on the dimension of the linear layer of the feedforward neural network of the model;

[0021] The model developer side performs matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set; performs matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set; performs matrix permutation processing on the second permutation matrix, the third permutation matrix, and the linear layer parameter set in the feedforward neural network to obtain a third hidden parameter set;

[0022] The model developer side sends the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to the cloud platform;

[0023] The client processes the original inference data based on the secret sharing algorithm to obtain first hidden inference data and second hidden inference data, sends the first hidden inference data to the model developer side, and sends the second hidden inference data to the cloud platform;

[0024] The model developer side performs first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a first inference result; and sends the first inference result to the client;

[0025] The cloud platform performs second privacy-preserving inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a second inference result; and sends the second inference result to the client;

[0026] The client performs a restoration and reconstruction process based on the first permutation matrix, the first inference result, and the second inference result to obtain a target inference result.

[0027] To achieve the above object, an embodiment of the fourth aspect of the present application provides a secure inference device for a model, which is applied to a model developer side, and the model is provided on the model developer side;

[0028] The device includes:

[0029] A generation module, configured to randomly generate a first permutation matrix based on a preset input sequence length; randomly generate a second permutation matrix based on the dimension of the model; randomly generate a third permutation matrix based on the dimension of the linear layer of the feedforward neural network of the model;

[0030] A parameter processing module, configured to perform matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set; perform matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set; perform matrix permutation processing on the second permutation matrix, the third permutation matrix, and the linear layer parameter set in the feedforward neural network to obtain a third hidden parameter set;

[0031] An inference module, configured to receive first hidden inference data sent by the client, and perform first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a first inference result;

[0032] A first sending module, configured to send the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to a cloud platform, so that the cloud platform performs second privacy-preserving inference based on second hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a second inference result, and send the second inference result to the client; wherein, the first hidden inference data and the second hidden inference data are obtained by the client processing original inference data based on a secret sharing algorithm;

[0033] A second sending module, configured to send the first permutation matrix and the first inference result to the client, so that the client performs a restoration and reconstruction process based on the first permutation matrix, the first inference result, and the second inference result to obtain a target inference result.

[0034] To achieve the above object, an embodiment of the fifth aspect of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the security inference method of the model described in any one of the embodiments of the first aspect, the second aspect, and the third aspect.

[0035] To achieve the above object, an embodiment of the sixth aspect of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the security inference method of the model described in any one of the embodiments of the first aspect, the second aspect, and the third aspect.

[0036] According to the security inference method, device, electronic device, and storage medium of the embodiments of the present application, the user side processes the original inference data based on the secret sharing algorithm to obtain the first hidden inference data and the second hidden inference data; the model developer side performs the first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the first inference result. During the process of performing the first privacy-preserving inference, since the model developer side cannot obtain the second hidden inference data; therefore, the model developer side cannot obtain the original inference data based on the first hidden inference data, so the model developer side cannot obtain the privacy of the user side. The cloud platform performs the second privacy-preserving inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the second inference result; the cloud platform cannot obtain the parameter set of the model based on each hidden parameter set, so the cloud platform cannot obtain the privacy of the model developer side; the cloud platform cannot obtain the original inference data based on the second hidden inference, so the cloud platform cannot obtain the privacy of the user side; the user side performs recovery and reconstruction processing based on the first permutation matrix, the first inference result, and the second inference result to obtain the target inference result. The user side cannot obtain the parameter set of the model, so the user side cannot obtain the privacy of the model developer side. In this way, the present application can protect the privacy of the model parameters and inference data, reduce the risk of privacy leakage, and thus improve the security of privacy data. Moreover, compared with the traditional security inference method, the present application does not need to convert the original inference data and the parameter set of the model into ciphertext using a traditional protocol, but performs matrix multiplication processing on the parameter set of the model and a randomly generated permutation matrix, and performs secret sharing algorithm processing on the original inference data, saving the computational and communication overhead caused by ciphertext calculation. Therefore, the security inference method of the present application has a faster processing speed and a smaller communication overhead.

[0037] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The following further describes the present application in conjunction with the accompanying drawings and embodiments, where:

[0039] Figure 1 is a schematic structural diagram of the security inference system according to an embodiment of the present application;

[0040] Figure 2 is a flowchart of the steps of a security inference method for a model applied to the model developer side;

[0041] Figure 3 is Figure 2 a specific flowchart of step S230 in

[0042] Figure 4 is Figure 3 a specific flowchart of step S340 in

[0043] Figure 5 is Figure 4 a specific flowchart of the steps of step S460 in

[0044] Figure 6 is Figure 5 a specific flowchart of step S570 in

[0045] Figure 7 is a security inference method for a model applied to a cloud platform according to an embodiment of the present application;

[0046] Figure 8 is for an embodiment of the present application applied to Figure 1 a flowchart of a security inference method for a model of a security inference system;

[0047] Figure 9 is a schematic diagram of functional modules of a security inference device for a model according to an embodiment of the present application;

[0048] Figure 10 is a schematic hardware structure diagram of an electronic device according to an embodiment of the present application. Specific Embodiments

[0049] The following details the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application.

[0050] In the description of the present application, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present application.

[0051] In the description of the present application, the meaning of "several" is more than one, the meaning of "multiple" is more than two, and understandings such as "greater than", "less than", "exceeding", etc. do not include the recited number, and understandings such as "above", "below", "within", etc. include the recited number. If there is a description of "first", "second", etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0052] In the description of the present application, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.

[0053] In the description of the present application, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0054] First, analyze several nouns involved in the present application:

[0055] Transformer Model: The Transformer model is an architecture that is revolutionary in the field of natural language processing. The Transformer model includes an attention mechanism and a feed-forward neural network. Self-Attention Mechanism: The self-attention mechanism allows the model to focus on different parts of the input sequence at different positions, thereby capturing the dependencies within the sequence. Specifically, for each position in the input sequence, the self-attention mechanism calculates the correlation (i.e., attention weights) between that position and other positions in the sequence, and then performs a weighted sum on the input sequence based on these weights to generate a new representation. Feed-Forward Neural Network (FFNN): It is the simplest type of neural network. Its neurons are arranged in layers, and each neuron is only connected to the neurons in the previous layer, receiving the output of the previous layer and outputting it to the next layer. There is no feedback between layers. This network structure enables information to flow unidirectionally, from the input layer through the hidden layer and finally to the output layer.

[0056] Secure Multi-Party Computation (SMPC) Protocol: The SMPC protocol is a cryptographic technique that allows multiple parties to perform computations without revealing their private inputs to each other and only returns the computation results to the parties without disclosing any other private information. The goal of SMPC is to achieve computation among multiple parties while protecting data privacy.

[0057] Secret sharing algorithm is an important research topic in the fields of cryptography and information security. It allows the owner of a secret to share the secret with multiple participants while ensuring that a single participant cannot obtain any information about the secret, and the secret can only be reconstructed when multiple participants cooperate. The secret sharing algorithm can split data A into two secret data, namely A0 and A1, [[A]] = [A0] + [A1]. Therefore, A can also be reconstructed based on A0 and A1.

[0058] More and more models are deployed to the cloud platform to provide high-quality services for customers, such as chatting, virtual assistants, and code generation. However, this service model requires the model developer side to upload model parameters to the cloud platform, and the user side to upload inference data to the cloud platform. This process may lead to the leakage of model parameters and inference data, resulting in a serious risk of privacy leakage. To reduce the risk of privacy leakage, traditional technologies usually directly call existing SMPC protocols to achieve privacy-preserving Transformer Inference (PPTI). However, this method requires converting the original inference data into ciphertext, and ciphertext calculation requires high computational overhead and communication overhead, resulting in a very slow privacy-preserving inference process.

[0059] Based on this, the embodiments of this application propose a secure inference method, device, electronic device, and storage medium for a model, which can protect the privacy of the parameter set and inference data of the model, reduce the risk of privacy leakage, and save the computational and communication overhead caused by ciphertext calculation, so that the processing speed is faster and the communication overhead is smaller.

[0060] First, the secure inference system of the embodiments of this application is introduced. Refer to Figure 1 , Figure 1 which is the structural schematic diagram of the secure inference system of the embodiments of this application. The secure inference system includes a model developer side P0, a user side P2, and a cloud platform P1. The model developer side deploys a model, and the model is a Transformer model. The model developer side randomly generates a first permutation matrix, a second permutation matrix, and a third permutation matrix. In Figure 1 , the first permutation matrix is π, the second permutation matrix is π1, and the third permutation matrix is π2. The model developer side multiplies each parameter set of the Transformer model with π, π1, and π2 respectively to obtain multiple hidden parameter sets (this process is the "parameter permutation" in Figure 1 ), and in Figure 1 , represents multiple hidden parameter sets. Then π is sent to the user side, and multiple hidden parameter sets are sent to the cloud platform. The user side processes the original inference data X based on the secret sharing algorithm to obtain a first hidden inference data [X]0 and a second inference hidden data [X]1. The model developer side performs privacy-preserving inference based on multiple hidden parameter sets and [X]0 to obtain a first inference result [Yπ]0. The cloud platform performs privacy-preserving inference based on [X]1 and multiple hidden parameter sets to obtain a second inference result [Yπ]1. Then the user side reconstructs and recovers the inference result based on [Yπ]0 and [Yπ]1 to obtain Yπ, and then permutes Yπ based on π to obtain the target inference result Y.

[0061] It should be noted that the model developer side, the cloud platform, and the user side can be terminals or servers. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a router, a programmable switch, a network card, etc.; the server can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the vehicle recognition model training method, etc., but is not limited to the above forms.

[0062] This application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific operations or implement specific abstract data types. This application can also be practiced in distributed computing environments where operations are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0063] An embodiment of the first aspect of this application provides a method for secure inference of a model. The secure inference method of the embodiment of the first aspect is applied to Figure 1 the model developer side therein. Referring to Figure 2 , Figure 2 is the flowchart of the steps of the method for secure inference of the model applied to the model developer side. The method for secure inference of the model applied to the model developer side includes but is not limited to the following steps:

[0064] Step S210, randomly generate a first permutation matrix based on a preset input sequence length; randomly generate a second permutation matrix based on the dimension of the model; randomly generate a third permutation matrix based on the dimension of the linear layer of the feedforward neural network of the model;

[0065] Specifically, the model developer side randomly generates a set of permutation matrices:

[0066] ;

[0067] The set of permutation matrices is used for permuting model parameters of different dimensions. Among them, π is the first permutation matrix, π1 is the second permutation matrix, and π2 is the third permutation matrix; n is the preset input sequence length, d is the dimension of the model; and k is the dimension of the linear layer of the feed-forward neural network of the model.

[0068] It should be noted that the original inference data at the user end is the input of the Transformer model. Generally, the sequence length of the original inference data is the input sequence length. When the sequence length of the original inference data is less than the input sequence length, the original inference data is supplemented to make the sequence length of the original inference data the same as the input sequence length. When the sequence length of the original inference data is greater than the input sequence length, the original inference data is truncated to make the sequence length of the original inference data the same as the input sequence length. It should be noted that the user can input the original inference data at the user end through devices such as a mouse, keyboard, or touch screen.

[0069] Step S220: Perform matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain the first hidden parameter set; perform matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain the second hidden parameter set; perform matrix permutation processing on the first permutation matrix, the second permutation matrix, the third permutation matrix, and the linear layer parameter set in the feed-forward neural network to obtain the third hidden parameter set;

[0070] Specifically, the attention mechanism parameter set includes , and the first permutation matrix is used to permute the attention mechanism parameter set to obtain the first hidden parameter set, which is expressed as:

[0071] ;

[0072] The embedding layer parameter set includes , and the first permutation matrix is used to permute to obtain the second hidden parameter set, which is .

[0073] The linear layer parameter set in the feed-forward neural network includes , and the third permutation matrix is used to permute the linear layer parameter set in the feed-forward neural network to obtain the third hidden parameter set, which is expressed as:

[0074] .

[0075] It should be noted that the specific illustration of the above attention mechanism parameter set, embedding layer parameter set, and linear layer parameter set in the feed-forward neural network is only an example and should not be construed as a limitation to this application. Additionally, this application does not specifically limit the parameters in the attention mechanism parameter set, embedding layer parameter set, and linear layer parameter set in the feed-forward neural network. The parameters in the attention mechanism parameter set, embedding layer parameter set, and linear layer parameter set in different Transformer models are different.

[0076] Step S230: Receive the first hidden inference data sent by the user side, and perform first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a first inference result.

[0077] Step S240: Send the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to the cloud platform, so that the cloud platform performs second privacy-preserving inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a second inference result, and send the second inference result to the user side; wherein, the first hidden inference data and the second hidden inference data are obtained by the user side processing the original inference data based on the secret sharing algorithm.

[0078] It should be noted that since the cloud platform cannot obtain the first permutation matrix, the second permutation matrix, and the third permutation matrix, the cloud platform cannot obtain the attention mechanism parameter set, the embedding layer parameter set, and the linear layer parameter set in the feed-forward neural network based on the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set. Therefore, the cloud platform cannot obtain the model parameters, protecting the model parameters at the model developer side and reducing the risk of model parameter leakage.

[0079] It is worth noting that to reconstruct the original inference data, it is necessary to simultaneously rely on the first hidden inference data and the second hidden inference data. In step S230, during the process of performing the first privacy-preserving inference, the model developer side cannot obtain the second hidden inference data, so the model developer side cannot reconstruct the original inference data. In step S240, during the process of performing the second privacy-preserving inference by the cloud platform, the cloud platform cannot obtain the first hidden inference data, so the cloud platform cannot reconstruct the original inference data. Therefore, neither the cloud platform nor the model developer side can obtain the original inference data, protecting the original inference data of the user side and reducing the risk of original inference data leakage.

[0080] It should be noted that secret sharing is based on the integer ring , and secret sharing is performed on the ring Z L . The specific process is illustrated as follows:

[0081] , ;

[0082] Wherein, X is the original inference data, [X]0 is the first hidden inference data, [X]1 is the second hidden inference data, L is a preset modulus, and L is 2 32 or 2 64 . X can be reconstructed based on [X]0 and [X]1.

[0083] Step S250: Send the first permutation matrix and the first inference result to the client, so that the client performs recovery and reconstruction processing based on the first permutation matrix, the first inference result, and the second inference result to obtain the target inference result.

[0084] Specifically, in step S250, after the client obtains the first inference result [Yπ]0 and the second inference result [Yπ]1, according to the secret sharing algorithm, Yπ can be reconstructed based on [Yπ]0 and [Yπ]1, and then based on the inverse matrix π of the first permutation matrix π T , such that Yπ is multiplied by π T to obtain the target inference result Y.

[0085] The embodiments of the present application implement the privacy-preserving inference process of the model through the above steps S210 to S250. In this process, the client processes the original inference data based on the secret sharing algorithm to obtain the first hidden inference data and the second hidden inference data; the model developer side performs the first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the first inference result. During the first privacy-preserving inference process, since the model developer side cannot obtain the second hidden inference data; therefore, the model developer side cannot obtain the original inference data based on the first hidden inference data, so the model developer side cannot obtain the privacy of the client. The cloud platform performs the second privacy-preserving inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the second inference result; the cloud platform cannot obtain the parameter set of the model based on each hidden parameter set, so the cloud platform cannot obtain the privacy of the model developer side; the cloud platform cannot obtain the original inference data based on the second hidden inference, so the cloud platform cannot obtain the privacy of the client; the client performs recovery and reconstruction processing based on the first permutation matrix, the first inference result, and the second inference result to obtain the target inference result. The client cannot obtain the parameter set of the model, so the client cannot obtain the privacy of the model developer side. In this way, the present application can protect the privacy of the model parameters and inference data, reduce the risk of privacy leakage, and thus improve the security of privacy data. Moreover, compared with the traditional secure inference method, the present application does not need to convert the original inference data and the parameter set of the model into ciphertext, but performs matrix multiplication processing on the parameter set of the model and the randomly generated permutation matrix, and performs secret sharing algorithm processing on the original inference data, saving the computational and communication overhead caused by ciphertext calculation. Therefore, the secure inference method of the present application has a faster processing speed and a smaller communication overhead.

[0086] In some embodiments, referring to Figure 3 , Figure 3 is Figure 2 a specific process schematic diagram of step S230 in

[0087] Step S310, perform matrix multiplication processing on the second hidden parameter set and the first hidden inference data to obtain the first intermediate matrix;

[0088] It should be noted that the matrix multiplication processing includes MatMul multiplication processing and ScalMul multiplication processing. In mathematics and computer science, matrix multiplication (Matrix Multiplication, MatMul) is a binary operation that multiplies two matrices to obtain a new matrix. Scalar multiplication (Scalar Multiplication, ScalMul) is the multiplication operation of a scalar (plaintext) and a vector or matrix (ciphertext).

[0089] In step S310, ScalMul multiplication is used to process with [X]0, and the first intermediate matrix obtained is . In ScalMul multiplication, is used as the plaintext and [X]0 is used as the ciphertext.

[0090] Step S320: Send the first intermediate matrix to the cloud platform so that the cloud platform performs reconstruction processing based on the first intermediate matrix and the second intermediate matrix to obtain a first reconstruction matrix, input the first reconstruction matrix into a normalization function to obtain a normalized matrix, and process the normalized matrix based on a secret sharing algorithm to obtain a first normalized matrix and a second normalized matrix; wherein, the second intermediate matrix is obtained by the cloud platform based on the second hidden inference data and the second hidden parameter set;

[0091] Specifically, in step S320, the cloud platform performs matrix multiplication processing with [X]1 to obtain the first intermediate matrix as . Then the cloud platform reconstructs with to obtain ; input into the normalization function, specifically:

[0092] ;

[0093] where Zπ is the normalized matrix, LayerNorm is the normalization function. Then, based on the secret sharing algorithm, Zπ is processed to obtain a first normalized matrix [Zπ]0 and a second normalized matrix [Zπ]1.

[0094] Step S330: Receive the first normalized matrix sent by the cloud platform and the first hidden matrix sent by the user side; perform matrix multiplication processing on the first normalized matrix and the first hidden matrix to obtain a first normalized hidden matrix; wherein, the first hidden matrix is obtained by the user side processing a randomly generated fourth permutation matrix based on the secret sharing algorithm;

[0095] Specifically, the user side randomly generates a fourth permutation matrix [α], and then the user side processes the randomly generated fourth permutation matrix based on the secret sharing algorithm to obtain a first hidden matrix [Zπ]0 and a second hidden matrix [Zπ]1. The first hidden matrix is sent to the model developer side, and the second hidden matrix is sent to the cloud platform. Then the model developer side multiplies the first normalized matrix [Zπ]0 and the first hidden matrix α0 through matrix multiplication MatMul to obtain a first normalized hidden matrix [α0Zπ]0.

[0096] It should be noted that the second privacy protection inference process on the cloud platform is carried out simultaneously with and in cooperation with the first privacy protection inference process on the model developer side. Specifically, in the second privacy protection inference process on the cloud platform, the cloud platform multiplies the second normalization matrix [Zπ]1 by the second hidden matrix α1 based on matrix multiplication to obtain the second normalization hidden matrix [α1Zπ]1.

[0097] Step S340: Obtain a first inference result based on the first normalization hidden matrix, the first set of hidden parameters, the second set of hidden parameters, and the third set of hidden parameters.

[0098] Through steps S310 to S330 in the embodiments of the present application, operations of looking up a table and layer normalization on the first hidden inference data are realized, and the first hidden inference data is converted into a vector, that is, the first normalization hidden matrix. The first normalization hidden matrix can be used as the input of the attention mechanism, so that in step S340 on the model developer side, a first inference result is obtained based on the first normalization hidden matrix, the first set of hidden parameters, the second set of hidden parameters, and the third set of hidden parameters.

[0099] In some embodiments, refer to Figure 4 , Figure 4 For Figure 3 a specific process schematic diagram of step S340 in

[0100] Step S410: Perform matrix multiplication processing on the first set of hidden parameters and the first normalization hidden matrix to obtain a first query representation matrix, a first key representation matrix, and a first value representation matrix;

[0101] It should be noted that the ScalMul multiplication is used to process the first set of hidden parameters and the first normalization hidden matrix. Specifically:

[0102] ;

[0103] ;

[0104] ;

[0105] where [Q]0 is the first query representation matrix, [K]0 is the first key representation matrix, and [V]0 represents the first value representation matrix.

[0106] Step S420: Perform matrix multiplication processing on the first value representation matrix and the first hidden matrix to obtain a first attention hidden matrix;

[0107] Specifically, the ScalMul multiplication is used to process the first value representation matrix [V]0 and the first hidden matrix α0 to obtain the first attention hidden matrix as [Vα0]0.

[0108] It should be noted that in the second privacy protection inference process of the cloud platform, the ScalMul multiplication is used to process the first hidden parameter set and the second normalized hidden matrix [α1Zπ]1, specifically as follows:

[0109] ;

[0110] ;

[0111] ;

[0112] Among them, [Q]1 is the second query representation matrix, [K]1 is the second key representation matrix, and [V]1 represents the second value representation matrix. Then the cloud platform uses the ScalMul multiplication to process the second value representation matrix [V]1 and the second hidden matrix α1 to obtain the second attention hidden matrix as [Vα1]1.

[0113] Step S430, obtaining the first intermediate attention matrix based on the first query representation matrix and the first key representation matrix; performing matrix multiplication processing on the first intermediate attention matrix and the first hidden matrix to obtain the second attention hidden matrix;

[0114] Specifically, the ScalMul multiplication is used to process the first query representation matrix and the first key representation matrix to obtain the first intermediate attention matrix, specifically as follows:

[0115] ;

[0116] The first intermediate attention matrix is [O]0, , represents the dimension of the Transformer model, represents the number of heads in the multi-head attention mechanism in the Transformer model, represents the mask matrix.

[0117] Then the ScalMul multiplication is used to perform α0 processing on the first intermediate attention matrix [O]0 and the first hidden matrix to obtain the second attention hidden matrix [Oα0]0.

[0118] It should be noted that in the second privacy protection inference process of the cloud platform, the ScalMul multiplication is used to process the second query representation matrix and the second key representation matrix to obtain the second intermediate attention matrix, specifically as follows:

[0119] ;

[0120] The second intermediate attention matrix is [O]1, , denotes the dimension of the Transformer model, denotes the number of heads in the multi-head attention mechanism in the Transformer model, denotes the mask matrix. Then, the ScalMul multiplication is used to perform α2 processing on the second intermediate attention matrix [O]1 and the second hidden matrix to obtain the third attention hidden matrix [Oα1]1.

[0121] Step S440: Send the second attention hidden matrix to the cloud platform so that the cloud platform can perform reconstruction based on the second attention hidden matrix and the third attention hidden matrix to obtain the attention reconstruction hidden matrix, input the attention reconstruction hidden matrix into the classification function to obtain the classification hidden matrix, and process the classification hidden matrix based on the secret sharing algorithm to obtain the first classification hidden matrix and the second classification hidden matrix; wherein, the third attention hidden matrix is obtained by the cloud platform based on the second normalization matrix and the second hidden matrix, and the second hidden matrix is obtained by the user side processing the fourth permutation matrix based on the secret sharing algorithm;

[0122] Specifically, after obtaining the third attention hidden matrix [Oα1]1, the cloud platform performs reconstruction based on the third attention hidden matrix [Oα1]1 and the second attention hidden matrix [Oα0]0 to obtain the attention reconstruction hidden matrix O1π, and then inputs the attention reconstruction hidden matrix O1π into the classification function, specifically:

[0123] ;

[0124] where O2π is the classification hidden matrix and Softmax is the classification function. Then, the cloud platform processes the classification hidden matrix based on the secret sharing algorithm to obtain the first classification hidden matrix [O2π]0 and the second classification hidden matrix [O2π]1.

[0125] Step S450: Receive the first classification hidden matrix sent by the cloud platform, and obtain the first target attention hidden matrix based on the first classification hidden matrix, the first attention hidden matrix, and the first normalization hidden matrix;

[0126] Specifically, multiply the first classification hidden matrix [O2π]0 by the first attention hidden matrix [Vα0]0 to obtain [O3]0, and then calculate [O4π]0, specifically:

[0127] ;

[0128] Then, make the first target attention hidden matrix [O4]0 = [O4π]0 + [α0Zπ]0, where [α0Zπ]0 is the first normalization hidden matrix.

[0129] It should be noted that, in the second privacy protection inference process of the cloud platform, the second classification hidden matrix [O2π]1 is multiplied by the second attention hidden matrix [Vα1]1 to obtain [O3]1, and then [O4π]1 is calculated as follows:

[0130] ;

[0131] Then, the second target attention hidden matrix [O4]1 = [O4π]1 + [α1Zπ]1, where [α1Zπ]1 is the second normalized hidden matrix.

[0132] Step S460: Obtain the first inference result based on the first target attention hidden matrix, the second hidden parameter set, and the third hidden parameter set.

[0133] In some embodiments, referring to Figure 5 , Figure 5 is Figure 4 a specific step flow diagram of step S460 in

[0134] Step S510: Send the first target attention hidden matrix to the cloud platform, so that the cloud platform reconstructs based on the first target attention hidden matrix and the second target attention hidden matrix to obtain an attention hidden reconstruction matrix; input the attention hidden reconstruction matrix into a normalization function to obtain an attention hidden normalization matrix; and process the attention hidden normalization matrix based on the secret sharing algorithm to obtain a first attention hidden normalization matrix and a second attention hidden normalization matrix; where the second target attention hidden matrix is obtained by the cloud platform based on the second classification hidden matrix and the second hidden matrix.

[0135] Specifically, [O4]0 is sent to the cloud platform, and the cloud platform reconstructs [O4]1 and [O4]0 to obtain O4. Where [O4]1 is the second target attention hidden matrix, and [O4]1 = [O4π]1 + [α1Zπ]1. The cloud platform inputs O4 into the normalization function to obtain an attention hidden normalization matrix L1π, and processes the attention hidden normalization matrix L1π based on the secret sharing algorithm to obtain a first attention hidden normalization matrix [L1π]0 and a second attention hidden normalization matrix [L1π]1.

[0136] Step S520: Receive the first attention hidden normalization matrix sent by the cloud platform; perform matrix multiplication processing on the first attention hidden normalization matrix, the second hidden parameter set, and the third hidden parameter set to obtain a first linear transformation hidden matrix.

[0137] Specifically, the process of obtaining the first linear transformation hidden matrix is expressed as:

[0138] ;

[0139] Among them, [O5π2]0 is the first linear transformation hidden matrix.

[0140] It should be noted that in the second privacy protection inference process of the cloud platform, the process of obtaining the second linear transformation hidden matrix is as follows:

[0141] ;

[0142] Among them, [O5π2]1 is the second linear transformation hidden matrix.

[0143] Step S530, send the first linear transformation hidden matrix to the cloud platform, so that the cloud platform reconstructs based on the first linear transformation hidden matrix and the second linear transformation hidden matrix to obtain a linear transformation hidden reconstruction matrix; and input the linear transformation hidden reconstruction matrix into the Gaussian error linear function to obtain a Gaussian error hidden matrix; process the Gaussian error hidden matrix based on the secret sharing algorithm to obtain a first Gaussian error hidden matrix and a second Gaussian error hidden matrix; among them, the second linear transformation hidden matrix is obtained by the cloud platform based on the second attention hidden normalization matrix;

[0144] Specifically, the cloud platform reconstructs the first linear transformation hidden matrix [O5π2]0 and the second linear transformation hidden matrix [O5π2]1 to obtain a linear transformation hidden reconstruction matrix O5π2. Then input O5π2 into the Gaussian error linear function, specifically:

[0145] ;

[0146] Among them, Gπ2 is the Gaussian error hidden matrix, and GeLU represents the Gaussian error linear function. Then the cloud platform processes Gπ2 based on the secret sharing algorithm to obtain a first Gaussian error hidden matrix [Gπ2]0 and a second Gaussian error hidden matrix [Gπ2]1. Then the cloud platform sends the first Gaussian error hidden matrix [Gπ2]0 to the model developer side.

[0147] Step S540, receive the first Gaussian error hidden matrix sent by the cloud platform; obtain a first addition hidden matrix based on the first Gaussian error hidden matrix and the first attention hidden normalization matrix;

[0148] Specifically, the first addition hidden matrix is expressed as:

[0149] ;

[0150] Among them, [O6π2]0 represents the first addition hidden matrix.

[0151] It should be noted that in the second privacy protection inference process of the cloud platform, the second additive hidden matrix is represented as:

[0152] ;

[0153] where [O6π2]1 represents the second additive hidden matrix.

[0154] Step S550: Send the first additive hidden matrix to the cloud platform so that the cloud platform can reconstruct based on the first additive hidden matrix and the second additive hidden matrix to obtain an additive reconstruction matrix; input the additive reconstruction matrix into a normalization function to obtain an additive hidden normalization matrix; and process the additive hidden normalization matrix based on the secret sharing algorithm to obtain a first additive hidden normalization matrix and a second additive hidden normalization matrix; where the second additive hidden matrix is obtained by the cloud platform based on the second Gaussian error hidden matrix.

[0155] Specifically, the cloud platform reconstructs O6π2 based on [O6π2]0 and [O6π2]1, then inputs O6π2 into the normalization function to obtain an additive hidden normalization matrix L2π, and processes L2π based on the secret sharing algorithm to obtain a first additive hidden normalization matrix [L2π]0 and a second additive hidden normalization matrix [L2π]1. Then the cloud platform sends the first additive hidden normalization matrix [L2π]0 to the model developer side.

[0156] Step S560: Receive the first additive hidden normalization matrix sent by the cloud platform.

[0157] Step S570: Perform parameter adaptation processing on the first additive hidden normalization matrix to obtain a first inference result.

[0158] Through the above steps S510 to S570 in the embodiments of the present application, the present application realizes attention processing on the first hidden inference data by using the attention mechanism to capture information in different subspaces and improve the expression ability of the model. Similarly, in the second privacy protection inference process of the cloud platform, the attention mechanism is used to perform attention processing on the second hidden inference data to capture information in different subspaces and improve the expression ability of the model. Moreover, in steps S510 to S570, the cloud platform cannot obtain specific model parameters nor specific original inference data; the model developer side cannot obtain specific original inference data, so the risk of privacy leakage can be reduced.

[0159] In some embodiments, refer to Figure 6 , Figure 6 is Figure 5 a specific process schematic diagram of step S570. Step S570 includes but is not limited to the following steps:

[0160] Step S610: Perform matrix permutation processing on the first permutation matrix and the linear parameter set of the parameter adaptation layer of the model to obtain a fourth hidden parameter set; perform matrix permutation processing on the second permutation matrix and the non-linear parameter set of the parameter adaptation layer of the model to obtain a fifth hidden parameter set.

[0161] It should be noted that the parameter adaptation layer (Adaptation Layer): In some cases, in order to adapt to different tasks or data sets, one or more parameter adaptation layers may be added to the Transformer model. These layers usually contain linear layers and non-linear layers, which are used to adjust the output of the model to better match the requirements of specific tasks, and combine linear transformation and non-linear activation functions to achieve more complex feature mapping and data transformation.

[0162] Specifically, the linear parameter set of the parameter adaptation layer of the model is Wp, the non-linear Wc of the parameter adaptation layer of the model, the fourth hidden parameter set is Wpπ, and the fifth hidden parameter set is Wcπ.

[0163] Step S620: Perform matrix multiplication processing on the first added hidden normalization matrix and the fourth hidden parameter set to obtain a first adapted hidden matrix.

[0164] Specifically, use MatMul multiplication to multiply the first added hidden normalization matrix [L2π]0 by the fourth hidden parameter set Wpπ to obtain the first adapted hidden matrix [Sπ]0.

[0165] It should be noted that in the second privacy-preserving inference process of the cloud platform, use MatMul multiplication to multiply the second added hidden normalization matrix [L2π]1 by the fourth hidden parameter set Wpπ to obtain the second adapted hidden matrix [Sπ]1.

[0166] Step S630: Send the first adapted hidden matrix, the fourth hidden parameter set, and the fifth hidden parameter set to the cloud platform, so that the cloud platform can reconstruct based on the first adapted hidden matrix and the second adapted hidden matrix to obtain an adapted hidden reconstruction matrix, and input the adapted hidden reconstruction matrix into the hyperbolic tangent function to obtain a hyperbolic hidden matrix; and process the hyperbolic hidden matrix based on the secret sharing algorithm to obtain a first hyperbolic hidden matrix and a second hyperbolic hidden matrix; where the second adapted hidden matrix is obtained by the cloud platform based on the fourth hidden parameter set and the second added hidden normalization matrix.

[0167] Specifically, the cloud platform reconstructs based on the first adapted hidden matrix [Sπ]0 and the second adapted hidden matrix [Sπ]1 to obtain an adapted hidden reconstruction matrix Sπ. Then input the adapted hidden reconstruction matrix Sπ into the hyperbolic tangent function, specifically:

[0168] ;

[0169] Among them, Tπ is the hyperbolic hidden matrix, and Tanh is the hyperbolic tangent function. Then, the cloud platform processes the hyperbolic hidden matrix Tπ based on the secret sharing algorithm to obtain the first hyperbolic hidden matrix [Tπ]0 and the second hyperbolic hidden matrix [Tπ]1, and sends the first hyperbolic hidden matrix [Tπ]0 to the model developer side.

[0170] Step S640: Receive the first hyperbolic hidden matrix sent by the cloud platform, and obtain the first inference result based on the first hyperbolic hidden matrix and the fifth set of hidden parameters.

[0171] Specifically, the model developer side uses ScalMul multiplication to multiply the first hyperbolic hidden matrix [Tπ]0 by the fifth set of hidden parameters Wcπ to obtain the first inference result [Yπ]0. In the second privacy-preserving inference process of the cloud platform, ScalMul multiplication is used to multiply the second hyperbolic hidden matrix [Tπ]1 by the fifth set of hidden parameters Wcπ to obtain the second inference result [Yπ]1.

[0172] Through the above steps S610 to S640 of this application, the first inference result is obtained based on the first addition-hidden normalization matrix. At the same time, the cloud platform obtains the second inference result based on the second addition-hidden normalization matrix. After the model developer side sends the first inference result [Yπ]0 to the user side and the cloud platform sends the second inference result [Yπ]1 to the user side, the user side can reconstruct and restore the inference result based on [Yπ]0 and [Yπ]1 to obtain Yπ, and then permute Yπ based on π to obtain the target inference result Y. In this way, the target inference result is obtained based on the original inference data of the user side, and the original inference data of the user side and the model parameters of the model developer side can be avoided from being leaked during the inference process.

[0173] The second aspect embodiment of this application provides a secure inference method for a model applied to a cloud platform. Refer to Figure 7 , Figure 7 which is the secure inference method for the model applied to the cloud platform in the embodiment of this application. Figure 7 The schematic method includes the following steps:

[0174] Step S710: Receive the first set of hidden parameters, the second set of hidden parameters, and the third hidden parameter sent by the model developer side; among them, the first set of hidden parameters is obtained by the model developer side based on the attention mechanism parameter set of the linear layer of the model and the randomly generated first permutation matrix; the second set of hidden parameters is obtained by the model developer side based on the embedding layer parameter set of the model and the randomly generated second permutation matrix; the third set of hidden parameters is obtained by the model developer side based on the linear layer parameter set of the prefrontal neural network of the model and the randomly generated third permutation matrix;

[0175] Step S720: Receive the second hidden inference data sent by the client;

[0176] Step S730: Randomly generate a first permutation matrix based on a preset input sequence length; randomly generate a second permutation matrix based on the dimension of the model; randomly generate a third permutation matrix based on the dimension of the linear layer of the model's feed-forward neural network;

[0177] Step S740: Perform matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set; perform matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set; perform matrix permutation processing on the first permutation matrix, the second permutation matrix, the third permutation matrix and the linear layer parameter set in the feed-forward neural network to obtain a third hidden parameter set;

[0178] Step S750: Receive the first hidden inference data sent by the client, and perform second privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to obtain a second inference result;

[0179] Step S760: Send the second inference result to the client so that the client performs recovery and reconstruction processing based on the first permutation matrix, the first inference result and the second inference result to obtain a target inference result; wherein, the first inference result is obtained by the model developer side performing first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set.

[0180] The security inference method of the model applied to the cloud platform in the embodiments of the present application realizes the privacy protection inference process through steps S710 to S760. In this process, the user side processes the original inference data based on the secret sharing algorithm to obtain the first hidden inference data and the second hidden inference data; the model developer side performs the first privacy protection inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the first inference result. During the process of performing the first privacy protection inference, since the model developer side cannot obtain the second hidden inference data, the model developer side cannot obtain the original inference data based on the first hidden inference data, so the model developer side cannot obtain the privacy of the user side. The cloud platform performs the second privacy protection inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the second inference result; the cloud platform cannot obtain the parameter set of the model based on each hidden parameter set, so the cloud platform cannot obtain the privacy of the model developer side; the cloud platform cannot obtain the original inference data based on the second hidden inference, so the cloud platform cannot obtain the privacy of the user side; the user side performs restoration and reconstruction processing based on the first permutation matrix, the first inference result, and the second inference result to obtain the target inference result. The user side cannot obtain the parameter set of the model, so the user side cannot obtain the privacy of the model developer side. In this way, the present application can protect the privacy of the model parameters and inference data and reduce the risk of privacy leakage. Compared with the traditional security inference method, the present application does not need to convert the original inference data and the parameter set of the model into ciphertext, but performs matrix multiplication processing on the parameter set of the model and the randomly generated permutation matrix, and performs secret sharing algorithm processing on the original inference data, saving the computing and communication overhead caused by ciphertext calculation. Therefore, the security inference method of the present application has a faster processing speed and a smaller communication overhead.

[0181] The third aspect of the present application provides a security inference method for a model of a security inference system applied to Figure 1 a Figure 8 , Figure 8 which is Figure 1 a schematic flowchart of the security inference method for a model of a security inference system applied to the embodiments of the present application. Figure 8 The method shown includes:

[0182] Step S810, the model developer side randomly generates a first permutation matrix based on a preset input sequence length; randomly generates a second permutation matrix based on the dimension of the model; randomly generates a third permutation matrix based on the dimension of the linear layer of the feedforward neural network of the model;

[0183] Step S820: The model developer side performs matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set; performs matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set; performs matrix permutation processing on the first permutation matrix, the second permutation matrix, the third permutation matrix and the linear layer parameter set in the feed-forward neural network to obtain a third hidden parameter set;

[0184] Step S830: The model developer side sends the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to the cloud platform;

[0185] Step S840: The user side processes the original inference data based on the secret sharing algorithm to obtain a first hidden inference data and a second hidden inference data, sends the first hidden inference data to the model developer side, and sends the second hidden inference data to the cloud platform;

[0186] Step S850: The model developer side performs first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to obtain a first inference result; and sends the first inference result to the user side;

[0187] Step S860: The cloud platform performs second privacy-preserving inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to obtain a second inference result; and sends the second inference result to the user side;

[0188] Step S870: The user side performs recovery and reconstruction processing based on the first permutation matrix, the first inference result and the second inference result to obtain the target inference result.

[0189] The security inference method of the model applied to the security inference system in the embodiments of the present application realizes the privacy protection inference process through steps S810 to S870. In this process, the client processes the original inference data based on the secret sharing algorithm to obtain the first hidden inference data and the second hidden inference data; the model developer side performs the first privacy protection inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the first inference result. During the process of performing the first privacy protection inference, since the model developer side cannot obtain the second hidden inference data, the model developer side cannot obtain the original inference data based on the first hidden inference data. Therefore, the model developer side cannot obtain the privacy of the client. The cloud platform performs the second privacy protection inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the second inference result; the cloud platform cannot obtain the parameter set of the model based on each hidden parameter set. Therefore, the cloud platform cannot obtain the privacy of the model developer side; the cloud platform cannot obtain the original inference data based on the second hidden inference. Therefore, the cloud platform cannot obtain the privacy of the client; the client performs recovery and reconstruction processing based on the first permutation matrix, the first inference result, and the second inference result to obtain the target inference result. The client cannot obtain the parameter set of the model. Therefore, the client cannot obtain the privacy of the model developer side. In this way, the present application can protect the privacy of model parameters and inference data and reduce the risk of privacy leakage. Compared with the traditional security inference method, the present application does not need to convert the original inference data and the parameter set of the model into ciphertext, but performs matrix multiplication processing on the parameter set of the model and the randomly generated permutation matrix, and performs secret sharing algorithm processing on the original inference data, saving the computational and communication overhead caused by ciphertext calculation. Therefore, the security inference method of the present application has a faster processing speed and a smaller communication overhead.

[0190] Table 1

[0191]

[0192] Referring to Table 1, Table 1 shows the time comparison of privacy protection inference for models with different Transformer architectures using the security inference method of the model of the present application and other security inference methods of models in related technologies. The models with different Transformer architectures include BERT series models, whose encoder structure is mainly used for natural language understanding (NLU) tasks, and GPT-2 series models, whose decoder structure is mainly used for natural language generation (NLG). The BERT series models include the BERT BASE model, the BERT LARGE model. The GPT-2 series models include the GPT-2 BASE model, the GPT-2 LARGE model.

[0193] Referring to Table 1, the security inference methods of other models in the related art include PUMA, MPCFormer, and SecFormer. PUMA realizes privacy-preserving inference through multi-party secure computation (MPC) technology. It adopts a 2-out-of-3 Replicated Secret Sharing scheme, similar to the ABY3 protocol, and can quickly and securely perform Transformer model inference in a three-party scenario. MPCFormer is a framework that combines multi-party secure computation (MPC) and knowledge distillation (KD) to achieve fast, efficient, and private Transformer model inference. SecFormer is a framework for privacy-preserving inference, aiming to provide fast, efficient, and accurate inference services for Transformer models. SecFormer improves the inference performance of Transformer models in privacy-preserving scenarios by combining secure multi-party computation (Secure Multi-Party Computation, SMPC) and optimized numerical calculation methods.

[0194] In Table 1, the Embedding layer is the embedding layer, and the Adaptation layer is the parameter adaptation layer. The unit of the data in Table 1 is seconds. The data in Table 1 was obtained from experiments on two servers equipped with A100 graphics cards, and the server bandwidth was set to 3 Gbps, with a round-trip delay of 0.8 milliseconds.

[0195] As can be seen from Table 1, for BERT series models, the speed of performing PPTI using the security inference method of the model proposed in this application is 5.1 - 30.3 times faster than the existing methods. For GPT-2 series models, the speed of performing PPTI using the security inference method of the model proposed in this application is 5.0 - 27.2 times faster than the existing methods.

[0196] The fourth aspect embodiment of this application provides a security inference device for a model, which is applied to the model developer side. Referring to Figure 9 , Figure 9 is a schematic diagram of the functional modules of the security inference device for the model in the embodiment of this application. The device includes:

[0197] A generation module 910, configured to randomly generate a first permutation matrix based on a preset input sequence length; randomly generate a second permutation matrix based on the dimension of the model; randomly generate a third permutation matrix based on the dimension of the linear layer of the feed-forward neural network of the model;

[0198] A parameter processing module 920, configured to perform matrix permutation processing on the attention mechanism parameter set of the first permutation matrix and the linear layer of the model to obtain a first hidden parameter set; perform matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set; perform matrix permutation processing on the second permutation matrix, the third permutation matrix and the linear layer parameter set in the feed-forward neural network to obtain a third hidden parameter set;

[0199] An inference module 930, configured to receive first hidden inference data sent by the user side, and perform first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to obtain a first inference result;

[0200] A first sending module 940, configured to send the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to the cloud platform, so that the cloud platform performs second privacy-preserving inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to obtain a second inference result, and send the second inference result to the user side; wherein, the first hidden inference data and the second hidden inference data are obtained by the user side processing the original inference data based on the secret sharing algorithm;

[0201] A second sending module 950, configured to send the first permutation matrix and the first inference result to the user side, so that the user side performs restoration and reconstruction processing based on the first permutation matrix, the first inference result and the second inference result to obtain a target inference result.

[0202] The security inference device of the model in the embodiment of the present application is used to execute the security inference method of the model in the first aspect embodiment of the present application. When executing the method, the client processes the original inference data based on the secret sharing algorithm to obtain the first hidden inference data and the second hidden inference data; the model developer side performs the first privacy-preserving inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the first inference result. During the process of performing the first privacy-preserving inference, since the model developer side cannot obtain the second hidden inference data; therefore, the model developer side cannot obtain the original inference data based on the first hidden inference data, so the model developer side cannot obtain the privacy of the client. The cloud platform performs the second privacy-preserving inference based on the second hidden inference data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain the second inference result; the cloud platform cannot obtain the parameter set of the model based on each hidden parameter set, so the cloud platform cannot obtain the privacy of the model developer side; the cloud platform cannot obtain the original inference data based on the second hidden inference, so the cloud platform cannot obtain the privacy of the client; the client performs recovery and reconstruction processing based on the first permutation matrix, the first inference result, and the second inference result to obtain the target inference result. The client cannot obtain the parameter set of the model, so the client cannot obtain the privacy of the model developer side. In this way, the present application can protect the privacy of the model parameters and inference data and reduce the risk of privacy leakage. And compared with the traditional security inference method, the present application does not need to convert the original inference data and the parameter set of the model into ciphertext, but performs matrix multiplication processing on the parameter set of the model and the randomly generated permutation matrix, and performs secret sharing algorithm processing on the original inference data, saving the computational and communication overhead caused by ciphertext calculation. Therefore, the security inference method of the present application has a faster processing speed and a smaller communication overhead.

[0203] It should be noted that the specific implementation manner of the security inference device of the model is basically the same as the specific embodiment of the privacy protection method of the model in the first aspect embodiment, and will not be elaborated here. On the premise of meeting the requirements of the embodiment of the present application, other functional modules can be set in the security inference device of the model to implement the privacy protection method of the model in the first aspect embodiment.

[0204] The embodiment of the fifth aspect of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the security inference method of the model in the above embodiment. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0205] In one embodiment, referring to Figure 10 , Figure 10 schematically shows the hardware structure of the electronic device in the embodiment of the present application. The electronic device includes:

[0206] The processor 101 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0207] The memory 102 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 102 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 102, and the processor 101 is called to execute the security inference method of the model in the embodiments of the present application;

[0208] The input / output interface 103 is used to implement information input and output;

[0209] The communication interface 104 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0210] The bus 105 transmits information between the various components of the device (such as the processor 101, the memory 102, the input / output interface 103, and the communication interface 104);

[0211] Among them, the processor 101, the memory 102, the input / output interface 103, and the communication interface 104 achieve communication connections with each other inside the device through the bus 105.

[0212] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the security inference method of the model in the embodiment of the first aspect.

[0213] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0214] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0215] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0216] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0217] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0218] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above figures are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0219] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0220] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0221] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0222] In addition, each functional unit in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0223] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0224] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall fall within the scope of the rights of the embodiments of this application.

Claims

1. A secure reasoning method for a model, characterized in that: Applied to a model developer end, the model developer end being provided with the model; The method comprises: randomly generating a first permutation matrix based on a preset input sequence length, randomly generating a second permutation matrix based on the dimension of the model, and randomly generating a third permutation matrix based on the dimension of a linear layer of a feedforward neural network of the model; Performing matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set, performing matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set, and performing matrix permutation processing on the first permutation matrix, the second permutation matrix, the third permutation matrix, and the linear layer parameter set in the feedforward neural network to obtain a third hidden parameter set; receiving first hidden reasoning data sent by a user terminal, performing first privacy protection reasoning based on the first hidden reasoning data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a first reasoning result; Sending the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to the cloud platform, so that the cloud platform performs a second privacy protection reasoning based on the second hidden reasoning data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a second reasoning result, and sending the second reasoning result to the user end; the first hidden reasoning data and the second hidden reasoning data are obtained by the user end processing the original reasoning data based on a secret sharing algorithm; The first permutation matrix and the first inference result are sent to the user terminal, so that the user terminal performs recovery and reconstruction processing based on the first permutation matrix, the first inference result and the second inference result to obtain a target inference result.

2. The security reasoning method of the model according to claim 1, characterized in that: The performing a first privacy protection reasoning based on the first hidden reasoning data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a first reasoning result includes: Performing matrix multiplication processing on the second hidden parameter set and the first hidden inference data to obtain a first intermediate matrix; Sending the first intermediate matrix to the cloud platform so that the cloud platform performs reconstruction processing based on the first intermediate matrix and the second intermediate matrix to obtain a first reconstructed matrix, and inputting the first reconstructed matrix into a normalization function to obtain a normalized matrix, and processing the normalized matrix based on a secret sharing algorithm to obtain a first normalized matrix and a second normalized matrix; wherein the second intermediate matrix is ​​obtained by the cloud platform based on the second hidden inference data and the second hidden parameter set; Receiving the first normalized matrix sent by the cloud platform and the first hidden matrix sent by the user terminal; performing matrix multiplication processing on the first normalized matrix and the first hidden matrix to obtain a first normalized hidden matrix; wherein the first hidden matrix is ​​obtained by the user terminal processing a randomly generated fourth permutation matrix based on a secret sharing algorithm; The first inference result is obtained based on the first normalized hidden matrix, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set.

3. The security reasoning method of the model according to claim 2, characterized in that: The obtaining the first inference result based on the first normalized hidden matrix, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set includes: Performing matrix multiplication processing on the first hidden parameter set and the first normalized hidden matrix to obtain a first query representation matrix, a first key representation matrix, and a first value representation matrix; Performing matrix multiplication processing on the first value representation matrix and the first hidden matrix to obtain a first attention hidden matrix; Obtaining a first intermediate attention matrix based on the first query representation matrix and the first key representation matrix; performing matrix multiplication processing on the first intermediate attention matrix and the first hidden matrix to obtain a second attention hidden matrix; Sending the second attention hidden matrix to the cloud platform so that the cloud platform reconstructs the second attention hidden matrix and the third attention hidden matrix to obtain an attention reconstruction hidden matrix, and inputting the attention reconstruction hidden matrix into a classification function to obtain a classification hidden matrix, and processing the classification hidden matrix based on a secret sharing algorithm to obtain a first classification hidden matrix and a second classification hidden matrix; wherein the third attention hidden matrix is ​​obtained by the cloud platform based on the second normalized matrix and the second hidden matrix, and the second hidden matrix is ​​obtained by the user end by processing the fourth permutation matrix based on a secret sharing algorithm; Receiving a first classification hidden matrix sent by the cloud platform, and obtaining a first target attention hidden matrix based on the first classification hidden matrix, the first attention hidden matrix, and the first normalized hidden matrix; The first inference result is obtained based on the first target attention hidden matrix, the second hidden parameter set and the third hidden parameter set.

4. The safety reasoning method of the model according to claim 3, characterized in that: The obtaining the first inference result based on the first target attention hidden matrix, the second hidden parameter set and the third hidden parameter set includes: Sending the first target attention hidden matrix to the cloud platform, so that the cloud platform reconstructs the first target attention hidden matrix and the second target attention hidden matrix to obtain an attention hidden reconstruction matrix; and inputting the attention hidden reconstruction matrix into a normalization function to obtain an attention hidden normalized matrix; and processing the attention hidden normalized matrix based on a secret sharing algorithm to obtain a first attention hidden normalized matrix and a second attention hidden normalized matrix; wherein the second target attention hidden matrix is ​​obtained by the cloud platform based on the second classification hidden matrix and the second hidden matrix; Receiving the first attention hidden normalization matrix sent by the cloud platform; performing matrix multiplication processing on the first attention hidden normalization matrix, the second hidden parameter set and the third hidden parameter set to obtain a first linear transformation hidden matrix; The first linear transformation hidden matrix is ​​sent to the cloud platform, so that the cloud platform reconstructs the first linear transformation hidden matrix and the second linear transformation hidden matrix to obtain a linear transformation hidden reconstruction matrix; and the linear transformation hidden reconstruction matrix is ​​input into a Gaussian error linear function to obtain a Gaussian error hidden matrix; the Gaussian error hidden matrix is ​​processed based on a secret sharing algorithm to obtain a first Gaussian error hidden matrix and a second Gaussian error hidden matrix; wherein the second linear transformation hidden matrix is ​​obtained by the cloud platform based on the second attention hidden normalization matrix; Receiving the first Gaussian error hidden matrix sent by the cloud platform; obtaining a first added hidden matrix based on the first Gaussian error hidden matrix and the first attention hidden normalization matrix; Sending the first additive hidden matrix to the cloud platform so that the cloud platform reconstructs the first additive hidden matrix and the second additive hidden matrix to obtain an additive reconstruction matrix; and inputting the additive reconstruction matrix into a normalization function to obtain an additive hidden normalized matrix; and processing the additive hidden normalized matrix based on a secret sharing algorithm to obtain a first additive hidden normalized matrix and a second additive hidden normalized matrix; wherein the second additive hidden matrix is ​​obtained by the cloud platform based on the second Gaussian error hidden matrix; Receiving the first additive hidden normalized matrix sent by the cloud platform; Perform parameter adaptation processing on the first added hidden normalized matrix to obtain the first inference result.

5. The security reasoning method of the model according to claim 4 is characterized in that: The performing parameter adaptation processing on the first added hidden normalized matrix to obtain the first inference result includes: Performing matrix permutation processing on the first permutation matrix and the linear parameter set of the parameter adaptation layer of the model to obtain a fourth hidden parameter set; performing matrix permutation processing on the second permutation matrix and the nonlinear parameter set of the parameter adaptation layer of the model to obtain a fifth hidden parameter set; Performing matrix multiplication processing on the first added hidden normalized matrix and the fourth hidden parameter set to obtain a first adapted hidden matrix; Sending the first adaptive hidden matrix, the fourth hidden parameter set and the fifth hidden parameter set to the cloud platform, so that the cloud platform reconstructs based on the first adaptive hidden matrix and the second adaptive hidden matrix to obtain an adaptive hidden reconstruction matrix, and inputs the adaptive hidden reconstruction matrix into a hyperbolic tangent function to obtain a hyperbolic hidden matrix; and processing the hyperbolic hidden matrix based on a secret sharing algorithm to obtain a first hyperbolic hidden matrix and a second hyperbolic hidden matrix; wherein the second adaptive hidden matrix is ​​obtained by the cloud platform based on the fourth hidden parameter set and the second additive hidden normalized matrix; The first hyperbolic hidden matrix sent by the cloud platform is received, and the first inference result is obtained based on the first hyperbolic hidden matrix and the fifth hidden parameter set.

6. A secure reasoning method for a model, characterized in that: Applied to a cloud platform, the method includes: Receive a first hidden parameter set, a second hidden parameter set, and a third hidden parameter set sent by a model developer; wherein the first hidden parameter set is obtained by the model developer based on the attention mechanism parameter set of the linear layer of the model and a randomly generated first permutation matrix; the second hidden parameter set is obtained by the model developer based on the embedding layer parameter set of the model and a randomly generated second permutation matrix; the third hidden parameter set is obtained by the model developer based on the linear layer parameter set of the model's pre-coronavirus neural network and a randomly generated third permutation matrix; receiving second hidden inference data sent by a user terminal; A first permutation matrix is ​​randomly generated based on a preset input sequence length; a second permutation matrix is ​​randomly generated based on the dimension of the model; and a third permutation matrix is ​​randomly generated based on the dimension of a linear layer of a feedforward neural network of the model; Performing matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set; performing matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set; performing matrix permutation processing on the first permutation matrix, the second permutation matrix, the third permutation matrix and the linear layer parameter set in the feedforward neural network to obtain a third hidden parameter set; receiving first hidden reasoning data sent by the user terminal, performing second privacy protection reasoning based on the first hidden reasoning data, the first hidden parameter set, the second hidden parameter set, and the third hidden parameter set to obtain a second reasoning result; The second reasoning result is sent to the user end, so that the user end performs recovery and reconstruction processing based on the first permutation matrix, the first reasoning result and the second reasoning result to obtain a target reasoning result; wherein the first reasoning result is obtained by the model developer end performing a first privacy protection reasoning based on the first hidden reasoning data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set.

7. A model security reasoning method, characterized in that: Applied to a security reasoning system, the security reasoning system includes a model developer end, a user end and a cloud platform; The method comprises: The model developer randomly generates a first permutation matrix based on a preset input sequence length; randomly generates a second permutation matrix based on the dimension of the model; and randomly generates a third permutation matrix based on the dimension of the linear layer of the feedforward neural network of the model; The model developer performs matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set; performs matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set; performs matrix permutation processing on the first permutation matrix, the second permutation matrix, the third permutation matrix and the linear layer parameter set in the feedforward neural network to obtain a third hidden parameter set; The model developer terminal sends the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to the cloud platform; The user end processes the original reasoning data based on the secret sharing algorithm to obtain first hidden reasoning data and second hidden reasoning data, and sends the first hidden reasoning data to the model developer end, and sends the second hidden reasoning data to the cloud platform; The model developer end performs a first privacy protection reasoning based on the first hidden reasoning data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to obtain a first reasoning result; and sends the first reasoning result to the user end; The cloud platform performs a second privacy protection reasoning based on the second hidden reasoning data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to obtain a second reasoning result; and sends the second reasoning result to the user end; The user end performs recovery and reconstruction processing based on the first permutation matrix, the first reasoning result, and the second reasoning result to obtain a target reasoning result.

8. A safety reasoning device for a model, characterized in that: Applied to a model developer end, the model developer end being provided with the model; The device comprises: A generation module is configured to randomly generate a first permutation matrix based on a preset input sequence length; randomly generate a second permutation matrix based on the dimension of the model; and randomly generate a third permutation matrix based on the dimension of a linear layer of a feedforward neural network of the model; A parameter processing module is configured to perform matrix permutation processing on the first permutation matrix and the attention mechanism parameter set of the linear layer of the model to obtain a first hidden parameter set; perform matrix permutation processing on the first permutation matrix and the embedding layer parameter set of the model to obtain a second hidden parameter set; perform matrix permutation processing on the second permutation matrix, the third permutation matrix and the linear layer parameter set in the feedforward neural network to obtain a third hidden parameter set; an inference module configured to receive first hidden inference data sent by a user terminal, perform first privacy protection inference based on the first hidden inference data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set, and obtain a first inference result; A first sending module is configured to send the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to the cloud platform, so that the cloud platform performs a second privacy protection reasoning based on the second hidden reasoning data, the first hidden parameter set, the second hidden parameter set and the third hidden parameter set to obtain a second reasoning result, and sends the second reasoning result to the user end; wherein the first hidden reasoning data and the second hidden reasoning data are obtained by the user end by processing the original reasoning data based on a secret sharing algorithm; The second sending module is configured to send the first permutation matrix and the first reasoning result to the user terminal, so that the user terminal performs recovery and reconstruction processing based on the first permutation matrix, the first reasoning result and the second reasoning result to obtain a target reasoning result.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the safety reasoning method of the model described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the safety reasoning method of the model described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Privacy reasoning method and system based on Transform network model, medium and electronic equipment

    CN117077162A

  • Privacy reasoning method, first node and second terminal

    CN119416892A