Data processing method, electronic equipment and computer readable storage medium
The lightweight data processing method using private multiplication and logistic regression techniques addresses data privacy and integrity issues in 5G+ industrial internet by encrypting data operations, ensuring secure and efficient data analysis and model training.
Patent Information
- Application Number
- CN202410057469.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-15
AI Technical Summary
In the 5G+ industrial Internet, data privacy leakage and model parameter leakage are serious problems. It is difficult for the existing technology to achieve lightweight data transmission and model training while ensuring data security, and the transmission data encryption cost is high.
Lightweight privacy point multiplication technology and logistic regression technology are adopted to generate random number data sets and basic data to perform operations, generate the first calculation result, and add it to the data request information, and encrypt and sign it using a shared key to ensure the privacy and integrity of data transmission.
It realizes data analysis and model training without leaking detailed data and model parameters, ensuring the privacy and integrity of the data. It is suitable for low-computing scenarios of 5G+ industrial Internet nodes, improving processing efficiency and maintaining model accuracy.
Smart Images

Figure CN120316807A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of terminal communication, and particularly to a data processing method, an electronic device, and a computer-readable storage medium. Background Art
[0002] 5G+(utilizing a new generation of information and communication technologies represented by 5G (the fifth generation of mobile communication technology)) industrial Internet, as an important part of the new industrial revolution, combines traditional industrial production with modern information technologies, realizing efficient connection and collaboration among devices, sensors, and systems. The rapid development of the industrial Internet has brought a large amount of data, which contains valuable information and potential value in the industrial production process and has important decision-making significance. However, the security issues in the use of industrial Internet data are becoming increasingly prominent, such as data privacy leakage, data abuse, etc., posing potential threats to the privacy of enterprises and individuals. With the continuous increase in data use and exchange, the security issues in data use have also attracted increasing attention. In the data use business process of 5G+ industrial Internet, in order to ensure the security and privacy of data, various technical means are needed to protect the confidentiality, integrity, and availability of data. Summary of the Invention
[0003] Embodiments of the present disclosure provide a data processing method, an electronic device, and a computer-readable storage medium.
[0004] In a first aspect, embodiments of the present disclosure provide a data processing method, which is applied to a data user, and the method includes:
[0005] Obtain basic data required for work;
[0006] Perform an operation based on the basic data to obtain a first operation result;
[0007] Add the first operation result to data request information;
[0008] Send the data request information to a data producer to obtain target data required for work from the data producer based on the first operation result.
[0009] In a second aspect, embodiments of the present disclosure provide a data processing method, which is applied to a data producer, and the method includes:
[0010] Receive data request information sent by a data user;
[0011] Perform an operation based on the target data to be returned and the data request information to obtain a second operation result;
[0012] Obtain data response information based on the second operation result;
[0013] Send the data response information to the data user.
[0014] In a third aspect, embodiments of the present disclosure provide an electronic device, which includes:
[0015] One or more processors;
[0016] A memory storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method;
[0017] One or more input / output (I / O) interfaces connected between the processor and the first memory, configured to implement information interaction between the processor and the memory.
[0018] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium storing a computer program thereon, and when the computer program is executed by a processor, the data processing method is implemented.
[0019] Before the data user in the embodiments of the present disclosure sends data request information to the data producer, operations are performed based on the basic data to obtain a first operation result; the first operation result is added to the data request information to obtain target data from the data producer based on the first operation result, so that the basic data is not directly provided to the data producer, and the data user can obtain the target data required for work from the data producer without disclosing the detailed basic data, ensuring the privacy of the basic data. Description of the Drawings
[0020] In the drawings of the embodiments of the present disclosure:
[0021] Figure 1 It is a flowchart of the data processing method applied to the data user provided by the embodiments of the present disclosure;
[0022] Figure 2 It is a schematic diagram of a data processing interaction among an administrator, a data user, and a data producer provided by the embodiments of the present disclosure;
[0023] Figure 3 It is another schematic diagram of a data processing interaction among an administrator, a data user, and a data producer provided by the embodiments of the present disclosure;
[0024] Figure 4 It is a flowchart of the data processing method applied to the data producer provided by the embodiments of the present disclosure;
[0025] Figure 5 It is a block diagram of the composition of the electronic device provided by the embodiments of the present disclosure;
[0026] Figure 6 Block diagram of the computer-readable storage medium provided by the embodiments of the present disclosure. Detailed implementation manners
[0027] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the communication perception data processing method and the computer-readable storage medium provided by the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0028] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, but the illustrated embodiments may be embodied in different forms and the present disclosure should not be construed as limited to the embodiments set forth hereinafter. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0029] The accompanying drawings of the embodiments of the present disclosure are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification, and are used to explain the present disclosure together with the detailed embodiments, and do not constitute a limitation to the present disclosure. By describing the detailed embodiments with reference to the accompanying drawings, the above and other features and advantages will become more apparent to those skilled in the art.
[0030] The present disclosure may be described with reference to the plan views and / or sectional views by means of the ideal schematic diagrams of the present disclosure. Therefore, the example illustrations may be modified according to the manufacturing techniques and / or tolerances.
[0031] In the case of no conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0032] The terms used in the present disclosure are only for describing specific embodiments and are not intended to limit the present disclosure. As used in the present disclosure, the term "and / or" includes any and all combinations of one or more related listed items. As used in the present disclosure, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. As used in the present disclosure, the terms "comprising", "made of...", specify the presence of the stated features, wholes, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their groups.
[0033] Unless otherwise defined, all terms (including technical and scientific terms) used in the present disclosure have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless the present disclosure clearly defines so.
[0034] The present disclosure is not limited to the embodiments shown in the drawings, but includes modifications to the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the drawings have schematic attributes, and the shapes of the regions shown in the figures illustrate the specific shapes of the regions of the components, but are not intended to be restrictive.
[0035] 5G+ (utilizing the new generation of information and communication technologies represented by 5G (the fifth generation of mobile communication technology)) industrial Internet, as an important part of the new generation of industrial revolution, combines traditional industrial production with modern information technologies, achieving efficient connection and collaboration among devices, sensors, and systems. The rapid development of the industrial Internet has brought a large amount of data, which contains valuable information and potential value in the industrial production process and has important decision-making significance. However, the issue of data usage security in the industrial Internet has become increasingly prominent, such as data privacy leakage, data abuse, etc., posing potential threats to the privacy of enterprises and individuals. With the continuous increase in data usage and exchange, the issue of data usage security has also received increasing attention. In the data usage business process of 5G+ industrial Internet, in order to ensure the security and privacy of data, various technical means need to be adopted to protect the confidentiality, integrity, and availability of data.
[0036] Due to concerns that the data stored in the cloud may be lost or leaked, data producers will store some sensitive industrial data in their own hard disks, forming data islands. At this time, data users must analyze and process the data without obtaining the industrial data, so as to ensure the confidentiality of sensitive industrial data. That is, the data is available but not visible.
[0037] At the same time, during the process of training the model and using the model for inference, data users do not want to disclose their model parameters. At this time, data users must complete the training and inference work without sending the model parameters to data producers. That is, the privacy of model parameters.
[0038] Finally, during the process of learning, training, and inference, data producers and data users must also ensure the confidentiality and integrity of the interaction data. Since a large amount of data needs to be transmitted during the training and inference processes, but the connection nodes in 5G+ industrial Internet may not have strong computing capabilities, the algorithms need to be as lightweight as possible. At the same time, we do not want to reduce the accuracy of the model and affect the performance while ensuring data security. Therefore, lightweight integrity and confidentiality protection mechanisms must be designed to reduce the computing overhead and data transmission overhead of data producers and data users.
[0039] The privacy of industrial data, the privacy of model parameters, the integrity and privacy of transmitted data are the current research hotspots in the security of data usage in 5G + industrial Internet, and also the basic security requirements. Existing data usage security technologies have problems such as leakage of industrial data and model parameters, and too high encryption costs for transmitted data.
[0040] The embodiments of the present disclosure design a lightweight 5G + industrial Internet data usage security solution, and the solution of this embodiment can ensure the privacy of industrial data and model parameters. Moreover, the solution of this embodiment also designs a lightweight integrity and privacy protection mechanism to ensure the security of the information being interacted, with high efficiency, and is suitable for the analysis and processing of a large amount of data.
[0041] Before the data user in the embodiments of the present disclosure sends data request information to the data producer, perform operations based on the basic data to obtain a first operation result; add the first operation result to the data request information, so as to obtain the target data from the data producer based on the first operation result, thereby preventing the basic data from being directly provided to the data producer, and the data user can obtain the target data required for work from the data producer without disclosing the detailed basic data, ensuring the privacy of the basic data.
[0042] The data processing method of the embodiments of the present disclosure can be executed by any electronic device that needs data transmission, such as a terminal device or a server. The terminal device may include, but is not limited to: in-vehicle devices, user equipment (UE), mobile devices, computing devices, wearable devices, etc. For example, it includes, but is not limited to, cellular phones, cordless phones, personal digital assistants (PDAs), portable computers, etc. The data processing method can be implemented by the processor calling computer-readable program instructions stored in the memory, or can be implemented by the server.
[0043] The solution of the embodiments of the present disclosure will be introduced in detail below.
[0044] The embodiments of the present disclosure provide a data processing method, as Figure 1 shown, applied to a data user, and the method may include steps S11 - S14:
[0045] S11. Obtain the basic data required for work.
[0046] S12. Perform operations according to the basic data to obtain a first operation result;
[0047] S13. Add the first operation result to the data request information;
[0048] S14. Send the data request information to the data producer to obtain the target data required for work from the data producer based on the first operation result.
[0049] In the embodiments of the present disclosure, in order to overcome the problem of insufficient security in the process of using 5G + industrial Internet data, the embodiments of the present disclosure propose a data usage security algorithm and protocol based on the private dot product technology and the logistic regression technology. The protocol includes the following entities in 5G + industrial Internet: administrators, data producers, and data users. The administrator is responsible for configuring the security key, the data producer produces industrial data, and the data user analyzes and processes the data, which may include but is not limited to analyzing and processing the data using a machine learning model.
[0050] In the embodiments of the present disclosure, the solutions of the embodiments of the present disclosure can be applied to any scenario where data needs to be transmitted among multiple parties and the data needs to be kept confidential from the other party. The data can be any data transmitted over the Internet, including but not limited to the data required for model work.
[0051] In the embodiments of the present disclosure, performing operations on the basic data to obtain the first operation result includes:
[0052] Generate a random number data set;
[0053] Perform operations on the basic data and the random numbers in the random number data set to obtain the first operation result.
[0054] In the embodiments of the present disclosure, performing operations on the basic data based on random numbers has a simple scheme and realizes lightweight operations, thereby improving the processing efficiency. It can be applied to the scenario of low computing power of 5G + industrial Internet nodes. And after the above operations, even if the processed basic data is sent to the data producer, the data producer will not know what the original basic data is, thus protecting the detailed content of the basic data from being leaked and ensuring the privacy of the data.
[0055] In the embodiments of the present disclosure, as Figure 2 、 Figure 3 shown, the solutions of the embodiments of the present disclosure will be described in detail below taking the data required for model work as an example.
[0056] In the embodiments of the present disclosure, in the scenario where the data user needs to analyze and process the data using a machine learning model, the basic data may include but is not limited to model data. For example, the model data may include but is not limited to the model parameters of the old model before model training, the error signals required during model inference, etc. The target data may include but is not limited to the sample data in the sample set required for model training, the feature vectors included in each sample, etc.
[0057] In the embodiments of the present disclosure, the data processing interaction solution of the embodiments of the present disclosure may include, but is not limited to, three stages: an initialization stage, a forward inference stage, and a backward parameter tuning stage.
[0058] In the embodiments of the present disclosure, in the initialization stage, the administrator pre-configures a shared key sk ∈ Z p (Z p is the set of secret keys) for protecting the integrity and privacy of the interaction data between the two. And, the data user initializes the model data required for the work, and the data producer prepares the dataset required for the model work.
[0059] In the embodiments of the present disclosure, initializing the model data may include: initializing the parameters θ = (θ1,..., θ n ) of the model (for example, a logistic regression model; logistic regression, also known as logistic regression analysis, is a generalized linear regression analysis model, commonly used in data mining, automatic disease diagnosis, economic prediction, etc. Logistic regression estimates the probability of an event occurring based on the given independent variable dataset. Since the result is a probability, the range of the dependent variable is between 0 and 1), where θ is the set of model parameters, θ1,..., θ n are the specific model parameters, and n is the dimension of the sample.
[0060] In the embodiments of the present disclosure, the dataset may include l samples X = {X1,..., X l}), where X is the set of samples, and each sample X i ∈ X contains n feature components X i = {X i1 ,..., X in}}, X i1 ,..., X in are n feature vectors. In addition, the label of each sample X i ∈ X is y i , and y i represents the classification to which the sample belongs.
[0061] In the embodiments of the present disclosure, in the forward inference stage and the backward parameter tuning stage, the data user may respectively send data request information to the data producer based on the model data it has to obtain the target data required for the model work.
[0062] In the embodiments of the present disclosure, before sending the data request information, the data producer may perform operations on the random numbers in the basic data and the random number dataset to obtain a first operation result; and add the first operation result to the data request information.
[0063] In the embodiments of the present disclosure, the basic data may include, but is not limited to, model data, and the data request information may include, but is not limited to: model inference data request information or model parameter tuning data request information; the random number data set includes, but is not limited to, general random numbers and specific random numbers.
[0064] In the embodiments of the present disclosure, performing operations on the random numbers in the basic data and the random number data set to obtain a first operation result may include:
[0065] For each model data, performing operations according to the general random number, the specific random number, the model data, and a first pre-designed formula respectively to obtain a plurality of first operation results.
[0066] In the embodiments of the present disclosure, in the forward inference stage and the backward parameter tuning stage, since the data request information and the model data are different, the detailed operation schemes for the model data are also different. The detailed operation schemes for the forward inference stage and the backward parameter tuning stage are introduced below respectively.
[0067] In the embodiments of the present disclosure, as Figure 2 shown, the forward inference stage includes three links: the inference process request link, the inference process response link, and the inference link. In the inference process request link, the data user generates model inference data request information, encrypts the model inference data request information and generates a digital signature, and sends the encrypted model inference data request information and the signature to the data producer; in the inference process response link, after receiving the model inference data request information, the data producer decrypts and verifies the integrity of the model inference data request information, injects the target data to be returned into the model inference data request information to obtain model inference data response information, and returns it to the data user; in the data user inference link, the data user parses the information required for the inference process (i.e., the target information) from the model inference data response information and executes the inference process to obtain an inference result. The processing schemes for the above-mentioned each link are introduced in detail below respectively.
[0068] In the embodiments of the present disclosure, in the forward inference stage, the data request information sent by the data producer is model inference data request information, and the model data includes the model parameters of the old model before model training; the number of the model parameters of the old model is the same as the total number of feature vectors included in each sample in the sample set required for model training, both being n.
[0069] In the embodiments of the present disclosure, as Figure 3 shown, the first pre-designed formula includes, but is not limited to:
[0070] θ ij ′=θ j +R ij R0;
[0071] Among them, θ ij ′ is the first operation result corresponding to the j-th feature vector included in the i-th sample in the sample set required for model training, and θ j is the j-th model parameter of the old model, and R ij is the specific random number corresponding to the j-th feature vector included in the i-th sample; R0 is a general random number; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of samples, and n is the total number of feature vectors included in each sample; each first operation result θ ij ′ corresponds to the specific random number R ij .
[0072] In the embodiments of the present disclosure, the data user randomly generates R = {R0, R ij , 0 < i < l + 1, 0 < j < n + 1}. Where R is a set of random numbers, and R0, R ij are the generated random numbers.
[0073] In the embodiments of the present disclosure, for each θ j ∈θ, the data user calculates θ ij ′ = θ j +R ij R0, where 0 < i < l + 1.
[0074] In the embodiments of the present disclosure, through the above calculations, the first operation result θ ij ′ in the forward inference stage can be obtained.
[0075] In the embodiments of the present disclosure, the first operation result is added to the data request information, including:
[0076] Adding each first operation result and the specific random number corresponding to each first operation result to the data request information.
[0077] In the embodiments of the present disclosure, as Figure 3 shown, ReqI = {(θ ij ′, R ij ), 0 < i < l + 1, 1 ≤ j ≤ n}, ReqI is the data request information, that is, the model inference data request information, and is composed of θ ij ′ (the first operation result) and R ij (the specific random number corresponding to the first operation result).
[0078] In the embodiments of the present disclosure, before sending the data request information to the data producer, the method may further include:
[0079] Encrypting the data request information and generating a digital signature.
[0080] In an embodiment of the present disclosure, encrypting data request information and generating a digital signature includes:
[0081] Encrypting the data request information using a pre-configured shared key and a preset encryption algorithm to obtain an encryption result; and,
[0082] Generating a digital signature for the data request information using the shared key and a preset signature function to obtain a signature result.
[0083] In an embodiment of the present disclosure, for model inference data request information, the encryption result may be represented by a first encryption result, and the signature result may be represented by a first signature result.
[0084] In an embodiment of the present disclosure, as Figure 3 shown, a data user uses sk ∈ Z p to encrypt ReqI = {(θ ij ′, R ij ), 0 < i < l + 1, 1 ≤ j ≤ n} to obtain C ReqI = Enc sk (ReqI), where C ReqI is the ciphertext of ReqI, Enc is a symmetric encryption algorithm (AES), and Enc sk (ReqI) means encrypting ReqI using a symmetric encryption algorithm based on the shared key sk ∈ Z p .
[0085] In an embodiment of the present disclosure, the preset encryption algorithm may include, but is not limited to, a symmetric encryption algorithm (the symmetric encryption algorithm uses the encryption method of a single-key cryptosystem. The same key can be used for both information encryption and information decryption. This encryption method is called symmetric encryption, also known as single-key encryption). Any implementable encryption algorithm is applicable to the embodiments of the present disclosure.
[0086] In an embodiment of the present disclosure, as Figure 3 shown, a data user uses sk ∈ Z p to generate a digital signature σ ReqI = h(ReqI, sk) for ReqI, where h(.) is a hash function and σ ReqI is the digital signature of ReqI. A digital signature (also known as a public key digital signature) is a digital string that can only be generated by the sender of the information and cannot be forged by others. This digital string is also a valid proof of the authenticity of the information sent by the sender of the information. It is a kind of physical signature similar to an ordinary signature written on paper and is implemented using techniques in the field of public key encryption for authenticating the integrity of digital information. A hash function refers to a function that maps the key value of an element in a hash table to the storage location of the element.
[0087] In an embodiment of the present disclosure, sending data request information to a data producer includes:
[0088] Sending an encryption result and a signature result to the data producer.
[0089] In an embodiment of the present disclosure, that is, sending a first encryption result and a first signature result to the data producer.
[0090] In an embodiment of the present disclosure, as Figure 3 shown, a data user sends (C ReqI , σ ReqI ) to the data producer.
[0091] In an embodiment of the present disclosure, the data producer receives the data request information sent by the data user, performs an operation according to the target data to be returned and the data request information to obtain a second operation result, obtains data response information according to the second operation result, and sends the data response information to the data user.
[0092] In an embodiment of the present disclosure, after receiving the data request information sent by the data user, the data producer decrypts the data request information and verifies the integrity of the data request information.
[0093] In an embodiment of the present disclosure, the data producer can decrypt the data request information (here it is model inference data request information) based on a shared key preconfigured by the administrator, and use the shared key and a preset signature function to perform a digital signature on the decrypted data request information, detect whether the obtained signature result is the same as the first signature result, and verify the integrity of the data request information according to the detection result. Among them, if the obtained signature result is the same as the first signature result, it indicates that the data request information is complete; if the obtained signature result is different from the first signature result, it indicates that the data request information is incomplete.
[0094] In an embodiment of the present disclosure, after the data producer receives the ciphertext C ReqI , it decrypts to obtain ReqI = Dec sk (C ReqI ), and then verifies whether it satisfies σ ReqI = h(ReqI, sk) to confirm the integrity of the data.
[0095] In an embodiment of the present disclosure, based on the foregoing, it can be known that the data request information received by the data producer includes multiple first operation results calculated by the data user according to model data and a specific random number corresponding to each first operation result.
[0096] In an embodiment of the present disclosure, performing an operation according to the target data to be returned and the data request information to obtain a second operation result includes:
[0097] Perform an operation based on the target data, the first operation result corresponding to the target data, and the second pre-designed formula to obtain a first sub-operation result;
[0098] Perform an operation based on the target data, the specific random number corresponding to the first operation result corresponding to the target data, and the third pre-designed formula to obtain a second sub-operation result;
[0099] Use the first sub-operation result and the second sub-operation result as the second operation result.
[0100] In the embodiments of the present disclosure, obtaining data response information based on the second operation result includes:
[0101] Use the first sub-operation result as the first data response information and the second sub-operation result as the second data response information;
[0102] The data response information is composed of the first data response information and the second data response information.
[0103] In the embodiments of the present disclosure, when the data request information is model inference data request information, the data response information returned by the data producer is model inference data response information, and the target data that the data producer needs to return includes the sample data in the sample set required for model training.
[0104] In the embodiments of the present disclosure, as Figure 3 shown, the second pre-designed formula includes:
[0105]
[0106] where, InferResp i1 is the first sub-operation result corresponding to the i-th sample, θ ij ′ is the first operation result corresponding to the j-th feature vector included in the i-th sample, X ij is the j-th feature vector of the i-th sample; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of samples, and n is the total number of feature vectors included in each sample.
[0107] In the embodiments of the present disclosure, as Figure 3 shown, the third pre-designed formula includes:
[0108]
[0109] where, InferResp i2 is the second sub-operation result corresponding to the i-th sample, and R ij is the specific random number corresponding to the first operation result corresponding to the j-th feature vector included in the i-th sample.
[0110] In the embodiments of the present disclosure, for each sample X i ∈X, the data producer calculates and to obtain InferResp = {InferResp i1 , InferResp i2 , 0 < i < l + 1}, where InferResp is data response information (substantially a set of response information), InferResp i1 and InferResp i2 are the first data response information (i.e., the first sub - operation result) and the second data response information (i.e., the second sub - operation result), respectively.
[0111] In the embodiments of the present disclosure, after the data producer obtains the data response information according to the second operation result, the method further includes: encrypting the data response information and generating a digital signature to send the encrypted and signed data response information to the data user.
[0112] In the embodiments of the present disclosure, as Figure 3 shown, the data producer can use the shared key sk ∈ Z p to encrypt the above - mentioned data response information InferResp = {InferResp i1 , InferResp i2 , 0 < i < l + 1} to obtain C InferResp = Enc sk (InferResp), where Enc is a symmetric encryption algorithm (such as AES), and C InferResp is the ciphertext of InferResp.
[0113] In the embodiments of the present disclosure, as Figure 3 shown, the data producer uses sk ∈ Z p to generate a digital signature σ InferResp = h(InferResp, sk) for InferResp, where h(.) is a hash function, and σ InferResp is the digital signature of InferResp.
[0114] In the embodiments of the present disclosure, the data producer sends (C InferResp , σ InferResp ) as the model inference data response information to the data user.
[0115] In the embodiments of the present disclosure, after the data user sends the data request information to the data producer, the method further includes:
[0116] Receiving the data response information of the data request information returned by the data producer;
[0117] Obtain target data based on the data response information.
[0118] In the embodiments of the present disclosure, the target data may include, but is not limited to: data required for model operation;
[0119] The data response information may include, but is not limited to: model inference data response information or model parameter tuning data response information.
[0120] In the embodiments of the present disclosure, after the data user receives the model inference data response information returned by the data producer, the data required for model operation can be obtained based on the model inference data response information.
[0121] In the embodiments of the present disclosure, when the data response information is model inference data response information, the data required for model operation includes the error between the inference output result corresponding to each sample in the sample set required for model inference and the label corresponding to the sample; wherein, each sample corresponds to a label for indicating the classification to which the sample belongs; the data response information includes: first data response information and second data response information.
[0122] In the embodiments of the present disclosure, obtaining target data based on the data response information includes: performing the following operations for each sample respectively:
[0123] Subtract the product of the second data response information corresponding to the sample and a pre-generated universal random number from the first data response information corresponding to the sample to obtain a first difference;
[0124] Calculate the inference output result corresponding to the sample based on the first difference;
[0125] Subtract the label corresponding to the sample from the inference output result to obtain the error corresponding to the sample.
[0126] In the embodiments of the present disclosure, the data user decrypts the model inference data response information to obtain InferResp = Enc sk (C InferResp ), and verifies whether it satisfies σ InferResp = h(InferResp, sk) to confirm the integrity of InferResp.
[0127] In the embodiments of the present disclosure, as Figure 3 shown, after the data user confirms the integrity of InferResp, for each sample X i ∈ X, calculate g(X i ) = InferResp i1 -R0InferResp i2 , then g(X i ) = θ T·X i , where g(X i ) is the result of the inner product of the model parameters and the sample. Data users can calculate Get sample X i The inference output result h θ (X i ); for each sample X i ∈X, the data user calculates τ i =h θ (X i )-y i , where τ i is the error between the inference result and the label (or sample error), and finally we get τ={τ i ,0 <i<l+1},其中τ为所有样本误差的集合。
[0128] In the disclosed embodiment, the above scheme can realize data interaction between data users and data producers in the forward reasoning stage. Through the above scheme, data users can realize the training process of the model without leaking model parameters and use the model to reason about data.
[0129] In the embodiments of the present disclosure, Figure 2 As shown in the figure, the backward parameter adjustment phase includes three links: parameter adjustment request link, parameter adjustment response link and parameter adjustment link. In the parameter adjustment request link, the data user generates model parameter adjustment data request information, encrypts the model parameter adjustment data request information and generates a digital signature, and sends the encrypted model parameter adjustment data request information and signature to the data producer; in the parameter adjustment response link, after receiving the model parameter adjustment data request information, the data producer decrypts and verifies the integrity of the model parameter adjustment data request information, injects the target data to be returned into the model parameter adjustment data request information, and obtains the model parameter adjustment data response information; in the parameter adjustment link of the data user, the data user parses the information required for the parameter adjustment process (i.e., the target information) from the model parameter adjustment data response information, and executes the parameter adjustment process to obtain the parameter adjustment result. The following is a detailed introduction to the processing schemes of each of the above links.
[0130] In the embodiment of the present disclosure, in the backward parameter adjustment stage, the data request information sent by the data user is model parameter adjustment data request information, and the model data includes but is not limited to error data; the number of error data is the same as the total number of samples.
[0131] In the embodiment of the present disclosure, at this stage, before sending the model parameter adjustment data request information, the data producer also operates on the basic data and the random numbers in the random number data set to obtain a first operation result, which may specifically include:
[0132] For each model data, operations are respectively performed according to a general random number, a specific random number, the model data, and a first pre-designed calculation formula to obtain a plurality of first operation results.
[0133] In an embodiment of the present disclosure, as Figure 3 shown, the first pre-designed calculation formula may include, but is not limited to:
[0134] τ ij ′ = τ i +S ij S0;
[0135] Wherein, τ ij ′ is the first operation result corresponding to the j-th feature vector included in the i-th sample, τ i is the i-th error data, S ij is the specific random number corresponding to the j-th feature vector included in the i-th sample in the sample set required for model training, S0 is the general random number; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of samples, and n is the total number of feature vectors included in each sample; each first operation result τ ij ′ corresponds to the specific random number S ij .
[0136] In an embodiment of the present disclosure, a data user randomly generates a set of data S = {S0, S ij , 0 < i < l + 1, 1 ≤ j ≤ n}. Wherein S is a set of random numbers, S0, S ij are the generated random numbers.
[0137] In an embodiment of the present disclosure, for each τ i , the data user calculates τ ij ′ = τ i +S ij S0, where 0 < i < l + 1.
[0138] In an embodiment of the present disclosure, through the above calculations, the first operation result τ ij ′ in the backward parameter adjustment stage can be obtained.
[0139] In an embodiment of the present disclosure, adding the first operation result to the data request information includes:
[0140] Adding each first operation result and the specific random number corresponding to each first operation result to the data request information.
[0141] In an embodiment of the present disclosure, as Figure 3 shown, ReqP = {(τ ij ′, S ij), 0 < i < l + 1, 1 ≤ j ≤ n}, ReqP is the data request information, that is, the model parameter tuning data request information, which is composed of τ ij ′ (the first operation result) and S ij (the specific random number corresponding to the first operation result).
[0142] In the embodiment of the present disclosure, before sending the data request information to the data producer, the method may further include:
[0143] Encrypt the data request information and generate a digital signature.
[0144] In the embodiment of the present disclosure, encrypting the data request information and generating a digital signature includes:
[0145] Using a pre-configured shared key and a preset encryption algorithm to encrypt the data request information to obtain an encryption result; and,
[0146] Using the shared key and a preset signature function to perform a digital signature on the data request information to obtain a signature result.
[0147] In the embodiment of the present disclosure, for the model parameter tuning data request information, the encryption result can be represented by a second encryption result, and the signature result can be represented by a second signature result.
[0148] In the embodiment of the present disclosure, as Figure 3 shown, the data user uses sk ∈ Z p to encrypt ReqP = {(τ ij ′, S ij ), 0 < i < l + 1, 1 ≤ j ≤ n} to obtain C ReqP = Enc sk (ReqP), where C ReqP is the ciphertext of ReqP, Enc is the symmetric encryption algorithm (AES), and Enc sk (ReqP) means encrypting ReqP using the symmetric encryption algorithm based on the shared key sk ∈ Z p .
[0149] In the embodiment of the present disclosure, the preset encryption algorithm may include but is not limited to the symmetric encryption algorithm, and any implementable encryption algorithm is applicable to the embodiment of the present disclosure.
[0150] In the embodiment of the present disclosure, the data user uses sk ∈ Z p to generate a digital signature σ ReqP = h(ReqP, sk) for ReqP, where h(.) is the hash function and σ ReqP is the digital signature of ReqP.
[0151] In the embodiments of the present disclosure, sending data request information to a data producer includes:
[0152] Sending an encryption result and a signature result to the data producer.
[0153] In the embodiments of the present disclosure, that is, sending a second encryption result and a second signature result to the data producer.
[0154] In the embodiments of the present disclosure, a data user sends (C ReqP , σ ReqP ) to the data producer.
[0155] In the embodiments of the present disclosure, the data producer receives the data request information sent by the data user, performs an operation according to the target data to be returned and the data request information to obtain a second operation result, obtains data response information according to the second operation result, and sends the data response information to the data user.
[0156] In the embodiments of the present disclosure, after receiving the data request information sent by the data user, the data producer decrypts the data request information and verifies the integrity of the data request information.
[0157] In the embodiments of the present disclosure, the data producer can decrypt the data request information (here it is the model tuning parameter data request information) based on the shared key pre-configured by the administrator, and use the shared key and a preset signature function to perform a digital signature on the decrypted data request information, detect whether the obtained signature result is the same as the second signature result, and verify the integrity of the data request information according to the detection result. Among them, if the obtained signature result is the same as the second signature result, it indicates that the data request information is complete; if the obtained signature result is different from the second signature result, it indicates that the data request information is incomplete.
[0158] In the embodiments of the present disclosure, after the data producer receives the ciphertext C ReqP , it decrypts to obtain ReqP = Dec sk (C ReqP ), and then verifies whether σ ReqP = h(ReqP, sk) to confirm the integrity of the data.
[0159] In the embodiments of the present disclosure, based on the foregoing, it can be known that the data request information received by the data producer includes multiple first operation results calculated by the data user according to the model data and a specific random number corresponding to each first operation result.
[0160] In the embodiments of the present disclosure, performing an operation according to the target data to be returned and the data request information to obtain a second operation result includes:
[0161] Perform an operation based on the target data, the first operation result corresponding to the target data, and the second pre-designed formula to obtain a first sub-operation result;
[0162] Perform an operation based on the target data, the specific random number corresponding to the first operation result corresponding to the target data, and the third pre-designed formula to obtain a second sub-operation result;
[0163] Use the first sub-operation result and the second sub-operation result as the second operation result.
[0164] In the embodiments of the present disclosure, obtaining data response information based on the second operation result includes:
[0165] Use the first sub-operation result as the first data response information and the second sub-operation result as the second data response information;
[0166] The data response information is composed of the first data response information and the second data response information.
[0167] In the embodiments of the present disclosure, when the data request information is model tuning data request information, the data response information returned by the data producer is model tuning data response information, and the target data that the data producer needs to return includes, but is not limited to, the feature vectors included in the samples in the sample set required for model training.
[0168] In the embodiments of the present disclosure, as Figure 3 shown, the second pre-designed formula includes, but is not limited to:
[0169]
[0170] where ParaResp 1j is the first sub-operation result corresponding to the jth feature vector included in the ith sample, τ ij ′ is the first operation result corresponding to the jth feature vector included in the ith sample, X ij is the jth feature vector of the ith sample; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of samples, and n is the total number of feature vectors included in each sample.
[0171] In the embodiments of the present disclosure, as Figure 3 shown, the third pre-designed formula includes, but is not limited to:
[0172]
[0173] where ParaResp 2j is the second sub-operation result corresponding to the jth feature vector included in the ith sample, S ijis a specific random number corresponding to the first operation result corresponding to the j-th feature vector included in the i-th sample.
[0174] In the embodiments of the present disclosure, for each feature component j, the data producer calculates and to obtain ParaResp = {ParaResp 1j , ParaResp 2j , 1 ≤ j ≤ n}. Wherein, ParaResp is data response information (substantially a set of response information), ParaResp 1j and ParaResp 2j are the first data response information (i.e., the first sub-operation result) and the second data response information (i.e., the second sub-operation result), respectively.
[0175] In the embodiments of the present disclosure, after the data producer obtains the data response information according to the second operation result, the method further includes: encrypting the data response information and generating a digital signature to send the encrypted and signed data response information to the data user.
[0176] In the embodiments of the present disclosure, as Figure 3 shown, the data producer can use the shared key sk ∈ Z p to encrypt the above data response information ParaResp = {ParaResp 1j , ParaResp 2j , 1 ≤ j ≤ n} to obtain C ParaResp = Enc sk (ParaResp), where Enc is a symmetric encryption algorithm (such as AES), and C ParaResp is the ciphertext of ParaResp.
[0177] In the embodiments of the present disclosure, the data producer uses sk ∈ Z p to generate a digital signature σ ParaResp = h(ParaResp, sk) for ParaResp, where h(.) is a hash function, and σ ParaResp is the digital signature of ParaResp.
[0178] In the embodiments of the present disclosure, as Figure 3 shown, the data producer sends (C ParaResp , σ ParaResp ) to the data user by adding them to the model inference data response information.
[0179] In the embodiments of the present disclosure, after the data user sends the data request information to the data producer, the method further includes:
[0180] Receive the data response information for the data request information returned by the data producer;
[0181] Obtain the target data based on the data response information.
[0182] In the embodiments of the present disclosure, the target data may include, but is not limited to: data required for model operation;
[0183] The data response information may include, but is not limited to: model inference data response information or model parameter tuning data response information.
[0184] In the embodiments of the present disclosure, after the data user receives the model parameter tuning data response information returned by the data producer, the data user may obtain the data required for model operation based on the model parameter tuning data response information.
[0185] In the embodiments of the present disclosure, when the data response information is model parameter tuning data response information, the data required for model operation includes the model parameters of the new model obtained after model training; the data response information includes: first data response information and second data response information.
[0186] In the embodiments of the present disclosure, obtaining the target data based on the data response information includes: performing the following operations on each feature vector included in each sample in the sample set required for model training respectively:
[0187] Subtract the product of the second data response information corresponding to the feature vector and a pre-generated general random number from the first data response information corresponding to the feature vector to obtain a second difference;
[0188] Calculate each model parameter of the new model according to the second difference and each model parameter of the old model.
[0189] In the embodiments of the present disclosure, the data user decrypts the model parameter tuning data response information to obtain ParaResp = Enc sk (C ParaResp ), and verifies whether it satisfies σ ParaResp = h(ParaResp, sk) to confirm the integrity of ParaResp.
[0190] In the embodiments of the present disclosure, as Figure 3 shown, after the data user confirms the integrity of ParaResp, for each feature component j, the data user calculates f(θ j ) = ParaResp 1j - S0InferResp 2j , then f(θ j ) = τ T ·J, where J = <X 1j ,..., X lj>, is an l-dimensional new vector composed of the j-th feature vectors of all samples.
[0191] In the embodiments of the present disclosure, for each θ j ∈θ, calculate θ jnew = θ j -αf(θ j ). Where θ jnew is the model parameter of the new model, θ j is the model parameter of the old model, and α is the learning rate.
[0192] In the embodiments of the present disclosure, through the above solution, data interaction between the data user and the data producer in the backward parameter tuning stage can be achieved. Through the above solution, the data user can use industrial data to train the logistic regression model without knowing the sample data set, realizing that the data is available but invisible.
[0193] The embodiments of the present disclosure also provide a data processing method, which is applied to a data producer. As Figure 4 shown, the method may include steps S21-S24:
[0194] S21. Receive the data request information sent by the data user.
[0195] S22. Perform an operation according to the target data to be returned and the data request information to obtain a second operation result.
[0196] S23. Obtain data response information according to the second operation result.
[0197] S24. Send the data response information to the data user.
[0198] In the embodiments of the present disclosure, after receiving the data request information sent by the data user, the method may further include: decrypting the data request information and verifying the integrity of the data request information.
[0199] In the embodiments of the present disclosure, after obtaining the data response information according to the second operation result, the method may further include: encrypting the data response information and generating a digital signature to send the encrypted and signed data response information to the data user.
[0200] In the embodiments of the present disclosure, the data request information includes multiple first operation results calculated by the data user according to the model data and a specific random number corresponding to each first operation result;
[0201] Performing an operation according to the target data to be returned and the data request information to obtain a second operation result includes:
[0202] Perform an operation based on the target data, the first operation result corresponding to the target data, and the second pre-designed formula to obtain the first sub-operation result;
[0203] Perform an operation based on the target data, the specific random number corresponding to the first operation result corresponding to the target data, and the third pre-designed formula to obtain the second sub-operation result;
[0204] Use the first sub-operation result and the second sub-operation result as the second operation result.
[0205] In the embodiments of the present disclosure, when the data response information is model inference data response information, the target data includes the sample data in the sample set required for model training;
[0206] The second pre-designed formula may include but is not limited to:
[0207]
[0208] Among them, InferResp i1 is the first sub-operation result corresponding to the i-th sample, θ ij ′ is the first operation result corresponding to the j-th feature vector included in the i-th sample, X ij is the j-th feature vector of the i-th sample; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of samples, and n is the total number of feature vectors included in each sample.
[0209] In the embodiments of the present disclosure, when the data response information is model inference data response information, the target data includes the sample data in the sample set required for model training;
[0210] The third pre-designed formula may include but is not limited to:
[0211]
[0212] Among them, InferResp i2 is the second sub-operation result corresponding to the i-th sample, R ij is the specific random number corresponding to the first operation result corresponding to the j-th feature vector included in the i-th sample.
[0213] In the embodiments of the present disclosure, when the data response information is model tuning data response information, the target data includes the feature vectors included in the samples in the sample set required for model training;
[0214] The second pre-designed formula may include but is not limited to:
[0215]
[0216] Among them, ParaResp 1j is the first sub - operation result corresponding to the j - th feature vector included in the i - th sample, and τ ij ′ is the first operation result corresponding to the j - th feature vector included in the i - th sample.
[0217] In the embodiments of the present disclosure, when the data response information is model inference data response information, the target data includes the feature vectors included in the samples in the sample set required for model training;
[0218] The third pre - designed calculation formula may include but is not limited to:
[0219]
[0220] Among them, ParaResp 2j is the second sub - operation result corresponding to the j - th feature vector included in the i - th sample, S ij is the specific random number corresponding to the first operation result corresponding to the j - th feature vector included in the i - th sample, X ij is the j - th feature vector of the i - th sample; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of samples, and n is the total number of feature vectors included in each sample.
[0221] In the embodiments of the present disclosure, obtaining data response information according to the second operation result includes:
[0222] Taking the first sub - operation result as the first data response information and taking the second sub - operation result as the second data response information;
[0223] The data response information is composed of the first data response information and the second data response information.
[0224] In the embodiments of the present disclosure, any of the above - mentioned data processing methods applied to data users is applicable to the data processing method of this data producer, and will not be elaborated here one by one.
[0225] In the embodiments of the present disclosure, it has at least the following advantages:
[0226] 1. Without knowing the sample data set, the data user can use industrial data to train the logistic regression model, realizing that the data is available but invisible.
[0227] 2. Without leaking the model parameters, the data user can realize the training process of the model and use the model to infer data.
[0228] 3. The data transmitted between the data user and the data producer is encrypted, ensuring the privacy of the transmitted data.
[0229] 4. The data transmitted between the data user and the data producer is signed, ensuring the integrity of the transmitted data.
[0230] 5. The solution of the present disclosure embodiment only uses lightweight symmetric encryption algorithms, hash functions, addition, multiplication operations, etc., and does not use high-complexity cryptographic algorithms. Therefore, it has high efficiency and is suitable for the low computing power scenario of 5G + industrial Internet nodes.
[0231] 6. The private dot product technology used in this solution has almost no impact on the accuracy of the machine learning model without overflow, has higher accuracy than protection technologies such as differential privacy, and is more suitable for precise processing under 5G + industrial Internet.
[0232] The present disclosure embodiment also provides an electronic device 100, as Figure 5 shown, the electronic device 100 includes:
[0233] One or more processors 101;
[0234] A memory 102, on which one or more programs are stored. When the one or more programs are executed by the one or more processors 101, the one or more processors 101 implement a data processing method applied to data users;
[0235] One or more input / output I / O interfaces 103, connected between the processor 101 and the memory 102, configured to implement information interaction between the processor 101 and the memory 102.
[0236] Among them, the processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the first memory 102 is a device with data storage capabilities, including but not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the I / O interface (read / write interface) 103 is connected between the processor 101 and the memory 102 and can implement information interaction between the processor 101 and the memory 102, including but not limited to a data bus (Bus), etc.
[0237] In some embodiments, the processor 101, the memory 102, and the I / O interface 103 are interconnected through a bus 104 and are further connected to other components of the computing device.
[0238] The present disclosure embodiment also provides a computer-readable storage medium 200, as Figure 6As shown, a computer program is stored on the computer-readable storage medium 200, and when the computer program is executed by a processor, it implements a data processing method applied to data users.
[0239] Those of ordinary skill in the art can understand that all or some of the functional modules / units disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations.
[0240] In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation.
[0241] Some or all physical components can be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH), or other magnetic disk memories; compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical disc memories; magnetic cartridges, tapes, magnetic disk storage, or other magnetic memories; and any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0242] The present disclosure has disclosed exemplary embodiments, and although specific terms are used, they are only used and should only be construed as having a general illustrative meaning and not for the purpose of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, the features, characteristics, and / or elements described in connection with a particular embodiment can be used alone or in combination with the features, characteristics, and / or elements described in connection with other embodiments. Accordingly, those skilled in the art will understand that various forms and details can be changed without departing from the scope of the present disclosure as set forth by the appended claims.
Claims
1. A data processing method, characterized in that, Applied to a data user, the method includes: Obtain the basic data required for work; Perform operations based on the basic data to obtain a first operation result; Add the first operation result to the data request information; Send the data request information to the data producer to obtain the target data required for work from the data producer based on the first operation result.
2. The data processing method according to claim 1, wherein The performing operations based on the basic data to obtain a first operation result includes: Generate a random number data set; Perform operations on the basic data and the random numbers in the random number data set to obtain the first operation result.
3. The data processing method according to claim 2, characterized in that The basic data includes model data, The data request information includes: model inference data request information or model tuning data request information; The random number data set contains general random numbers and specific random numbers; The performing operations on the basic data and the random numbers in the random number data set to obtain the first operation result includes: For each model data, perform operations respectively according to the general random number, the specific random number, the model data and a first pre-designed formula to obtain a plurality of the first operation results.
4. The data processing method according to claim 3, wherein In the case where the data request information is model inference data request information, the model data includes the model parameters of the old model before model training; the number of the model parameters of the old model is the same as the total number of feature vectors included in each sample in the sample set required for model training; The first pre-designed formula includes: θ ij ' = θ j + R ij R0; where, θ ij ′ is the first operation result corresponding to the j-th feature vector included in the i-th sample in the sample set required for model training, θ j is the j-th model parameter of the old model, R ij is the specific random number corresponding to the j-th feature vector included in the i-th sample, R0 is the general random number; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of the samples, n is the total number of the feature vectors included in each sample; each of the first operation results θ ij ′ corresponds to the specific random number R ij correspondingly.
5. The data processing method according to claim 3, wherein In the case where the data request information is model tuning data request information, the model data includes error data; the number of the error data is the same as the total number of the samples; The first pre-designed formula includes: τ ij ′ = τ i + S ij S0; Among them, τ ij ′ is the first operation result corresponding to the j-th feature vector included in the i-th sample, and τ i is the i-th error data, and S ij is the specific random number corresponding to the j-th feature vector included in the i-th sample in the sample set required for model training, and S0 is the general random number; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of the samples, and n is the total number of the feature vectors included in each sample; each of the first operation results τ ij ′ corresponds to the specific random number S ij correspondingly.
6. The data processing method according to any one of claims 3-5, characterized in that, The adding the first operation result to the data request information includes: Add each of the first operation results and the specific random number corresponding to each of the first operation results to the data request information.
7. The data processing method according to claim 1, wherein After sending the data request information to the data producer, the method further includes: Receive the data response information of the data producer to the data request information; Obtain the target data based on the data response information.
8. The data processing method according to claim 7, wherein, The target data includes: data required for model work; The data response information includes: model inference data response information or model tuning data response information.
9. The data processing method according to claim 8, characterized in that, In the case where the data response information is model inference data response information, the data required for model work includes the error between the inference output result corresponding to each sample in the sample set required for model inference and the label corresponding to the sample; wherein, each sample corresponds to a label for indicating the classification to which the sample belongs; the data response information includes: first data response information and second data response information; The obtaining the target data based on the data response information includes: performing the following operations respectively for each sample: Subtract the product of the second data response information corresponding to the sample and a pre-generated general random number from the first data response information corresponding to the sample to obtain a first difference; Calculate the inference output result corresponding to the sample based on the first difference; Subtract the label corresponding to the sample from the inference output result to obtain the error corresponding to the sample.
10. The data processing method according to claim 8, wherein When the data response information is model parameter adjustment data response information, the data required for the model to work includes the model parameters of the new model obtained after model training; The data response information includes: first data response information and second data response information; Obtaining the target data based on the data response information includes: performing the following operations for each feature vector included in each sample in the sample set required for model training: Subtract the product of the second data response information corresponding to the feature vector and a pre-generated general random number from the first data response information corresponding to the feature vector to obtain a second difference; Calculate each model parameter of the new model according to the second difference and each model parameter of the old model.
11. A data processing method, characterized in that, Applied to a data producer, the method includes: Receive data request information sent by a data user; Perform an operation according to the target data to be returned and the data request information to obtain a second operation result; Obtain data response information according to the second operation result; Send the data response information to the data user.
12. The data processing method according to claim 11, wherein The data request information includes multiple first operation results calculated by the data user based on model data and a specific random number corresponding to each first operation result; The performing an operation according to the target data to be returned and the data request information to obtain a second operation result includes: Perform an operation according to the target data, the first operation result corresponding to the target data, and a second pre-designed formula to obtain a first sub-operation result; Perform an operation according to the target data, the specific random number corresponding to the first operation result corresponding to the target data, and a third pre-designed formula to obtain a second sub-operation result; Use the first sub-operation result and the second sub-operation result as the second operation result.
13. The data processing method according to claim 12, wherein When the data response information is model inference data response information, the target data includes sample data in the sample set required for model training; The second pre-designed formula includes: Among them, InferResp i1 is the first sub-operation result corresponding to the i-th sample, and θ ij ′ is the first operation result corresponding to the j-th feature vector included in the i-th sample, and X ij is the j-th feature vector of the i-th sample; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of the samples, and n is the total number of the feature vectors included in each sample.
14. The data processing method according to claim 12, characterized in that, When the data response information is model inference data response information, the target data includes sample data in the sample set required for model training; The third pre-designed formula includes: Among them, InferResp i2 is the second sub-operation result corresponding to the i-th sample, R ij is the specific random number corresponding to the first operation result corresponding to the j-th feature vector included in the i-th sample, X ij is the j-th feature vector of the i-th sample; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of the samples, and n is the total number of the feature vectors included in each sample.
15. The data processing method according to claim 12, wherein When the data response information is model parameter adjustment data response information, the target data includes feature vectors included in samples in the sample set required for model training; The second pre-designed formula includes: Among them, ParaResp 1j is the first sub-operation result corresponding to the j-th feature vector included in the i-th sample, τ ij ′ is the first operation result corresponding to the j-th feature vector included in the i-th sample, X ij is the j-th feature vector of the i-th sample; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of the samples, and n is the total number of the feature vectors included in each sample.
16. The data processing method according to claim 12, wherein When the data response information is model inference data response information, the target data includes feature vectors included in samples in the sample set required for model training; The third pre-designed formula includes: Among them, ParaResp 2j is the second sub-operation result corresponding to the j-th feature vector included in the i-th sample, S ij is the specific random number corresponding to the first operation result corresponding to the j-th feature vector included in the i-th sample, X ij is the j-th feature vector of the i-th sample; i and j are positive integers, 0 < i < l + 1, 0 < j < n + 1, l is the total number of the samples, and n is the total number of the feature vectors included in each sample.
17. The data processing method according to claim 12, wherein The obtaining data response information according to the second operation result includes: Use the first sub-operation result as the first data response information and the second sub-operation result as the second data response information; The data response information is composed of the first data response information and the second data response information.
18. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory storing one or more programs which, when executed by the one or more processors, cause the one or more processors to implement the data processing method according to any one of claims 1-10, or claims 11-17; One or more input / output (I / O) interfaces connected between the processor and the memory and configured to enable information interaction between the processor and the memory.
19. A computer-readable storage medium storing a computer program which, when executed by a processor, implements the data processing method according to any one of claims 1-10, or claims 11-17.