Longitudinal federal GRU model training method and system based on secret sharing and differential privacy

By adopting secret sharing and differential privacy methods in vertical federated learning, the sample data is encrypted, split and noise processed, which solves the problems of low training efficiency and privacy leakage of the GRU model and achieves efficient data protection and model updating.

CN120675707AActive Publication Date: 2025-09-19杭州金智塔科技有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510818264.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-19
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In vertical federated learning, the training efficiency of the GRU model is low and there is a risk of data privacy leakage. The existing encryption method consumes too many computing resources and cannot effectively protect sensitive data.

Method used

Using secret sharing and differential privacy methods, the sample data is encrypted and split and predicted at participating nodes. After calculating the initial model gradient, a noise vector is added, and the GRU model is updated through the central node to avoid direct data sharing and encryption of intermediate results.

Benefits of technology

It improves the federated learning efficiency of the GRU model, reduces computing resource usage, and ensures the privacy and security of training data through secondary encryption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675707A_ABST
    Figure CN120675707A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a longitudinal federated GRU model training method and system based on secret sharing and differential privacy, and the method comprises the steps: determining encrypted sample data, splitting the encrypted sample data, and obtaining a plurality of encrypted sample data fragments; the multiple encrypted sample data fragments are input into a first GRU model, prediction result fragments corresponding to the multiple encrypted sample data fragments are obtained, and the first GRU model is deployed at the participation node; calculating an initial model gradient of a first GRU model according to the prediction result fragment and an encrypted label data fragment sent by the center node; and obtaining a target noise vector, adding the target noise vector to the initial model gradient, obtaining an updated model gradient, and sending the updated model gradient to the center node, so that the center node updates a second GRU model according to the updated model gradient, and the second GRU model is deployed at the center node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a vertical federated GRU model training method and system based on secret sharing and differential privacy. Background Art

[0002] With the development of the Internet of Things, FinTech, and healthcare, the demand for analyzing and modeling time series data has increased dramatically. The Gated Recurrent Unit (GRU) has become a mainstream model for processing time series data due to its high parameter efficiency and fast training speed. However, in real-world applications, time series data is often dispersed among different parties. Due to data privacy and security concerns, some data cannot be directly shared centrally, resulting in the formation of "data silos." For example, banks hold user transaction records, while e-commerce platforms hold user shopping behavior data. Joint modeling of these two is necessary to improve the effectiveness of credit scoring models.

[0003] This led to the emergence of federated learning strategies for models. The core idea of ​​this federated learning strategy is that multiple participants can collaboratively train a global model using their own data, without sharing the original data. However, when using federated learning for GRUs, participants may hold sensitive data (such as medical records and transaction records), which must be protected from leakage. Current data encryption methods, applied to the GRU parameter calculation process, consume significant computing resources, resulting in low federated learning efficiency. Therefore, an effective solution is urgently needed that can maintain the efficiency of federated learning for GRUs while protecting data privacy. Summary of the Invention

[0004] In view of this, embodiments of this specification provide two methods for training vertical federated GRU models based on secret sharing and differential privacy. One or more embodiments of this specification also involve two apparatuses for training vertical federated GRU models based on secret sharing and differential privacy, a system for training vertical federated GRU models based on secret sharing and differential privacy, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a longitudinal federated GRU model training method based on secret sharing and differential privacy is provided, which is applied to participating nodes and includes: Determining encrypted sample data, and splitting the encrypted sample data to obtain a plurality of encrypted sample data fragments; Inputting the plurality of encrypted sample data slices into a first GRU model to obtain prediction result slices corresponding to the plurality of encrypted sample data slices, wherein the first GRU model is deployed on the participating node; Calculating the initial model gradient of the first GRU model according to the prediction result slice and the encrypted label data slice sent by the central node; Obtain a target noise vector, add the target noise vector to the initial model gradient, obtain an updated model gradient, and send the updated model gradient to the central node so that the central node updates a second GRU model according to the updated model gradient, wherein the second GRU model is deployed at the central node.

[0006] According to a second aspect of the embodiments of this specification, a vertical federated GRU model training device based on secret sharing and differential privacy is provided, which is applied to participating nodes and includes: a splitting module configured to determine encrypted sample data and split the encrypted sample data to obtain a plurality of encrypted sample data fragments; an input module configured to input the plurality of encrypted sample data slices into a first GRU model to obtain prediction result slices corresponding to the plurality of encrypted sample data slices, wherein the first GRU model is deployed on the participating node; a calculation module configured to calculate an initial model gradient of the first GRU model based on the prediction result slice and the encrypted label data slice sent by the central node; A sending module is configured to obtain a target noise vector, add the target noise vector to the initial model gradient, obtain an updated model gradient, and send the updated model gradient to the central node so that the central node updates the second GRU model according to the updated model gradient, wherein the second GRU model is deployed at the central node.

[0007] According to a third aspect of the embodiments of this specification, another longitudinal federated GRU model training method based on secret sharing and differential privacy is provided, which is applied to a central node, including: Determining encrypted label data, and splitting the encrypted label data to obtain a plurality of encrypted label data fragments; Sending the multiple encrypted label data slices to participating nodes, and receiving updated model gradients of a first GRU model sent by the participating nodes, wherein the first GRU model is deployed on the participating nodes, and the updated model gradients are obtained by adding a target noise vector to an initial model gradient, and the initial model gradients are calculated based on the multiple encrypted label data slices; The second GRU model is updated according to the updated model gradient to obtain an updated second GRU model, and the updated model parameters of the updated second GRU model are sent to the participating nodes, wherein the second GRU model is deployed at the central node.

[0008] According to a fourth aspect of the embodiments of this specification, another longitudinal federated GRU model training device based on secret sharing and differential privacy is provided, which is applied to a central node, including: a splitting module configured to determine encrypted label data and split the encrypted label data to obtain a plurality of encrypted label data fragments; a communication module configured to send the plurality of encrypted label data slices to a participating node, and receive an updated model gradient of a first GRU model sent by the participating node, wherein the first GRU model is deployed on the participating node, and the updated model gradient is obtained by adding a target noise vector to an initial model gradient, and the initial model gradient is calculated based on the plurality of encrypted label data slices; An update module is configured to update the second GRU model according to the updated model gradient, obtain an updated second GRU model, and send the updated model parameters of the updated second GRU model to the participating nodes, wherein the second GRU model is deployed at the central node.

[0009] According to a fifth aspect of the embodiments of this specification, a vertical federated GRU model training system based on secret sharing and differential privacy is provided, including participating nodes and a central node, wherein: The central node is configured to determine the encrypted label data, split the encrypted label data to obtain a plurality of encrypted label data fragments, and send the plurality of encrypted label data fragments to the participating nodes; The participating node is configured to determine encrypted sample data, split the encrypted sample data to obtain multiple encrypted sample data slices; input the multiple encrypted sample data slices into a first GRU model to obtain prediction result slices corresponding to the multiple encrypted sample data slices; calculate the initial model gradient of the first GRU model based on the prediction result slices and the encrypted label data slices sent by the central node; obtain a target noise vector, add the target noise vector to the initial model gradient, obtain an updated model gradient, and send the updated model gradient to the central node, wherein the first GRU model is deployed on the participating node; The central node is configured to update the second GRU model according to the updated model gradient, obtain an updated second GRU model, and send the updated model parameters of the updated second GRU model to the participating nodes, wherein the second GRU model is deployed at the central node.

[0010] According to a sixth aspect of the embodiments of this specification, there is provided a computing device, including: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above method are implemented.

[0011] According to a seventh aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, and the steps of the above method are implemented when the computer program / instruction is executed by a processor.

[0012] According to an eighth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0013] One embodiment of the present specification provides a longitudinal federated GRU model training method based on secret sharing and differential privacy, which is applied to participating nodes, including: determining encrypted sample data, splitting the encrypted sample data to obtain multiple encrypted sample data shards; inputting the multiple encrypted sample data shards into a first GRU model to obtain prediction result shards corresponding to the multiple encrypted sample data shards, wherein the first GRU model is deployed on the participating nodes; calculating the initial model gradient of the first GRU model based on the prediction result shards and the encrypted label data shards sent by the central node; obtaining a target noise vector, adding the target noise vector to the initial model gradient, obtaining an updated model gradient, and sending the updated model gradient to the central node, so that the central node updates the second GRU model according to the updated model gradient, wherein the second GRU model is deployed on the central node.

[0014] In the above method, when the GRU model is federated trained, the encrypted sample data can be split to obtain multiple encrypted sample data shards, so that no party can restore the original encrypted sample data based on a single encrypted sample data shard. The multiple encrypted sample data shards are input into the first GRU model to obtain prediction result shards corresponding to the multiple encrypted sample data shards. Based on the prediction result shards and the encrypted label data shards sent by the central node, the initial model gradient of the first GRU model is calculated, and a target noise vector is obtained. The target noise vector is added to the initial model gradient to obtain an updated model gradient to achieve secondary encryption of the model gradient. The secondary encrypted updated model gradient is sent to the central node so that the central node updates the second GRU model based on the updated model gradient to achieve subsequent federated training of the second GRU model. The process of encrypting the sample data is pre-placed so that the subsequent parameter calculation process in the first GRU model is directly calculated based on the encrypted sample data shards, avoiding the large amount of computing resources occupied by encrypting the intermediate results, thereby improving the federated learning efficiency of the GRU model and ensuring the privacy and security of the training data through secondary encryption. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flowchart of a longitudinal federated GRU model training method based on secret sharing and differential privacy provided in one embodiment of this specification; Figure 2 This is a flowchart of a longitudinal federated GRU model training method based on secret sharing and differential privacy provided in one embodiment of this specification applied to participating nodes and a central node; Figure 3 This is a flowchart of another longitudinal federated GRU model training method based on secret sharing and differential privacy provided in one embodiment of this specification; Figure 4 This is a flowchart of a processing process of a longitudinal federated GRU model training method based on secret sharing and differential privacy provided in one embodiment of this specification; Figure 5 This is a structural diagram of a vertical federated GRU model training device based on secret sharing and differential privacy provided in one embodiment of this specification; Figure 6 This is a structural diagram of another longitudinal federated GRU model training device based on secret sharing and differential privacy provided in one embodiment of this specification; Figure 7 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0016] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0017] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0018] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0019] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0020] First, the terms involved in one or more embodiments of this specification are explained.

[0021] GRU (Gated Recurrent Unit) is an improved recurrent neural network (RNN) architecture for processing time series data. It introduces a gating mechanism to alleviate the vanishing gradient problem of traditional RNNs, enabling the model to better capture long-term dependencies. It uses two gates to control the flow of information: Update Gate: This determines how much historical information is retained in the current state. Reset Gate: This determines whether to ignore the previous hidden state.

[0022] RNN: Recurrent Neural Network, a recurrent neural network, is a neural network structure designed specifically for processing sequence data. It is suitable for tasks with time dependencies, such as text, speech, or time series modeling.

[0023] MPC (Multi-Party Computation) is a branch of cryptography that involves multiple parties working together to compute the result of a function without revealing their private inputs. It is widely used in privacy-preserving computing.

[0024] SS protocol: Secret Sharing, also known as secret sharing, is a commonly used technology in MPC. Its core idea is to split a secret value into multiple shares. Only when enough shares are combined together can the secret be recovered; otherwise, no information can be obtained.

[0025] Differential privacy is a mathematically defined privacy protection framework designed to ensure the security of personal data during data analysis and machine learning. Its core idea is to add appropriate noise to query results or model training to protect individual data from being leaked while still allowing useful information to be obtained from the overall dataset.

[0026] In practical applications, with the development of fields such as the Internet of Things, FinTech, and healthcare, the demand for analyzing and modeling time series data has increased dramatically. The Gated Recurrent Unit (GRU) has become a mainstream model for processing time series data due to its high parameter efficiency and fast training speed. However, in real-world applications, time series data is often dispersed across different institutions or parties, and due to data privacy and security regulations, data cannot be directly shared centrally, resulting in the formation of "data silos." For example, banks hold user transaction records, while e-commerce platforms hold user shopping behavior data. Joint modeling of these two is required to improve the effectiveness of credit scoring models.

[0027] Federated learning allows multiple participants to collaboratively train models without sharing the original data. This can be categorized into horizontal federated learning, vertical federated learning, and federated transfer learning. Vertical federated learning involves scenarios where all participants have access to the same samples but different feature spaces. However, it still faces the following challenges: privacy protection: Participants in model training may hold sensitive data (such as medical records and transaction records), which must be prevented from leaking. The model's gradient transfer can be reverse-engineered to reveal the training data. High computational complexity: The GRU activation function involves nonlinear operations, making existing secure computation methods for deep learning (such as homomorphic encryption) inefficient.

[0028] Specifically, current vertical federated learning methods rely on logistic regression or tree models, which are unable to capture the dynamic characteristics of time series data. Federated GRU solutions, on the other hand, only support horizontal data partitioning and are unsuitable for vertical scenarios where financial institutions and third-party data platforms complement each other. Furthermore, traditional homomorphic encryption protocols are technically inefficient in GRU gating parameter aggregation, necessitating further exploration of more efficient algorithms. It is mainly reflected in the following aspects: high overhead of nonlinear operations: the activation functions of GRU (such as sigmoid and tanh) need to be approximated by high-order polynomials (such as Taylor expansion) under homomorphic encryption, which makes the single activation calculation time too long and introduces significant approximation errors; defects in the combination of vertical federated learning and GRU: the current vertical federated learning cannot process time series data, and the current federated GRU solution only supports horizontal partitioning, and does not solve the scenario of vertical distribution of features across institutions; explosion of communication rounds: the gate calculation of each time step requires multiple ciphertext multiplication interactions. For example, the calculation of an activation function requires multiple rounds of communication, resulting in an exponential increase in the communication overhead of long sequence training; insufficiency of secret sharing (SS): the current SS protocol has communication bottlenecks in GRU gate calculations (such as multiple rounds of interaction for each iteration) or accuracy loss (such as gradient deviation caused by noise introduction).

[0029] In this specification, two vertical federated GRU model training methods based on secret sharing and differential privacy are provided. This specification also involves two vertical federated GRU model training devices based on secret sharing and differential privacy, a vertical federated GRU model training system based on secret sharing and differential privacy, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0030] See also Figure 1 , Figure 1 A flowchart of a longitudinal federated GRU model training method based on secret sharing and differential privacy provided according to an embodiment of this specification is shown, which is applied to participating nodes and specifically includes the following steps.

[0031] Step 102: Determine encrypted sample data, split the encrypted sample data, and obtain multiple encrypted sample data fragments.

[0032] Specifically, the longitudinal federated GRU model training method based on secret sharing and differential privacy provided in the embodiments of this specification can be applied to longitudinal federated training of GRU models. Longitudinal federated learning is a distributed machine learning method suitable for scenarios where different data holders have minimal overlap in the sample dimension but significant overlap in the feature dimension. For example, two different companies may share some of the same customers but collect different types of data (one may have customer transaction records, while the other may have information about their social media activity). Longitudinal federated learning allows these entities to jointly train models without sharing the original data, thereby protecting data privacy. The GRU (Gated Recurrent Unit) is a variant of the recurrent neural network (RNN) specifically designed to process sequential data. By introducing update and reset gates to control information flow, it effectively addresses the vanishing gradient problem in traditional RNNs, enabling the model to capture long-term dependencies. GRUs are widely used in fields such as natural language processing and time series prediction. The longitudinal federated GRU model combines these two approaches. A longitudinal federated GRU model refers to a GRU model trained within the longitudinal federated learning framework. This setup is particularly suitable for scenarios that require processing sequential data and involve multi-party collaboration. Specifically, suppose multiple organizations wish to jointly train a time series-based prediction model, but each possesses different attribute data for users or objects. For example, in healthcare, hospital A may possess a patient's electronic health record, while hospital B possesses the patient's genomic data. A longitudinal federated GRU model can be used to predict disease progression while ensuring the security of sensitive information.

[0033] In practice, each participating node uploads intermediate results (such as gradients) from local computations rather than raw data. These intermediate results are typically encrypted to further protect privacy. All participating nodes jointly maintain a global GRU model. In each iteration, each participating node calculates local gradients based on its own data and sends them to the central node. The central node aggregates these local gradients to update the global model parameters and then distributes the updated parameters back to each participating node. Throughout this process, since only encrypted gradients are exchanged, no one party can directly access the other's data, effectively protecting the data privacy of all parties. By avoiding direct data sharing, the risk of data leakage is significantly reduced. Integrating features from different sources can build more comprehensive and accurate models, helping to meet increasingly stringent privacy regulations.

[0034] Among them, the encrypted sample data can be understood as the encrypted sample data held and provided by the participating nodes. The sample data can be, for example, the transaction record data of users held by banks, the electronic medical record data of users held by hospitals, etc.

[0035] Specifically, the sample data held and provided by the participating nodes can be encrypted to obtain encrypted sample data, and the encrypted sample data can be split to obtain multiple encrypted sample data fragments, so that the sample data is encrypted and stored in a dispersed manner, and no party can restore the original sample data alone.

[0036] In practical applications, sample data can be time series data, which refers to a series of observations recorded in chronological order. Each observation is associated with a specific point in time or time period. This type of data is widely found in fields such as finance, meteorology, healthcare, and the Internet of Things. Understanding time series data is crucial for analyzing trends, predicting future events, and making data-driven decisions. The basic characteristics of time series data include: Time dependency: A key characteristic of time series data is its sequential nature and temporal continuity. This means that there is some form of dependency between data points, and current values ​​are often influenced by past values. Timestamp: Each record has a distinct time identifier, which can be a specific point in time (e.g., 2023-06-06 09:21) or a fixed time interval (e.g., the first day of each day or month). Periodicity / Seasonality: Many time series data exhibit periodic patterns, such as daytime temperature fluctuations that follow diurnal patterns or stock market fluctuations that follow weekly or annual patterns. Trend: Over the long term, time series data may exhibit an upward or downward trend, reflecting the ongoing influence of internal or external factors on the data. Noise: Real time series data often contains random fluctuations or noise, which are short-term changes caused by unpredictable factors.

[0037] The GRU model trained using the longitudinal federated GRU model training method based on secret sharing and differential privacy provided in this specification can be applied to time series data analysis scenarios, such as financial market analysis: stock prices, exchange rates, interest rates, and other typical time series data can help investors make buying and selling decisions by analyzing historical data; sales forecasting scenarios: retailers can use past sales records to predict future sales, thereby optimizing inventory management and marketing strategies; environmental monitoring scenarios: weather stations collect data such as temperature, humidity, and wind speed for weather forecasting and climate research; industrial monitoring scenarios: sensors in manufacturing plants continuously record equipment status parameters (such as temperature and pressure) for preventive maintenance; and health monitoring scenarios: wearable devices record information such as heart rate and step count, which can contribute to personal health management and medical research. The embodiments of this specification are not limited to these scenarios.

[0038] In a specific implementation, the step of determining the encrypted sample data includes: Obtaining sample data stored in the participating nodes; The sample data is encrypted according to a preset encryption algorithm to obtain encrypted sample data.

[0039] In actual applications, the preset encryption algorithm can be a symmetric encryption algorithm, an asymmetric encryption algorithm, a stream encryption algorithm, a homomorphic encryption algorithm, a timestamp encryption algorithm or a Beaver triple encryption algorithm, etc., or the sample data can also be encrypted according to the MPC protocol parameters negotiated between the participating nodes and the central node. The embodiments of this specification do not limit this.

[0040] Specifically, Beaver triples can be understood as a technique in multi-party secure computation (MPC) for performing secure two-party or multi-party computations without exposing the private inputs of participating nodes. Beaver triples can be used to implement secure multiplication operations. A Beaver triple consists of three random numbers a, b, and c, where c = a × b.

[0041] In summary, by encrypting the sample data to obtain encrypted sample data, the initial encryption of the sample data is achieved, and there is no need to encrypt it in the intermediate calculation process of the subsequent GRU model, thereby reducing the occupation of computing resources.

[0042] Furthermore, after obtaining the sample data stored in the participating nodes, the method further includes: performing standardization processing on the sample data to obtain standardized sample data; The step of encrypting the sample data according to a preset encryption algorithm to obtain encrypted sample data includes: The standardized sample data is encrypted according to a preset encryption algorithm to obtain encrypted sample data.

[0043] Specifically, the sample data may be standardized to obtain standardized sample data, and the standardized sample data may be encrypted according to a preset encryption algorithm to obtain encrypted sample data, which may then be split to obtain multiple encrypted sample data fragments.

[0044] In practical applications, the sample data can be the local time series data of the participating node, such as the time series feature matrix. Then, the time series data can be standardized, such as the time series data can be Z-score standardized to obtain standardized sample data. The standardized sample data can be encrypted and split into multiple encrypted sample data shards (i.e., secret sharing shards), and the multiple encrypted sample data shards can be dispersedly stored in the participating node.

[0045] In summary, by standardizing the sample data and achieving normalized processing of the sample data, the quality of the sample data can be improved, making the subsequent GRU model training process more effective, speeding up the model training speed, reducing the number of iterations required to reach convergence, and improving the model accuracy; in addition, the sample data is encrypted and stored in a decentralized manner, so that no other participating nodes or central nodes can restore the original sample data alone, thereby ensuring the privacy and security of the data.

[0046] Step 104: Input the multiple encrypted sample data slices into a first GRU model to obtain prediction result slices corresponding to the multiple encrypted sample data slices, wherein the first GRU model is deployed on the participating node.

[0047] Among them, the first GRU model can be understood as a GRU model deployed locally on the participating node. The first GRU can be used to calculate the encrypted sample data provided by the participating node in the participating node, obtain the prediction result shards, and be used for subsequent calculation of the model gradient of the first GRU model.

[0048] Specifically, multiple encrypted sample data shards can be input into the first GRU model, and calculations can be performed in the first GRU model to obtain prediction result shards corresponding to the multiple encrypted sample data shards. Furthermore, there can be multiple prediction result shards, and each encrypted sample data shard corresponds to a prediction result shard.

[0049] In practical applications, during the joint training of the GRU model, the chanting state of the current time step and the prediction result slice outputted at the current time step can be calculated during the forward propagation process. The specific implementation method is as follows: the multiple encrypted sample data slices are inputted into the first GRU model to obtain the prediction result slices corresponding to the multiple encrypted sample data slices, including: Input the multiple encrypted sample data slices into a first GRU model, and calculate the update gate and reset gate of the first GRU model according to the multiple encrypted sample data, the polynomial corresponding to the first activation function, and the hidden state of the previous time step; Calculating a candidate hidden state corresponding to the first GRU model according to the reset gate, the polynomial corresponding to the second activation function, and the hidden state of the previous time step; Calculating the hidden state of the current time step according to the update gate and the candidate hidden state; According to the hidden state of the current time step, the prediction result fragments corresponding to the multiple encrypted sample data fragments are calculated and output.

[0050] The GRU model controls information flow by introducing update and reset gates, effectively managing the process of memorization and forgetting. The update and reset gates can be understood as the update gate vector and reset gate vector, respectively. The update gate determines the extent to which new information at the current moment should be used to update the hidden state. It weights the ratio of new and old information with a value between 0 and 1. The update gate helps the GRU model decide how much old information to retain and how much new input to accept. If the update gate is close to 1, more old information is retained; if the update gate is close to 0, more new information is adopted.

[0051] The reset gate determines how much old information to ignore when calculating new candidate hidden states. It can be thought of as a filter, selectively discarding some historical information. A low reset gate indicates a preference for ignoring some old information to better adapt to new input. Conversely, a high reset gate indicates a preference for prior states.

[0052] The hidden state is one of the core outputs of the GRU model. It contains all the sequence information processed by the GRU model. At each time step, the hidden state is adjusted according to the guidance of the update gate and reset gate. The hidden state can be regarded as the memory unit of the GRU, storing all the relevant information of the GRU model from the beginning to the current moment.

[0053] The first activation function and the second activation function are both gating parameters of the GRU model. The first activation function can be a sigmoid function, and the second activation function can be a tanh function. The polynomial corresponding to the first activation function can be an approximate representation of the first activation function, and the polynomial corresponding to the second activation function can be an approximate representation of the second activation function.

[0054] Specifically, after multiple encrypted sample data slices are input into the first GRU model, the first GRU model can calculate an update gate and a reset gate for the first GRU model based on the multiple encrypted sample data, a polynomial corresponding to the sigmoid function, and the hidden state of the previous time step. In the case of an initial calculation, the hidden state of the previous time step can be an initial hidden state (i.e., an all-zero vector). Based on the reset gate, the polynomial corresponding to the tanh function, and the hidden state of the previous time step, candidate hidden states corresponding to the first GRU model are calculated. Based on the update gate and the candidate hidden states, the hidden state of the current time step is calculated. Finally, based on the hidden state of the current time step, prediction result slices corresponding to the multiple encrypted sample data slices are calculated and output.

[0055] In practical applications, during the forward propagation phase of the jointly trained GRU model, each participating node can set its initial hidden state to an all-zero vector and store it in a secret-sharing format. This initial hidden state can be encrypted and sharded, allowing for decentralized storage. Each participating node then calculates an update gate and a reset gate using secure matrix multiplication based on the current hidden state (i.e., the hidden state at the previous time step, which serves as the initial hidden state during the first calculation) and multiple encrypted sample data shards. The candidate hidden state is then calculated by combining the reset gate and the current hidden state. The new hidden state (i.e., the hidden state at the current time step) is then blended with the previous hidden state using the update gate. During the backward propagation phase of the jointly trained GRU model, each participating node can perform a linear transformation on the hidden state at the current time step t to obtain a prediction result shard.

[0056] Specifically, the update gate can be calculated according to the following formula.

[0057]

[0058] Where t is the current time step, is the encrypted sample data shard currently input, is the hidden state at the previous time step, is the polynomial corresponding to the first activation function. It is the update gate. is the bias term of the update gate, is the weight matrix of the update gate.

[0059] The reset gate can be calculated according to the following formula.

[0060]

[0061] in, It is the reset gate. is the weight matrix of the reset gate, is the bias term for the reset gate.

[0062] The candidate hidden states can be calculated according to the following formula.

[0063]

[0064] in, is a candidate hidden state, is the polynomial corresponding to the second activation function, is the weight of the hidden state, is the bias term for the hidden state.

[0065] The hidden state of the current time step can be calculated according to the following formula .

[0066]

[0067] The prediction result sharding can be calculated according to the following formula .

[0068]

[0069] in, It is the bias term of the prediction result output by the first GRU model.

[0070] In summary, through the secret sharing protocol, the encrypted sample data shards are calculated to achieve privacy-preserving collaborative computing of the training data of each participating node, ensuring that the original data of each participating node is not leaked during the training process, thereby ensuring data privacy security.

[0071] Furthermore, before calculating the update gate and the reset gate of the first GRU model according to the multiple encrypted sample data, the polynomial corresponding to the first activation function, and the hidden state of the previous time step, the method further includes: Approximately expressing the first activation function as a polynomial corresponding to the first activation function; The second activation function is approximately represented as a polynomial corresponding to the second activation function.

[0072] Specifically, the first activation function can be approximated and decomposed into multiple low-order polynomial terms to obtain the low-order polynomial corresponding to the first activation function; accordingly, the second activation function can be approximated and decomposed into multiple low-order polynomial terms to obtain the low-order polynomial corresponding to the second activation function.

[0073] In practical applications, the polynomial corresponding to the first activation function is shown in the following formula.

[0074]

[0075] The polynomial corresponding to the second activation function is shown in the following formula:

[0076] In summary, by decomposing the first activation function and the second activation function of the GRU model into multiple low-order polynomial terms and performing distributed computing on each participating node through a secret sharing protocol, the computational complexity of the gating function of the GRU model is reduced from the polynomial level to the linear level, which is particularly suitable for long sequence data scenarios; in terms of model performance, the gating structure and time series modeling capabilities of the GRU are fully retained, which is better than replacing the GRU with a linear model.

[0077] Step 106: Calculate the initial model gradient of the first GRU model based on the prediction result slice and the encrypted label data slice sent by the central node.

[0078] The central node can be understood as a participant in the federated training of the GRU model. This central node can hold labeled data, such as default labels in financial risk control or disease labels in medical prediction. Furthermore, the central node can also function as a central server, aggregating model gradients sent by participating nodes and performing model training for the GRU model.

[0079] Specifically, the central node can encrypt the label data it holds to obtain encrypted label data, split the encrypted label data to obtain multiple encrypted label data fragments, and send the multiple encrypted label data fragments to each participating node. Each participating node can calculate the initial model gradient of the first GRU model deployed locally on each participating node based on the prediction result fragments calculated by itself and the multiple encrypted label data fragments sent by the central node. It can be understood that the process of encrypting and splitting the label data by the central node is similar to the process of encrypting and splitting the sample data by the participating nodes described above, and this embodiment of the specification will not be repeated.

[0080] In practical applications, the central node can negotiate MPC protocol parameters with at least one participating node. The MPC protocol parameters may include parameters such as finite field size and random number seed, which can be used for data encryption and data transmission. Set the participant set ,in The central node holds the label data, and the rest are participating nodes. All parties negotiate the MPC protocol parameters (such as the size of the finite field of secret sharing, random number seed), and perform local time series data Standardize and generate secret sharing value< > (i.e., encrypted sample data shards or encrypted tag data shards), where: Represents a sample or feature matrix of time series data, indicating that the dimension of the time series data is T×di.

[0081] In a specific implementation, the initial model gradient of the first GRU model is calculated based on the prediction result slice and the encrypted label data slice sent by the central node, including: Calculating a model loss function based on the prediction result shards and the encrypted label data shards sent by the central node; Calculate the initial model gradient of the first GRU model according to the model loss function.

[0082] Specifically, during the joint training of the GRU model, in back propagation, after each participating node calculates the prediction result slice, it can receive the encrypted label data slice sent by the central node, and calculate the model loss function based on the prediction result slice and the encrypted label data slice, and calculate the initial model gradient of the first GRU model based on the model loss function. The initial model gradient can be understood as the local gradient of the GRU model.

[0083] In practical applications, the mean squared error can be used as the model loss function, and the initial model gradient of the first GRU model can be calculated using secure matrix multiplication. Secure matrix multiplication can be optimized using Beaver triples. The model loss function can be expressed as follows.

[0084]

[0085] in, It is the encrypted label data fragment sent by the central node.

[0086] In summary, for each participating node, by calculating the initial model gradient based on the prediction result shards and the encrypted label data shards, a training basis is provided for the subsequent federated training of the GRU model while ensuring data privacy and security.

[0087] Step 108: Obtain a target noise vector, add the target noise vector to the initial model gradient, obtain an updated model gradient, and send the updated model gradient to the central node so that the central node updates the second GRU model according to the updated model gradient, wherein the second GRU model is deployed on the central node.

[0088] The second GRU model and the first GRU model can be the same GRU model, deployed on different nodes. The second GRU model deployed on the central node can be trained on the central node. The model parameters of the second GRU model obtained after training can be sent to the participating nodes, and the model parameters of the first GRU model local to the participating nodes are updated to achieve federated learning of the GRU model. The target noise vector can be a noise vector whose noise direction is most similar to the gradient direction of the initial model gradient.

[0089] In practical applications, when there are multiple participating nodes, for each participating node, a target noise vector can be obtained, the target noise vector can be added to the initial model gradient, the updated model gradient can be obtained, and the updated model gradient can be sent to the central node. The central node aggregates the updated model gradient sent by each participating node and then trains the second GRU model.

[0090] Furthermore, obtaining the target noise vector includes: Determining a plurality of candidate noise vectors, wherein each candidate noise vector has a different noise direction; A similarity between the initial model gradient and the multiple candidate noise vectors is calculated, and a target noise vector is selected from the multiple candidate noise vectors according to the similarity.

[0091] Specifically, multiple candidate noise vectors with different noise directions can be determined. The similarity between the initial model gradient and each candidate noise vector is calculated, and the candidate noise vector with the highest similarity is determined as the target noise vector. The target noise vector can then be used to update the initial model gradient to obtain an updated model gradient.

[0092] In practical applications, a noise direction pool can be created , the noise direction pool may include multiple candidate noise vectors of different noise directions, the candidate noise vectors may be randomly generated using a Gaussian distribution and normalized to unit vectors, so that each candidate noise vector has a unit norm (i.e., normalized), where, , indicating that each candidate noise vector has a mean of 0 and a variance of Normal distribution (i.e. Gaussian distribution).

[0093] The cosine similarity between each candidate noise vector and the initial model gradient can be calculated by the following formula.

[0094]

[0095] in, is the initial model gradient, is the cosine similarity, For each candidate noise vector.

[0096] The formula for updating the initial model gradient using the target noise vector to obtain the updated model gradient is as follows.

[0097]

[0098] in, To update the model gradient, .

[0099] In summary, differential privacy processing of the initial model gradient is achieved by adding a target noise vector to the initial model gradient. The target noise vector considers not only the amplitude but also the noise direction. The noise direction of the target noise vector is consistent with the gradient direction of the initial model gradient, which protects the gradient while avoiding the performance degradation of the GRU model caused by excessive gradient deviation.

[0100] Furthermore, obtaining the target noise vector includes: The intensity of the target noise vector is adjusted according to the data sensitivity of the encrypted sample data and the encrypted label data fragment to obtain an adjusted target noise vector.

[0101] Specifically, during each round of iteration of the first GRU model and the second GRU model, the intensity of the acquired target noise vector (i.e., the mask noise intensity) can be adjusted according to the data sensitivity through a dynamic mask mechanism to obtain an adjusted target noise vector, thereby balancing the accuracy and security of the GRU model.

[0102] In practical applications, the hidden state of the current time step can also be encrypted and aggregated by adopting a dynamic mask mechanism, where the mask noise intensity is adaptively adjusted according to the feature sensitivity.

[0103] Furthermore, after sending the updated model gradient to the central node, the method further includes: receiving updated model parameters of the second GRU model sent by the central node, wherein the updated model parameters are obtained by the central node after updating the second GRU model according to the updated model gradient; Adjusting the model parameters of the first GRU model according to the updated model parameters of the second GRU model to obtain an adjusted first GRU model; Continue to perform the steps of determining the encrypted sample data, splitting the encrypted sample data, and obtaining a plurality of encrypted sample data fragments until a second GRU model that meets the training stop condition is obtained.

[0104] Specifically, the central node can aggregate the updated model gradients sent by each participating node according to the multi-party security protocol (i.e., MPC protocol parameters) to obtain a global gradient. The central node then uses the global gradient to adjust the model parameters of the second GRU model, obtaining the updated model parameters of the second GRU model. The updated model parameters of the second GRU model are then sent to each participating node. Each participating node can adjust the model parameters of the locally deployed first GRU model based on the updated model parameters of the second GRU model to obtain the adjusted first GRU model, after which the next round of training can be continued, thereby iteratively training the GRU model.

[0105] In practice, the updated model gradients calculated by each participating node are encrypted and transmitted to a central node through secret sharing, or gradients are aggregated by multiple participants. The MPC protocol ensures that each participating node can only see its own model gradients and cannot access the private data of other parties. Participants aggregate the encrypted model gradients through a secure computation protocol to obtain the global gradient.

[0106] The training stop condition can be that the number of training times reaches a preset threshold and / or the model loss function reaches convergence. Figure 2 , Figure 2 A flowchart of a longitudinal federated GRU model training method based on secret sharing and differential privacy provided in accordance with an embodiment of this specification is shown, which specifically includes the following steps.

[0107] Step 202: System initialization: the central node and each participating node negotiate MPC parameters, initialize the model weights of the GRU model, standardize the sample data and label data, and initialize the noise pool.

[0108] Step 204: For each participating node, calculate the update gate, reset gate, candidate hidden state and hidden state of the GRU model at the current time step.

[0109] Step 206: For each participating node, obtain the prediction result fragment output by the GRU model.

[0110] Step 208: The central node sends the encrypted label data fragments to each participating node.

[0111] Step 210: Each participating node calculates the model loss function based on the prediction result shards and the encrypted label data shards.

[0112] Step 212: Determine whether the model loss function has converged. If so, execute step 214; if not, execute step 216.

[0113] Step 214: End training.

[0114] Step 216: For each participating node, use the optimizer to calculate the initial model gradient of the GRU model.

[0115] Step 218: Perform differential privacy processing on the optimized initial model gradient to obtain an updated model gradient.

[0116] Step 220: Perform gradient aggregation on the updated model gradients calculated by each participating node to obtain a global gradient, and adjust the model parameters of the GRU model according to the global gradient.

[0117] In summary, the longitudinal federated GRU model training method based on secret sharing and differential privacy provided in the embodiments of this specification optimizes computational efficiency through a secret sharing protocol while ensuring privacy, and improves the adaptability of secret sharing and GRU. In the calculation of the reset gate / update gate of the GRU, a segmented secret sharing protocol is designed (i.e., the activation function is split into polynomial approximation + secret sharing), rather than directly applying the general SS protocol, and Beaver triples are used to reduce the computational complexity of model training. Regarding the secure aggregation of hidden states, a dynamic mask mechanism is introduced, and the noise intensity of the target noise vector of secret sharing is adaptively adjusted in each round of iteration according to the data sensitivity, balancing accuracy and security, and considering not only the noise amplitude but also the noise direction.

[0118] In summary, in the above method, when the GRU model is federated trained, the encrypted sample data can be split to obtain multiple encrypted sample data shards, so that no party can restore the original encrypted sample data based on a single encrypted sample data shard. The multiple encrypted sample data shards are input into the first GRU model to obtain prediction result shards corresponding to the multiple encrypted sample data shards. Based on the prediction result shards and the encrypted label data shards sent by the central node, the initial model gradient of the first GRU model is calculated, the target noise vector is obtained, and the target noise vector is added to the initial model gradient to obtain the updated model gradient to achieve secondary encryption of the model gradient. The secondary encrypted updated model gradient is sent to the central node so that the central node updates the second GRU model according to the updated model gradient to achieve subsequent federated training of the second GRU model. The process of encrypting the sample data is pre-placed so that the subsequent parameter calculation process in the first GRU model is directly calculated based on the encrypted sample data shards, avoiding the large amount of computing resources occupied by encrypting the intermediate results, thereby improving the federated learning efficiency of the GRU model and ensuring the privacy and security of the training data through secondary encryption.

[0119] Corresponding to the above method embodiment, this specification embodiment also provides a longitudinal federated GRU model training method based on secret sharing and differential privacy, which is applied to the central node, see Figure 3 , Figure 3 A flowchart of another longitudinal federated GRU model training method based on secret sharing and differential privacy provided according to an embodiment of this specification is shown, which specifically includes the following steps.

[0120] Step 302: Determine encrypted label data, split the encrypted label data, and obtain multiple encrypted label data fragments; Step 304: Send the multiple encrypted label data slices to the participating nodes, and receive updated model gradients of the first GRU model sent by the participating nodes, wherein the first GRU model is deployed on the participating nodes, and the updated model gradients are obtained by adding a target noise vector to the initial model gradients, and the initial model gradients are calculated based on the multiple encrypted label data slices; Step 306: Update the second GRU model according to the updated model gradient to obtain an updated second GRU model, and send the updated model parameters of the updated second GRU model to the participating nodes, wherein the second GRU model is deployed at the central node.

[0121] In an optional embodiment, there are multiple participating nodes; The updating of the second GRU model according to the updated model gradient to obtain an updated second GRU model includes: Aggregate the updated model gradients sent by multiple participating nodes to obtain the target model gradient; The second GRU model is updated according to the target model gradient to obtain an updated second GRU model.

[0122] The target model gradient can be understood as the global gradient. The central node and the participating nodes can be understood as the central node and the participating nodes in the vertical federated GRU model training method based on secret sharing and differential privacy, and the embodiments of this specification will not be repeated here.

[0123] The above is a schematic scheme of a vertical federated GRU model training method based on secret sharing and differential privacy in this embodiment. It should be noted that the technical solution of the vertical federated GRU model training method based on secret sharing and differential privacy and the technical solution of the vertical federated GRU model training method based on secret sharing and differential privacy belong to the same concept. For details not described in detail in the technical solution of the vertical federated GRU model training method based on secret sharing and differential privacy, please refer to the description of the technical solution of the vertical federated GRU model training method based on secret sharing and differential privacy.

[0124] The following combined Figure 4 Taking the application of the vertical federated GRU model training method based on secret sharing and differential privacy provided in this specification in the central node and participating nodes as an example, the vertical federated GRU model training method based on secret sharing and differential privacy is further explained. Figure 4 A flowchart of a processing process of a longitudinal federated GRU model training method based on secret sharing and differential privacy provided in an embodiment of this specification is shown, which specifically includes the following steps.

[0125] Step 402: The central node determines the encrypted label data, splits the encrypted label data, and obtains multiple encrypted label data fragments.

[0126] Step 404: The central node sends multiple encrypted tag data fragments to the participating nodes.

[0127] Step 406: The participating nodes determine the encrypted sample data, split the encrypted sample data, and obtain multiple encrypted sample data slices; input the multiple encrypted sample data slices into the first GRU model to obtain the prediction result slices corresponding to the multiple encrypted sample data slices; calculate the initial model gradient of the first GRU model based on the prediction result slices and the encrypted label data slices sent by the central node; obtain the target noise vector, add the target noise vector to the initial model gradient, and obtain the updated model gradient.

[0128] Step 408: The participating nodes send the updated model gradients to the central node.

[0129] Step 410: The central node updates the second GRU model according to the updated model gradient to obtain an updated second GRU model.

[0130] Step 412: The central node sends the updated model parameters of the updated second GRU model to the participating nodes.

[0131] Step 414: The participating nodes update the model parameters of the first GRU model according to the updated model parameters to obtain an updated first GRU model.

[0132] It can be understood that the above steps 402 to 414 can be understood as a round of federated training for the first GRU model and the second GRU model, and the above steps can be repeated to perform the next round of federated training.

[0133] In the above method, when the GRU model is federated trained, the encrypted sample data can be split to obtain multiple encrypted sample data shards, so that no party can restore the original encrypted sample data based on a single encrypted sample data shard. The multiple encrypted sample data shards are input into the first GRU model to obtain prediction result shards corresponding to the multiple encrypted sample data shards. Based on the prediction result shards and the encrypted label data shards sent by the central node, the initial model gradient of the first GRU model is calculated, and a target noise vector is obtained. The target noise vector is added to the initial model gradient to obtain an updated model gradient to achieve secondary encryption of the model gradient. The secondary encrypted updated model gradient is sent to the central node so that the central node updates the second GRU model based on the updated model gradient to achieve subsequent federated training of the second GRU model. The process of encrypting the sample data is pre-placed so that the subsequent parameter calculation process in the first GRU model is directly calculated based on the encrypted sample data shards, avoiding the large amount of computing resources occupied by encrypting the intermediate results, thereby improving the federated learning efficiency of the GRU model and ensuring the privacy and security of the training data through secondary encryption.

[0134] Corresponding to the above method embodiment, this specification also provides an embodiment of a vertical federated GRU model training device based on secret sharing and differential privacy, which is applied to participating nodes, Figure 5 FIG1 shows a schematic diagram of a longitudinal federated GRU model training device based on secret sharing and differential privacy provided by an embodiment of this specification. Figure 5 As shown, the device includes: A splitting module 502 is configured to determine encrypted sample data and split the encrypted sample data to obtain a plurality of encrypted sample data fragments; An input module 504 is configured to input the plurality of encrypted sample data slices into a first GRU model to obtain prediction result slices corresponding to the plurality of encrypted sample data slices, wherein the first GRU model is deployed on the participating node; A calculation module 506 is configured to calculate an initial model gradient of the first GRU model based on the prediction result slice and the encrypted label data slice sent by the central node; The sending module 508 is configured to obtain a target noise vector, add the target noise vector to the initial model gradient, obtain an updated model gradient, and send the updated model gradient to the central node so that the central node updates the second GRU model according to the updated model gradient, wherein the second GRU model is deployed at the central node.

[0135] In an optional embodiment, the sending module 508 is further configured to: Determining a plurality of candidate noise vectors, wherein each candidate noise vector has a different noise direction; A similarity between the initial model gradient and the multiple candidate noise vectors is calculated, and a target noise vector is selected from the multiple candidate noise vectors according to the similarity.

[0136] In an optional embodiment, the input module 504 is further configured to: Input the multiple encrypted sample data slices into a first GRU model, and calculate the update gate and reset gate of the first GRU model according to the multiple encrypted sample data, the polynomial corresponding to the first activation function, and the hidden state of the previous time step; Calculating a candidate hidden state corresponding to the first GRU model according to the reset gate, the polynomial corresponding to the second activation function, and the hidden state of the previous time step; Calculating the hidden state of the current time step according to the update gate and the candidate hidden state; According to the hidden state of the current time step, the prediction result fragments corresponding to the multiple encrypted sample data fragments are calculated and output.

[0137] In an optional embodiment, the input module 504 is further configured to: Approximately expressing the first activation function as a polynomial corresponding to the first activation function; The second activation function is approximately represented as a polynomial corresponding to the second activation function.

[0138] In an optional embodiment, the calculation module 506 is further configured to: Calculating a model loss function based on the prediction result shards and the encrypted label data shards sent by the central node; Calculate the initial model gradient of the first GRU model according to the model loss function.

[0139] In an optional embodiment, the splitting module 502 is further configured to: Obtaining sample data stored in the participating nodes; The sample data is encrypted according to a preset encryption algorithm to obtain encrypted sample data.

[0140] In an optional embodiment, the splitting module 502 is further configured to: performing standardization processing on the sample data to obtain standardized sample data; The step of encrypting the sample data according to a preset encryption algorithm to obtain encrypted sample data includes: The standardized sample data is encrypted according to a preset encryption algorithm to obtain encrypted sample data.

[0141] In an optional embodiment, the sending module 508 is further configured to: receiving updated model parameters of the second GRU model sent by the central node, wherein the updated model parameters are obtained by the central node after updating the second GRU model according to the updated model gradient; Adjusting the model parameters of the first GRU model according to the updated model parameters of the second GRU model to obtain an adjusted first GRU model; Continue to perform the steps of determining the encrypted sample data, splitting the encrypted sample data, and obtaining a plurality of encrypted sample data fragments until a second GRU model that meets the training stop condition is obtained.

[0142] In an optional embodiment, the sending module 508 is further configured to: The intensity of the target noise vector is adjusted according to the data sensitivity of the encrypted sample data and the encrypted label data fragment to obtain an adjusted target noise vector.

[0143] In the above-mentioned device, when the GRU model is federated trained, the encrypted sample data can be split to obtain multiple encrypted sample data shards, so that no party can restore the original encrypted sample data based on a single encrypted sample data shard. The multiple encrypted sample data shards are input into the first GRU model to obtain prediction result shards corresponding to the multiple encrypted sample data shards. The initial model gradient of the first GRU model is calculated based on the prediction result shards and the encrypted label data shards sent by the central node, and the target noise vector is obtained. The target noise vector is added to the initial model gradient to obtain the updated model gradient to achieve secondary encryption of the model gradient. The secondary encrypted updated model gradient is sent to the central node so that the central node updates the second GRU model according to the updated model gradient to achieve subsequent federated training of the second GRU model. The process of encrypting the sample data is pre-placed so that the subsequent parameter calculation process in the first GRU model is directly calculated based on the encrypted sample data shards, avoiding the large amount of computing resources occupied by encrypting the intermediate results, thereby improving the federated learning efficiency of the GRU model and ensuring the privacy and security of the training data through secondary encryption.

[0144] The above is a schematic scheme of a vertical federated GRU model training device based on secret sharing and differential privacy in this embodiment. It should be noted that the technical solution of the vertical federated GRU model training device based on secret sharing and differential privacy and the technical solution of the vertical federated GRU model training method based on secret sharing and differential privacy belong to the same concept. For details not described in detail in the technical solution of the vertical federated GRU model training device based on secret sharing and differential privacy, please refer to the description of the technical solution of the vertical federated GRU model training method based on secret sharing and differential privacy.

[0145] Corresponding to the above method embodiment, this specification also provides an embodiment of a vertical federated GRU model training device based on secret sharing and differential privacy, which is applied to a central node. Figure 6 FIG1 shows a structural diagram of another longitudinal federated GRU model training device based on secret sharing and differential privacy provided by an embodiment of this specification. Figure 6 As shown, the device includes: A splitting module 602 is configured to determine encrypted tag data and split the encrypted tag data to obtain multiple encrypted tag data fragments; a communication module 604 configured to send the plurality of encrypted label data slices to a participating node, and receive an updated model gradient of a first GRU model sent by the participating node, wherein the first GRU model is deployed on the participating node, and the updated model gradient is obtained by adding a target noise vector to an initial model gradient, and the initial model gradient is calculated based on the plurality of encrypted label data slices; The updating module 606 is configured to update the second GRU model according to the updated model gradient, obtain an updated second GRU model, and send the updated model parameters of the updated second GRU model to the participating nodes, wherein the second GRU model is deployed at the central node.

[0146] In an optional embodiment, there are multiple participating nodes; The updating module 606 is further configured to: Aggregate the updated model gradients sent by multiple participating nodes to obtain the target model gradient; The second GRU model is updated according to the target model gradient to obtain an updated second GRU model.

[0147] The above is a schematic scheme of a vertical federated GRU model training device based on secret sharing and differential privacy in this embodiment. It should be noted that the technical solution of the vertical federated GRU model training device based on secret sharing and differential privacy and the technical solution of the vertical federated GRU model training method based on secret sharing and differential privacy belong to the same concept. For details not described in detail in the technical solution of the vertical federated GRU model training device based on secret sharing and differential privacy, please refer to the description of the technical solution of the vertical federated GRU model training method based on secret sharing and differential privacy.

[0148] Corresponding to the above method embodiment, the embodiment of this specification also provides a vertical federated GRU model training system based on secret sharing and differential privacy, including participating nodes and a central node, wherein: The central node is configured to determine the encrypted label data, split the encrypted label data to obtain a plurality of encrypted label data fragments, and send the plurality of encrypted label data fragments to the participating nodes; The participating node is configured to determine encrypted sample data, split the encrypted sample data to obtain multiple encrypted sample data slices; input the multiple encrypted sample data slices into a first GRU model to obtain prediction result slices corresponding to the multiple encrypted sample data slices; calculate the initial model gradient of the first GRU model based on the prediction result slices and the encrypted label data slices sent by the central node; obtain a target noise vector, add the target noise vector to the initial model gradient, obtain an updated model gradient, and send the updated model gradient to the central node, wherein the first GRU model is deployed on the participating node; The central node is configured to update the second GRU model according to the updated model gradient, obtain an updated second GRU model, and send the updated model parameters of the updated second GRU model to the participating nodes, wherein the second GRU model is deployed at the central node.

[0149] The above is a schematic scheme of a vertical federated GRU model training system based on secret sharing and differential privacy in this embodiment. It should be noted that the technical scheme of the vertical federated GRU model training system based on secret sharing and differential privacy and the technical scheme of the vertical federated GRU model training method based on secret sharing and differential privacy are of the same concept. For details not described in detail in the technical scheme of the vertical federated GRU model training system based on secret sharing and differential privacy, please refer to the description of the technical scheme of the vertical federated GRU model training method based on secret sharing and differential privacy.

[0150] Figure 77 shows a block diagram of a computing device 700 according to one embodiment of the present disclosure. Components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.

[0151] Computing device 700 also includes an access device 740 that enables computing device 700 to communicate via one or more networks 760. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 740 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0152] In one embodiment of the present application, the above components of the computing device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 7 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.

[0153] Computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 700 can also be a mobile or stationary server.

[0154] The processor 720 is configured to execute the following computer program / instruction, which implements the steps of the above method when executed by the processor.

[0155] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the computing device embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the description of the method embodiment.

[0156] An embodiment of the present specification further provides a computer-readable storage medium storing a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0157] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the computer-readable storage medium embodiment is generally similar to the method embodiment, so its description is relatively simple. For relevant portions, refer to the description of the method embodiment.

[0158] An embodiment of the present specification further provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0159] The above is an illustrative solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the above method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above method.

[0160] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0161] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0162] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0163] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0164] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A longitudinal federated GRU model training method based on secret sharing and differential privacy, applied to participating nodes, including: Determining encrypted sample data, and splitting the encrypted sample data to obtain a plurality of encrypted sample data fragments; Inputting the plurality of encrypted sample data slices into a first GRU model to obtain prediction result slices corresponding to the plurality of encrypted sample data slices, wherein the first GRU model is deployed on the participating node; Calculating the initial model gradient of the first GRU model according to the prediction result slice and the encrypted label data slice sent by the central node; Obtain a target noise vector, add the target noise vector to the initial model gradient, obtain an updated model gradient, and send the updated model gradient to the central node so that the central node updates a second GRU model according to the updated model gradient, wherein the second GRU model is deployed at the central node.

2. The method according to claim 1, wherein obtaining the target noise vector comprises: Determining a plurality of candidate noise vectors, wherein each candidate noise vector has a different noise direction; A similarity between the initial model gradient and the multiple candidate noise vectors is calculated, and a target noise vector is selected from the multiple candidate noise vectors according to the similarity.

3. The method according to claim 1, wherein inputting the plurality of encrypted sample data slices into the first GRU model to obtain prediction result slices corresponding to the plurality of encrypted sample data slices comprises: Input the multiple encrypted sample data slices into a first GRU model, and calculate the update gate and reset gate of the first GRU model according to the multiple encrypted sample data, the polynomial corresponding to the first activation function, and the hidden state of the previous time step; Calculating a candidate hidden state corresponding to the first GRU model according to the reset gate, the polynomial corresponding to the second activation function, and the hidden state of the previous time step; Calculating the hidden state of the current time step according to the update gate and the candidate hidden state; According to the hidden state of the current time step, the prediction result fragments corresponding to the multiple encrypted sample data fragments are calculated and output.

4. The method according to claim 3, before calculating the update gate and reset gate of the first GRU model based on the plurality of encrypted sample data, the polynomial corresponding to the first activation function, and the hidden state of the previous time step, further comprising: Approximately expressing the first activation function as a polynomial corresponding to the first activation function; The second activation function is approximately represented as a polynomial corresponding to the second activation function.

5. The method according to claim 1, wherein calculating the initial model gradient of the first GRU model based on the prediction result slice and the encrypted label data slice sent by the central node comprises: Calculating a model loss function based on the prediction result shards and the encrypted label data shards sent by the central node; Calculate the initial model gradient of the first GRU model according to the model loss function.

6. The method according to any one of claims 1 to 5, wherein determining the encrypted sample data comprises: Obtaining sample data stored in the participating nodes; The sample data is encrypted according to a preset encryption algorithm to obtain encrypted sample data.

7. The method according to claim 6, after obtaining the sample data stored in the participating nodes, further comprising: performing standardization processing on the sample data to obtain standardized sample data; The step of encrypting the sample data according to a preset encryption algorithm to obtain encrypted sample data includes: The standardized sample data is encrypted according to a preset encryption algorithm to obtain encrypted sample data.

8. The method according to any one of claims 1 to 5, further comprising: after sending the updated model gradient to the central node: receiving updated model parameters of the second GRU model sent by the central node, wherein the updated model parameters are obtained by the central node after updating the second GRU model according to the updated model gradient; Adjusting the model parameters of the first GRU model according to the updated model parameters of the second GRU model to obtain an adjusted first GRU model; Continue to perform the steps of determining the encrypted sample data, splitting the encrypted sample data, and obtaining a plurality of encrypted sample data fragments until a second GRU model that meets the training stop condition is obtained.

9. The method according to any one of claims 1 to 5, wherein obtaining the target noise vector comprises: The intensity of the target noise vector is adjusted according to the data sensitivity of the encrypted sample data and the encrypted label data fragment to obtain an adjusted target noise vector.

10. A longitudinal federated GRU model training method based on secret sharing and differential privacy, applied to the central node, including: Determining encrypted label data, and splitting the encrypted label data to obtain a plurality of encrypted label data fragments; Sending the multiple encrypted label data slices to participating nodes, and receiving updated model gradients of a first GRU model sent by the participating nodes, wherein the first GRU model is deployed on the participating nodes, and the updated model gradients are obtained by adding a target noise vector to an initial model gradient, and the initial model gradients are calculated based on the multiple encrypted label data slices; The second GRU model is updated according to the updated model gradient to obtain an updated second GRU model, and the updated model parameters of the updated second GRU model are sent to the participating nodes, wherein the second GRU model is deployed at the central node.

11. The method according to claim 10, wherein there are multiple participating nodes; The updating of the second GRU model according to the updated model gradient to obtain an updated second GRU model includes: Aggregate the updated model gradients sent by multiple participating nodes to obtain the target model gradient; The second GRU model is updated according to the target model gradient to obtain an updated second GRU model.

12. A vertical federated GRU model training system based on secret sharing and differential privacy, including participating nodes and central nodes, wherein: The central node is configured to determine the encrypted label data, split the encrypted label data to obtain a plurality of encrypted label data fragments, and send the plurality of encrypted label data fragments to the participating nodes; The participating node is configured to determine encrypted sample data, split the encrypted sample data to obtain multiple encrypted sample data fragments; input the multiple encrypted sample data fragments into a first GRU model to obtain prediction result fragments corresponding to the multiple encrypted sample data fragments; Calculating the initial model gradient of the first GRU model according to the prediction result slice and the encrypted label data slice sent by the central node; Obtaining a target noise vector, adding the target noise vector to the initial model gradient, obtaining an updated model gradient, and sending the updated model gradient to the central node, wherein the first GRU model is deployed on the participating node; The central node is configured to update the second GRU model according to the updated model gradient, obtain an updated second GRU model, and send the updated model parameters of the updated second GRU model to the participating nodes, wherein the second GRU model is deployed at the central node.

13. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Enterprise electricity fee payment risk prediction method and system based on longitudinal federated logistic regression

    CN115392531A

  • Longitudinal federal learning user credit scoring method based on differential privacy

    CN117273901A

  • Label sharing-based longitudinal federated learning differential privacy protection method and system

    CN117579215A

  • Method and device for testing privacy calculation part of model reasoning in privacy calculation product

    CN119357050A

  • Vehicle data prediction method based on GRU and differential privacy federated learning

    CN119397593A