Data processing method and system for large-scale carbon data
By combining homomorphic encryption and secret sharing technology among multiple participants, efficient computing and secure sharing of large-scale carbon data processing are achieved, and the problems of inefficient computing efficiency and large communication overhead in the existing technology are solved.
Patent Information
- Application Number
- CN202510168907.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing homomorphic encryption and secret sharing technologies are inefficient in processing large-scale data and complex computing scenarios, and are difficult to meet actual needs.
By combining homomorphic encryption and secret sharing technologies among multiple participants, each participant updates the gradient information of the local carbon model separately, and restores the global encryption gradient through secret sharing and updates the global carbon model parameters.
This method improves the efficiency of large-scale carbon data processing, reduces communication overhead, and ensures safe participation of all participants in global computing.
Smart Images

Figure CN120017368A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and in particular to a data processing method and system for large-scale carbon data. Background Art
[0002] In recent years, with the increasing demand for data privacy protection, secure data sharing and distributed computing have become important research topics. Homomorphic encryption and secret sharing technologies, as key technologies, are widely used in scenarios to protect data privacy. Homomorphic encryption allows calculations to be performed directly on encrypted data without decryption, thereby avoiding the risk of data leakage during sharing and computing. Secret sharing greatly enhances data security in multi-party computing scenarios by dividing data into multiple shares and restoring the original data only when a certain threshold is reached. These technologies provide strong technical support for privacy protection and multi-party collaboration.
[0003] Existing homomorphic encryption technologies, such as Paillier encryption, support additive homomorphism and have been applied to scenarios such as federated learning. However, as the amount of data and computational complexity increase, the computational overhead and data expansion problems of homomorphic encryption gradually emerge, especially when processing complex operations such as matrix operations, where inefficiency becomes a bottleneck for applications.
[0004] Secret sharing technology divides secret data into several shares. Only when a certain number of shares are reached can the original data be restored. This technology is widely used in multi-party collaboration scenarios to ensure that the data of each party will not be leaked individually in multi-party computing. Although secret sharing has a wide range of application scenarios in multi-party secure computing, its efficiency problems are gradually exposed when processing high-dimensional sparse data, especially in large-scale data computing scenarios. The high computing and communication overhead limits its application in practical scenarios.
[0005] As an emerging privacy protection technology, federated learning allows each participant to train the model locally and aggregate the model update parameters to the central server, thus avoiding the leakage of original data. According to the distribution characteristics of data, federated learning can be divided into horizontal and vertical types. Vertical federated learning is suitable for data scenarios with complementary feature dimensions, especially for medical and financial fields. However, when traditional federated learning involves large-scale data processing, especially when combined with homomorphic encryption technology, it has high computing and communication costs and low processing efficiency, making it difficult to meet actual needs.
[0006] In short, with the development of homomorphic encryption, secret sharing technology, and parallel computing technology, these technologies are gradually integrated and applied to privacy-preserving computing and multi-party collaboration scenarios. By further optimizing these technologies, the problems of low computing efficiency and high communication overhead in existing technologies are solved, making it possible to apply them in large-scale data and complex computing scenarios. Summary of the invention
[0007] The present invention provides a data processing method and system for large-scale carbon data, which can solve at least one of the above technical problems.
[0008] According to one aspect of the present invention, a data processing method for large-scale carbon data is provided, comprising: Each of the multiple participants updates the local carbon model of each participant based on the carbon characteristic data of each participant to obtain the local model gradient of each participant; Each participant encrypts the local model gradient of each participant based on the homomorphic encryption public key of each participant to obtain the local encrypted model gradient of each participant; Each participant uses a secret sharing algorithm to split each participant's local encrypted model gradient into multiple encrypted sub-gradients and distribute them to other participants; A first participant among the multiple participants performs secret sharing recovery on the encrypted sub-gradient of the first participant and the encrypted sub-gradients from other participants to obtain a global encrypted gradient; The first participant updates a first global encrypted model parameter of the global carbon model based on the global encrypted gradient to obtain a second global encrypted model parameter of the global carbon model, and sends the second global encrypted model parameter to other participants; Each participant updates its local carbon model based on the second global encrypted model parameters.
[0009] According to another aspect of the present invention, there is provided a data processing system for large-scale carbon data, comprising a plurality of participants, wherein the plurality of participants include a first participant; Each of the participants is used to perform the following operations: Based on the carbon characteristic data of the participants, respectively update the local carbon models of the participants to obtain the local model gradients of the participants; Based on the homomorphic encryption public key of the participant, encrypt the local model gradients of the participant respectively to obtain the local encrypted model gradients of the participant; Using a secret sharing algorithm, the local encrypted model gradient of the participant is split into multiple encrypted sub-gradients and distributed to other participants; The first party is used to perform the following operations: Perform secret sharing recovery on the encrypted sub-gradient of the first participant and the encrypted sub-gradients from other participants to obtain a global encrypted gradient; Based on the global encrypted gradient, a first global encrypted model parameter of the global carbon model is updated to obtain a second global encrypted model parameter of the global carbon model, and the second global encrypted model parameter is sent to other participants; Each participant is also used to update the local carbon model of the participant based on the second global encrypted model parameters.
[0010] By adopting the technical solution of the present invention, each participant homomorphically encrypts the gradient information of the local carbon model and divides the local encrypted gradient by secret sharing and distributes it to other participants. In this way, one of the participants performs secret sharing recovery on the local encrypted sub-gradients provided by each participant to obtain the global encrypted gradient. Then, the global carbon model is updated using the global encrypted gradient to obtain the updated global encrypted model parameters, and the updated global encrypted model parameters are sent to each participant. Each participant can use the updated global encrypted model parameters to update the local carbon model. In this way, each participant can be guaranteed to participate in the global calculation safely.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention. Figure 1 is a flow chart of a data processing method for large-scale carbon data according to an embodiment of the present invention; Figure 2 is a flow chart of a data processing method for large-scale carbon data according to another embodiment of the present invention; Figure 3 is a structural block diagram of a data processing system for large-scale carbon data according to an embodiment of the present invention; Figure 4 The block diagram is a block diagram of an electronic device for implementing the method according to the embodiment of the present invention. DETAILED DESCRIPTION
[0013] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0014] Figure 1It is a flow chart of a data processing method for large-scale carbon data according to an embodiment of the present invention.
[0015] like Figure 1 As shown, the data processing method for large-scale carbon data may include: S110, each of the multiple participants updates the local carbon model of each participant based on the carbon characteristic data of each participant to obtain the local model gradient of each participant; S120, each participant encrypts the local model gradient of each participant based on the homomorphic encryption public key of each participant to obtain the local encrypted model gradient of each participant; S130, each participant uses a secret sharing algorithm to split each participant's local encrypted model gradient into multiple encrypted sub-gradients and distributes them to other participants; S140, a first participant among the multiple participants performs secret sharing recovery on the encrypted sub-gradient of the first participant and the encrypted sub-gradients from other participants to obtain a global encrypted gradient; S150, the first participant updates the first global encrypted model parameter of the global carbon model based on the global encrypted gradient to obtain the second global encrypted model parameter of the global carbon model, and sends the second global encrypted model parameter to other participants; S160, each participant updates its local carbon model based on the second global encrypted model parameters.
[0016] Exemplarily, the above steps S110 to S160 may be iterated multiple times to obtain the final local carbon model of each participant.
[0017] By way of example, the parties involved may be power companies, power stations, regulatory agencies, etc.
[0018] Exemplarily, each participant has its own key pair, including a public key and a private key. For example, if there are three participants, Party A, Party B, and Party C, the public keys are pk_A, pk_B, pk_C, and the private keys are sk_A, sk_B, sk_C. These keys can be pre-generated and stored in each participant.
[0019] Exemplarily, carbon data is preprocessed and chunked to support massively parallel computing and distributed collaboration.
[0020] For example, in the data preprocessing stage, the carbon feature vector is first explicitly divided into blocks. The carbon feature data held by each participant contains multiple time series, such as records of changes in electricity consumption in a certain area or equipment operating status. These data are represented in the form of columns, and each column represents a different time point or feature. The carbon feature data is strictly divided into blocks according to the time dimension. Specifically, the data is divided into blocks of 30 days, and each data block contains electricity consumption records at 30 time points. This block division method not only optimizes subsequent distributed computing, but also significantly reduces the size of a single data block during data transmission, thereby improving transmission efficiency.
[0021] In this example, each data block contains a relatively small amount of data, avoiding the communication bottleneck caused by large-scale data in a single transmission. The block-based data can more efficiently adapt to the needs of parallel computing, and each block can be assigned to different computing nodes for processing, thereby greatly improving the parallel processing capabilities of the system. For scenarios with high real-time requirements in carbon scenario operation, the size of the data block has been optimized according to the system computing resources and bandwidth conditions to ensure maximum system efficiency.
[0022] Exemplarily, after completing the segmentation of the feature vector, these segments can be further packaged for subsequent homomorphic encryption and collaborative computing. Packaging is the process of merging the feature data in multiple segments into a larger data block. This reduces the number of communications during transmission and reduces the computational overhead of each encryption operation. For example, in each 30-day segment, the power consumption data at these time points are packaged into a vector of 30 time points. Through this packaging method, multiple feature data can be merged into larger data blocks, which not only reduces the communication frequency in the system, but also effectively reduces the data expansion problem during encryption calculations, reducing the burden of encryption processing.
[0023] For example, data encryption and secret sharing in this example are the core links to ensure data security monitoring and data privacy protection. In order to solve the hidden dangers of data leakage in the process of cross-domain data sharing and ensure the data information security of each participant, this example adopts the combination of Paillier homomorphic encryption and Shamir's secret sharing scheme to achieve data encryption calculation and secure sharing.
[0024] First, after data preprocessing and packaging, all block data will be encrypted by each participant. Each participant uses Paillier homomorphic encryption technology to encrypt the data based on the feature data set it holds locally.
[0025] In the encryption process of the local model gradient, the participants generate a public key (n, g) and a private key (λ, μ), where n = pq, p and q are large prime numbers. The public key is used to encrypt the plaintext, and the private key is used to decrypt it. In this example, the participants use their block-based feature data mA as input and use the Paillier encryption algorithm to generate the ciphertext cA, which is calculated as follows: Among them, r is a random number used to ensure the randomness and security of the ciphertext. Through this encryption process, the privacy of data in calculation and transmission is ensured, so that in the subsequent collaborative computing process, each participant only operates on the encrypted data without directly exposing the original data.
[0026] After the local model gradient encryption is completed, each participant uses the multi-party secure computing framework of this example to perform cross-domain data collaboration and calculation. In this process, the additive homomorphism of Paillier homomorphic encryption allows the participants to perform direct addition operations on the encrypted data. For example, for two encrypted values cA = E(mA) and cB = E(mB), the addition operation can be performed by the following formula: c A ·c B =E(m A +m B ) This feature enables encrypted data to participate in collaborative computing without decryption. In addition, Paillier encryption also supports multiplication of encrypted data with constants, such as multiplication with constant k by exponential operation cAK on ciphertext cA, as shown in the following formula: In the process of collaborative computing among multiple participants, in order to further enhance data security and ensure the correctness of the calculation, this example introduces Shamir's secret sharing scheme. When the participants calculate the intermediate encryption results, that is, the local encrypted gradients, the system shards these encryption results to obtain multiple local encrypted sub-gradients. The participants use their encrypted calculation results, that is, the local encrypted gradients, as the secret cA, and construct a random polynomial f(x) to split the secret: f(x)=c A +a1x+a2x 2 +…+a t-1 x t-1 Among them, a1, a2, …, a t-1is a random coefficient, and t is the recovery threshold. Each participant Pi will get a share si=f(i), which is securely distributed through a secret sharing mechanism. Each participant only holds a share of part of the data. Even if the data of a participant is leaked, the original data cannot be derived from the share alone. Only when at least t participants combine their shares together can the original ciphertext cA be restored by the Lagrange interpolation method, as follows: In collaborative computing, each participant performs the above split calculation on its local encrypted gradient to obtain the intermediate calculation results, namely multiple local encrypted sub-gradients. Subsequently, these local encrypted sub-gradients are distributed to other participants using the Shamir secret sharing scheme to ensure that all parties can safely participate in the global computation. In this way, this example ensures that even if the data of one or more participants is leaked, the complete encryption result and original data cannot be recovered.
[0027] The following introduces the above data processing method by taking the system including three parties as an example.
[0028] Exemplarily, the three participants A, B and C each load their locally held carbon characteristic data. For example, the carbon characteristic data of party A is X_A, the carbon characteristic data of party B is X_B, and the carbon characteristic data of party C is X_C.
[0029] Exemplarily, the participants update the local carbon model based on the local carbon characteristic data to obtain the local model gradient of the local carbon model. For example, the local model gradient of party A is w_A=f(X_A,w_local), the local model gradient of party B is w_B=f(X_B,w_local), and the local model gradient of party C is w_C=f(X_C,w_local).
[0030] Exemplarily, the participant encrypts the local model gradient based on the local public key.
[0031] Party A encrypts the calculated gradient w_A, for example: E(w_A) = encrypted(w_A,pk_A).
[0032] Party B and Party C also encrypt their respective gradients w_B and w_C, for example: E(w_B) = encrypted(w_B, pk_B), E(w_C) = encrypted(w_C, pk_C).
[0033] Exemplarily, the participants use Shamir secret sharing to split the local encrypted gradient into multiple shares, that is, to obtain multiple local encrypted sub-gradients.
[0034] Party A divides E(w_A) into multiple shares and distributes them to other parties. For example, share_A = split(E(w_A)). Party B and Party C also split their respective encrypted gradients. For example, share_B = split(E(w_B)), share_C = split(E(w_C)). The splitting algorithm can use the above-mentioned random polynomial f(x) for splitting.
[0035] Exemplarily, after each participant receives the gradient share of other participants, Shamir secret sharing is used to restore the gradient share of each participant to obtain the first global encrypted gradient E(w_global).
[0036] For example, taking Party A as responsible for recovery: E(w_global)=recovery(share_A, share_B, share_C). The recovery algorithm may use the above-mentioned Lagrange interpolation method for recovery.
[0037] Exemplarily, for the global carbon model, in the encrypted state, homomorphically encrypted gradient updates can be performed as follows: the second global encrypted gradient is E(w_global_new)=E(w_global)-η*E(w_global), where η is the learning rate.
[0038] Exemplarily, after the second global encrypted gradient is updated through homomorphic encryption, this gradient can be distributed to other participants.
[0039] Exemplarily, each participant may use their own private key to decrypt the second global encrypted gradient to obtain a global gradient, and then use the global gradient to update the local carbon model.
[0040] Exemplarily, the model parameters of the local carbon model of party A are updated as: w_global_new=decrypted(E(w_global_new),sk_A). The model parameters of the local carbon model of party B are updated as w_global_new=decrypted(E(w_global_new),sk_B). The model parameters of the local carbon model of party C are updated as w_global_new=decrypted(E(w_global_new),sk_C).
[0041] According to the above implementation, each participant homomorphically encrypts the gradient information of the local carbon model and divides the local encrypted gradient by secret sharing and distributes it to other participants. In this way, one of the participants performs secret sharing recovery on the local encrypted sub-gradients provided by each participant to obtain the global encrypted gradient. Then, the global encrypted gradient is used to update the global carbon model to obtain the updated global encrypted model parameters, and the updated global encrypted model parameters are sent to each participant. Each participant can use the updated global encrypted model parameters to update the local carbon model. In this way, it is not necessary to use the original carbon data of other participants to calculate the local model, and it can ensure that each participant participates in the global model calculation safely.
[0042] In one embodiment, it also includes: each participant encodes the carbon sample data of each participant respectively to obtain the first matrix characteristic data corresponding to the carbon sample data of each participant; each participant diagonally encodes the first matrix characteristic data corresponding to each participant to obtain the second matrix characteristic data; each participant determines its own carbon characteristic data based on the second matrix characteristic data obtained by each participant.
[0043] For sparse matrices, there are usually a large number of zero elements, and the calculation of these elements is unnecessary. Through diagonal encoding, the non-zero elements in the matrix are concentrated near the main diagonal for storage and calculation, thereby optimizing the calculation process and reducing invalid calculations.
[0044] For example, for matrix A, the details are as follows: After diagonal encoding of matrix A, the matrix is obtained as follows:
[0045] Exemplarily, the second matrix characteristic data is used as the carbon characteristic data.
[0046] Exemplarily, the zero elements in the second matrix characteristic data are compressed to obtain the carbon characteristic data.
[0047] According to the above implementation, through diagonal coding, the non-zero elements in the matrix are concentrated near the main diagonal for storage and calculation, thereby optimizing the calculation process and reducing invalid calculations, thereby improving the efficiency of subsequent model calculations.
[0048] In one embodiment, each participant determines its own carbon characteristic data based on the second matrix characteristic data obtained by each participant, including: the participants extract non-zero elements from the second matrix characteristic data to obtain a first characteristic array; the participants determine the values of corresponding elements in the second characteristic array based on the column index of each element in the first characteristic array in the second matrix characteristic data; the participants determine the values of corresponding elements in the third characteristic array based on the row index of each element in the first characteristic array in the second matrix characteristic data; the participants determine the carbon characteristic data of the participants based on the first characteristic array, the second characteristic array and the third characteristic array.
[0049] For example, by caching intermediate calculation results, repeated calculations can be avoided. In each round of training, if the intermediate results have been calculated and have not changed, these results can be directly reused to avoid repeated operations. This mechanism is particularly important in multiple rounds of iterative calculations, which can significantly reduce the amount of calculations, especially in distributed computing, effectively reducing the communication and computing overhead between nodes.
[0050] For example, in traditional matrix multiplication, especially for high-dimensional sparse data, the presence of zero elements leads to a large number of redundant calculations, which not only wastes computing resources but also increases storage requirements. To address this problem, this example compresses the sparse matrix storage format, and the zero elements in the matrix are not stored, but only the non-zero elements and their corresponding index information are saved, thereby reducing storage space and computational complexity.
[0051] For example, for the following matrix X A : For example, from the matrix X A Extract the non-zero elements, so that the first characteristic array formed is V1 = [a, b, c, d].
[0052] Exemplarily, the second feature array is V2 = [2, 3, 1, 4], and the third feature array is V3 = [1, 2, 3, 4].
[0053] According to the above implementation, only non-zero elements are concerned during the calculation process, which avoids invalid calculations caused by zero elements and significantly improves the efficiency of matrix operations. Especially in scenarios where the matrix is large and sparse, this optimization method can reduce unnecessary calculations, thereby improving the overall calculation speed. At the same time, since each participant only needs to transmit non-zero elements and their index information, this reduces the communication overhead between participants in a distributed computing environment and reduces the consumption of network bandwidth, thereby optimizing the data exchange process and making the communication of the entire distributed system more efficient.
[0054] In one embodiment, the first participant updates the first global encrypted model parameter of the global carbon model based on the global encrypted gradient to obtain the second global encrypted model parameter, including: the first participant subtracts the product of the learning rate of the global carbon model and the global encrypted gradient from the first global encrypted model parameter of the global carbon model in the first participant to obtain the second global encrypted model parameter.
[0055] Exemplarily, the following formula is used to update the global model parameters.
[0056] E(w_global_new)=E(w_global)-η*E(w_global); Among them, E(w_global) is the first global encryption model parameter, E(w_global) is the global encryption gradient, η is the learning rate, and E(w_global_new) is the second global encryption model parameter.
[0057] Exemplarily, the global carbon model may be a local carbon model of the first party.
[0058] Exemplarily, the first global encrypted model parameters of the global carbon model may be: based on the carbon characteristic data of the first participant, updating the local carbon model of the first participant to obtain the model parameters of the updated local carbon model, encrypting the model parameters using the public key of the first participant to obtain the local encrypted model parameters of the first participant, and using the local encrypted model parameters as the first global encrypted model parameters of the global carbon model.
[0059] According to the above implementation, the encrypted gradient can be restored by one of the multiple participants to obtain a global encrypted gradient, and then the global carbon model can be updated using the global encrypted gradient to obtain updated global encrypted model parameters without decryption.
[0060] In one embodiment, based on the second global encrypted model parameters, the local carbon models of each participant are updated respectively, including: each participant decrypts the second global encrypted model parameters based on the homomorphic encryption private key of each participant to obtain the global model parameters; each participant updates the local carbon models of each participant based on the global model parameters.
[0061] It can be understood that the model parameters of the updated local carbon model are the global model parameters.
[0062] It can be understood that the local carbon models of each participant are finally obtained as the global carbon model, and their model parameters are the same.
[0063] According to the above implementation, each participant can train the model without using the original training data of other participants.
[0064] Figure 2 It is a flow chart of a data processing method for large-scale carbon data according to an embodiment of the present invention.
[0065] like Figure 2 As shown in the figure, assume that three participants (A, B, C) have carbon data from different regions. Since the data involves sensitive information, the participants do not want to share the original data directly, but still want to conduct collaborative analysis and model training without leaking privacy. Each participant can use the following steps to achieve secure collaborative computing.
[0066] Data preprocessing and packaging: Each participant preprocesses the carbon data held locally, and processes the time series data in blocks of 30 days. The processed feature data is combined into larger data blocks through packaging operations to reduce the computational burden during transmission and encryption.
[0067] Data encryption: After preprocessing, the three participants A, B, and C use Paillier homomorphic encryption to encrypt their packaged data blocks to ensure that the data is calculated in an encrypted state. Each participant generates its own public key and private key, performs encryption operations on the block data, and generates ciphertext.
[0068] Parallel processing and distributed computing: Each participant distributes the encrypted data blocks to multiple computing nodes and performs parallel processing through the MapReduce framework. In the Map phase, each computing node calculates the data blocks it receives and performs local gradient calculations and encryption operations. In the Reduce phase, the calculation results of each node are summarized to form a global gradient for global model updates.
[0069] Shamir secret sharing: During the calculation process, all participants share the encrypted intermediate calculation results in secret. Through Shamir secret sharing technology, the encrypted results are sharded and securely distributed to other participants. Only after collecting a sufficient number of shares can the parties restore the complete encrypted results, thus effectively ensuring the security of the data.
[0070] Global model update: The global gradients summarized in the Reduce phase are still encrypted. The global model parameters are updated directly in the encrypted state through the addition and multiplication characteristics of homomorphic encryption. After the model update is completed, each participant decrypts the new model parameters and synchronizes them to the local model to prepare for the next round of iteration.
[0071] Matrix optimization: During the collaborative computing process, participants use diagonal coding and lazy rotation and summation mechanisms to optimize sparse matrices, reducing invalid computing operations and improving the efficiency of large-scale matrix operations.
[0072] Iterative calculation and model training: Through multiple rounds of iterations, the system gradually trains a global model suitable for carbon data analysis. In each round of iteration, participants use local data to participate in the calculation while maintaining complete protection of data privacy. Through the MapReduce framework and parallel computing technology, the system can maintain high processing efficiency and computing performance in a large-scale data environment.
[0073] Figure 3 The invention is a data processing system for large-scale carbon data according to an embodiment of the present invention.
[0074] like Figure 3 As shown, the data processing system is characterized by comprising a plurality of participants (310 to 31N), wherein the plurality of participants include a first participant; Each of the participants is used to perform the following operations: Based on the carbon characteristic data of the participants, respectively update the local carbon models of the participants to obtain the local model gradients of the participants; Based on the homomorphic encryption public key of the participant, encrypt the local model gradients of the participant respectively to obtain the local encrypted model gradients of the participant; Using a secret sharing algorithm, the local encrypted model gradient of the participant is split into multiple encrypted sub-gradients and distributed to other participants; The first party is used to perform the following operations: Perform secret sharing recovery on the encrypted sub-gradient of the first participant and the encrypted sub-gradients from other participants to obtain a global encrypted gradient; Based on the global encrypted gradient, a first global encrypted model parameter of the global carbon model is updated to obtain a second global encrypted model parameter of the global carbon model, and the second global encrypted model parameter is sent to other participants; Each participant is also used to update the local carbon model of the participant based on the second global encrypted model parameters.
[0075] In one implementation, each participant is further configured to: Encoding the carbon sample data of the participant to obtain corresponding first matrix feature data; Performing diagonal encoding on the first matrix feature data to obtain second matrix feature data; Based on the second matrix characteristic data, the carbon characteristic data of the participant is determined.
[0076] In one embodiment, determining the carbon characteristic data of the participant based on the second matrix characteristic data includes: Extracting non-zero elements from the second matrix feature data to obtain a first feature array; Determine the value of the corresponding element in the second feature array based on the column index of each element in the first feature array in the second matrix feature data; Determine the value of the corresponding element in the third feature array based on the row index of each element in the first feature array in the second matrix feature data; Based on the first feature array, the second feature array and the third feature array, the carbon feature data of the participant is determined.
[0077] In one implementation, the first participant is used to update the first global encrypted model parameter of the global carbon model based on the global encrypted gradient to obtain the second global encrypted model parameter, specifically: The first participant is used to obtain the second global encrypted model parameter by subtracting the product of the learning rate of the global carbon model and the global encrypted gradient from the first global encrypted model parameter of the global carbon model in the first participant.
[0078] In one embodiment, the updating of the local carbon model of the participant based on the second global encrypted model parameter includes: Decrypting the second global encrypted model parameter based on the homomorphic encryption private key of the participant to obtain a global model parameter; Based on the global model parameters, the local carbon model of the participant is updated.
[0079] For the description of specific functions and examples of each module and submodule of the system in the embodiment of the present invention, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0080] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0081] According to an embodiment of the present invention, the present invention also provides a system and a readable storage medium.
[0082] Figure 4A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0083] like Figure 4 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0084] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0085] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a data processing method for large-scale carbon data. For example, in some embodiments, the data processing method for large-scale carbon data may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method for large-scale carbon data described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the data processing method for large-scale carbon data in any other appropriate manner (for example, by means of firmware).
[0086] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0087] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0088] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0089] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0090] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0091] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0092] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this document does not limit this.
[0093] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A data processing method for large-scale carbon data, characterized in that: include: Each of the multiple participants updates the local carbon model of each participant based on the carbon characteristic data of each participant to obtain the local model gradient of each participant; Each participant encrypts the local model gradient of each participant based on the homomorphic encryption public key of each participant to obtain the local encrypted model gradient of each participant; Each participant uses a secret sharing algorithm to split each participant's local encrypted model gradient into multiple encrypted sub-gradients and distribute them to other participants; A first participant among the multiple participants performs secret sharing recovery on the encrypted sub-gradient of the first participant and the encrypted sub-gradients from other participants to obtain a global encrypted gradient; The first participant updates the first global encrypted model parameter of the global carbon model based on the global encrypted gradient to obtain the second global encrypted model parameter of the global carbon model, and sends the second global encrypted model parameter to other participants; Each participant updates its local carbon model based on the second global encrypted model parameters.
2. The method according to claim 1, characterized in that Also includes: Each participant encodes the carbon sample data of each participant respectively to obtain first matrix feature data corresponding to the carbon sample data of each participant; Each participant performs diagonal encoding on the first matrix characteristic data corresponding to each participant to obtain second matrix characteristic data; each participant determines the carbon characteristic data based on the second matrix characteristic data obtained by each participant.
3. The method according to claim 2, characterized in that The respective participants determine their respective carbon characteristic data based on the respective obtained second matrix characteristic data, including: The participants respectively extract non-zero elements from the second matrix feature data to obtain a first feature array; The participant determines the value of the corresponding element in the second feature array based on the column index of each element in the first feature array in the second matrix feature data; The participant determines the value of the corresponding element in the third feature array based on the row index of each element in the first feature array in the second matrix feature data; The participant determines the carbon characteristic data of the participant based on the first characteristic array, the second characteristic array and the third characteristic array.
4. The method according to claim 1, characterized in that: The first participant updates the first global encrypted model parameter of the global carbon model based on the global encrypted gradient to obtain the second global encrypted model parameter, including: the first participant subtracts the product of the learning rate of the global carbon model and the global encrypted gradient from the first global encrypted model parameter of the global carbon model in the first participant to obtain the second global encrypted model parameter.
5. The method according to claim 1, characterized in that The updating of the local carbon model of each participant based on the second global encrypted model parameter includes: Each participant decrypts the second global encrypted model parameter based on the homomorphic encryption private key of each participant to obtain a global model parameter; Each participant updates its local carbon model based on the global model parameters.
6. A data processing system for large-scale carbon data, characterized in that: comprising a plurality of participants, wherein the plurality of participants include a first participant; Each of the participants is used to perform the following operations: Based on the carbon characteristic data of the participants, respectively update the local carbon models of the participants to obtain the local model gradients of the participants; Based on the homomorphic encryption public key of the participant, encrypt the local model gradients of the participant respectively to obtain the local encrypted model gradients of the participant; Using a secret sharing algorithm, the local encrypted model gradient of the participant is split into multiple encrypted sub-gradients and distributed to other participants; The first party is used to perform the following operations: Perform secret sharing recovery on the encrypted sub-gradient of the first participant and the encrypted sub-gradients from other participants to obtain a global encrypted gradient; Based on the global encrypted gradient, a first global encrypted model parameter of the global carbon model is updated to obtain a second global encrypted model parameter of the global carbon model, and the second global encrypted model parameter is sent to other participants; Each participant is also used to update the local carbon model of the participant based on the second global encrypted model parameters.
7. The system according to claim 6, characterized in that Each participant is also used to: Encoding the carbon sample data of the participant to obtain corresponding first matrix feature data; Performing diagonal encoding on the first matrix feature data to obtain second matrix feature data; Based on the second matrix characteristic data, the carbon characteristic data of the participant is determined.
8. The system according to claim 7, characterized in that Determining the carbon characteristic data of the participant based on the second matrix characteristic data includes: Extracting non-zero elements from the second matrix feature data to obtain a first feature array; Determine the value of the corresponding element in the second feature array based on the column index of each element in the first feature array in the second matrix feature data; Determine the value of the corresponding element in the third feature array based on the row index of each element in the first feature array in the second matrix feature data; Based on the first feature array, the second feature array and the third feature array, the carbon feature data of the participant is determined.
9. The system according to claim 6, characterized in that The first participant is used to update the first global encrypted model parameter of the global carbon model based on the global encrypted gradient to obtain the second global encrypted model parameter, which is specifically: The first participant is used to obtain the second global encrypted model parameter by subtracting the product of the learning rate of the global carbon model and the global encrypted gradient from the first global encrypted model parameter of the global carbon model in the first participant.
10. The system according to claim 6, characterized in that The updating of the local carbon model of the participant based on the second global encrypted model parameter includes: Decrypting the second global encrypted model parameter based on the homomorphic encryption private key of the participant to obtain a global model parameter; Based on the global model parameters, the local carbon model of the participant is updated.