Multi-party data processing method, system, electronic device and storage medium

By using data cutting functions and computational aggregation functions in multi-party computing, large-scale data is divided and allocated to multiple computing nodes for processing, the existing multi-party computing problem is solved and fast and efficient computing processing is achieved.

CN114296922BActive Publication Date: 2025-05-16HANGZHOU QULIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111631336.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-05-16
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

The existing multi-party computing technology is inefficient when processing large-scale data and has limited computing power, resulting in slow computing speed and increasing memory cannot effectively solve the problem.

Method used

The initiator and the participant respectively obtain their respective data sets and calculation operators, use the data cutting function to segment the data sets, allocate them to multiple calculation slave nodes for calculation logic execution, the initiator calculates the slave node summary results and sends them to the initiator to calculate the master node for data aggregation.

Benefits of technology

It improves the efficiency of multi-party computing, can avoid single-machine computing power limitations when processing large-scale data, and realizes rapid computing processing and data aggregation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114296922B_ABST
    Figure CN114296922B_ABST
Patent Text Reader

Abstract

Multi-party data processing method, system, electronic device and storage medium. The present application relates to a multi-party data processing method, wherein the multi-party data processing method comprises: obtaining respective data sets and computing operators through the initiator and the participants respectively; segmenting the data sets according to the data cutting function in the computing operator to obtain sub-data sets, and distributing the sub-data sets to each computing slave node of the party, and each computing slave node executes the corresponding computing logic according to the computing operator; the computing slave node of the initiator obtains the computing result according to the corresponding computing data provided by the computing slave nodes of each participant, and sends the computing result to the computing master node of the initiator; the computing master node of the initiator aggregates the computing result according to the computing operator to obtain aggregated data, thereby solving the problem of low efficiency of multi-party computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a multi-party data processing method, system, electronic device and storage medium. Background Art

[0002] Secure multi-party computation is about how to securely compute an agreed function without a trusted third party. In secure multi-party computation, each participant can complete the computation of an agreed function that requires the original data of multiple participants without revealing their original data to the other party or third party. Existing open source privacy computing frameworks such as SPDZ and ABY3 all use the computing strategy of loading data into memory and performing calculations directly. When the amount of data to be calculated increases, the only way to support it is to increase the memory. However, when the amount of data reaches TB or even PB, increasing the memory can no longer solve the problem. In addition, due to the limited computing power of a single machine, the computing speed will be very slow when the amount of data is too large.

[0003] Currently, no effective solution has been proposed to address the problem of low efficiency of multi-party computing in related technologies. Summary of the invention

[0004] The embodiments of the present application provide a multi-party data processing method, system, device and storage medium to at least solve the problem of low multi-party computing efficiency in the related art.

[0005] In a first aspect, an embodiment of the present application provides a multi-party data processing method, including:

[0006] The initiator and participants obtain their own data sets and computing operators respectively;

[0007] The data set is segmented according to the data cutting function in the computing operator to obtain sub-data sets, and the sub-data sets are distributed to each computing slave node of the party, and each computing slave node executes the corresponding computing logic according to the computing operator;

[0008] The initiator's computing slave node obtains a computing result based on the corresponding computing data provided by each participant's computing slave node, and sends the computing result to the initiator's computing master node;

[0009] The initiator computing master node aggregates the computing results according to the computing operator to obtain aggregated data.

[0010] In some embodiments, the method further comprises:

[0011] According to the data cutting function, the data set is segmented to obtain the sub-data sets and the sub-data set numbers are marked, and the sub-data sets are distributed to each computing slave node of the party, and the computing slave node of the initiator and the computing slave node of the participant having the same sub-data set number execute the corresponding computing logic;

[0012] The initiator's computing slave node obtains a computing result based on the corresponding computing data provided by each participant's computing slave node, and sends the computing result to the initiator's computing master node.

[0013] In some embodiments, the initiator computing slave node obtains the computing result according to the corresponding computing data provided by the computing slave nodes of each participant, including:

[0014] The participant computing slave node sends the computing data and the corresponding sub-data set number to the participant computing master node.

[0015] The participating party's computing master node sends the computing data and the corresponding sub-data set number to the initiating party's computing master node. The initiating party's computing master node sends the computing data to the corresponding initiating party's computing slave node according to the sub-data set number. The initiating party's computing slave node obtains the computing result according to the corresponding computing data provided by each participating party's computing slave node.

[0016] In some embodiments, the data segmentation of the data set to obtain the sub-data sets and labeling the sub-data sets according to the data segmentation function in the calculation operator includes:

[0017] When the data set is a vector data set, it is cut in segments, and the sub-data set number is the corresponding segment number;

[0018] When the data set is a set-type data set, each element in the data set is read, and the bucket number is calculated by modulus after hash operation and written into the corresponding sub-data set file. The sub-data set is labeled as the bucket number.

[0019] In some embodiments, the initiator and the participant respectively obtain their own computing operators including:

[0020] The initiator obtains the computing operator, parses the configuration file in the computing operator, reads the algorithm type and version number therein, and performs overwriting if the computing operator of the same algorithm type already exists;

[0021] The initiator distributes the computing operator to each of the participants.

[0022] In some of the embodiments, the calculation operator includes independently set data cutting functions, data aggregation functions and message types of the calculation data.

[0023] In a second aspect, an embodiment of the present application provides a multi-party data processing system, including an initiator and a participant, wherein the initiator includes an initiator computing master node and an initiator computing slave node, and the participant includes a participant computing master node and a participant computing slave node:

[0024] The initiator and the participant respectively obtain their own data sets and computing operators, perform data segmentation on the data sets according to the computing operators to obtain sub-data sets, and distribute the sub-data sets to their own computing slave nodes, and each computing slave node executes corresponding computing logic according to the computing operators;

[0025] The initiator's computing slave node obtains the calculation result according to the corresponding calculation data provided by each participant's computing slave node, and sends the calculation result to the initiator's computing master node. The initiator's computing master node aggregates the calculation result according to the calculation operator to obtain aggregated data.

[0026] In some embodiments, the computing node includes a scheduler, the scheduler includes a virtual machine and a network component, wherein the computing node includes a computing master node and a computing slave node,

[0027] The virtual machine executes the multi-party data processing method according to any one of claims 1 to 7;

[0028] The network component is used for data communication between the computing nodes.

[0029] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the multi-party data processing method as described in the first aspect above is implemented.

[0030] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, and when the program is executed by a processor, the multi-party data processing method as described in the first aspect above is implemented.

[0031] Compared with the related art, the multi-party data processing method provided in the embodiment of the present application is that the initiator and the participating parties respectively obtain their own data sets and computing operators; the data sets are segmented according to the data cutting function in the computing operator to obtain sub-data sets, and the sub-data sets are distributed to each computing slave node of the party, and each computing slave node executes the corresponding computing logic according to the computing operator; the initiator's computing slave node obtains the calculation result according to the corresponding calculation data provided by the computing slave nodes of each participating party, and sends the calculation result to the initiator's computing master node; the initiator's computing master node aggregates the calculation result according to the computing operator to obtain aggregated data, thereby solving the problem of low efficiency of multi-party computing.

[0032] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0034] Figure 1 is a flowchart of a multi-party data processing method according to an embodiment of the present application;

[0035] Figure 2 is a flowchart of a multi-party data processing method according to another embodiment of the present application;

[0036] Figure 3 is a flowchart of another multi-party data interaction method according to an embodiment of the present application;

[0037] Figure 4 is a flowchart of another multi-party data interaction method according to an embodiment of the present application;

[0038] Figure 5 is a schematic diagram of a multi-party data processing method according to a preferred embodiment of the present application;

[0039] Figure 6 is a structural block diagram of a multi-party data processing system according to an embodiment of the present application;

[0040] Figure 7 It is a structural block diagram of a computing node according to an embodiment of the present application. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for ordinary technicians in the field related to the contents disclosed in the present application, some changes such as design, manufacturing or production based on the technical contents disclosed in the present application are only conventional technical means, and should not be understood as insufficient contents disclosed in the present application.

[0042] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0043] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantitative limitation, and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to greater than or equal to two. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, "A and / or B" can represent: A exists alone, A and B exist at the same time, and B exists alone. The terms "first", "second", "third" and the like involved in the present application are merely used to distinguish similar objects and do not represent a specific ordering of the objects.

[0044] This embodiment provides a multi-party data processing method. Figure 1is a flowchart of a multi-party data processing method according to an embodiment of the present application. Figure 1 As shown, the process includes the following steps:

[0045] Step S101, the initiator and the participant obtain their own data sets and computing operators respectively. The initiator and the participant of multi-party computing obtain the computing operators sent to each party respectively, and the computing operators contain the segmentation functions and computing logic of each party's data sets. In addition, the initiator and the participant will obtain the data sets uploaded by the party for computing.

[0046] Step S102, data segmentation is performed on the data set according to the calculation operator to obtain sub-data sets, and the sub-data sets are distributed to each calculation slave node of the party, and each calculation slave node executes the corresponding calculation logic according to the calculation operator. After the initiator and the participant receive the uploaded data set, they will split their respective data sets according to the cutting function in their respective calculation operators to obtain sub-data sets. The initiator and each participant send the sub-data sets to each calculation slave node owned by the party. The number of calculation slave nodes owned by the initiator and each participant can be the same or different. The slave nodes of the initiator can participate in data calculation, or, in some embodiments, when the initiator has no data to upload, the initiator can serve as the management party, which is only used for the scheduling and aggregation of data calculations of each participant. After obtaining the sub-data set, the calculation slave nodes of each party perform corresponding calculations on the sub-data set according to the calculation logic in the calculation operator. It should be noted that the above calculations not only include simple mathematical calculations, but also include various data processing methods such as screening, comparison or set operations on data.

[0047] Step S103, the initiator's computing slave node obtains the calculation result according to the corresponding calculation data provided by each participant's computing slave node, and sends the calculation result to the initiator's computing master node. During the execution of the corresponding calculation logic, the computing slave node of the participant sends the calculation data to the computing slave node corresponding to the initiator, and each initiator's computing slave node completes the collection of the corresponding computing data of the participant, and then calculates with the data in the initiator's computing slave node to obtain the calculation result. Each initiator's computing slave node sends the calculation result to the initiator's computing master node.

[0048] Step S104: The initiator's computing master node aggregates the computation results according to the computing operator to obtain aggregated data. The initiator's computing master node calls the aggregation function in the computing operator to aggregate the computation results in all computing sub-nodes to obtain the final aggregated data, which is also the final result of the multi-party computation.

[0049] Through the above steps, a multi-party data processing method for multi-party privacy computing with high versatility is provided. Through the independently set cutting functions, calculation logic and aggregation functions in the calculation operator, data cutting, efficient calculation processing and aggregation are realized in the case of huge data volume. In addition, due to the block setting of cutting functions, calculation logic and aggregation functions, algorithm developers do not need to pay attention to writing data cutting logic, aggregation logic and data transmission between nodes, but only need to write and modify the calculation logic.

[0050] In some of these embodiments, Figure 2 is a flowchart of a multi-party data processing method according to another embodiment of the present application. Figure 2 As shown, the method further comprises the following steps:

[0051] Step S201, according to the data cutting function, the data set is divided into sub-data sets and marked with sub-data set numbers, and the sub-data sets are distributed to each computing slave node of the party, and the computing slave nodes of the initiator and the computing slave nodes of the participating parties with the same sub-data set number execute the corresponding computing logic. In this step, the initiator and all participating parties split their own data sets into multiple sub-data sets according to the data cutting function of the operator, and return the id of each sub-data set, that is, the sub-data set number.

[0052] Optionally, an optional data cutting function is provided in the data cutting function. When the data set is a vector data set, it is cut in segments, and the sub-data set number is the corresponding segment number. Vector operation is to calculate the data with the same sub-data set number, such as vector addition, multiplication, etc.;

[0053] When the data set is a set-type data set, read each element in the data set, calculate the bucket number by taking the modulus after hash operation, and write it into the corresponding sub-data set file. The sub-data set number is the bucket number. Usually, the hash function can also be called a hash function. The function of the hash function is to obtain the hash value of the target through a mapping method, or a function operation f. The function f here is called a hash function or hash function. The hash bucket algorithm is to resolve hash conflicts. That is, different target keys get the same hash value after mapping. The so-called hash bucket algorithm is actually a method to resolve chain address conflicts. For example, the number of buckets is set to 5, which is the number of f(key) sets. In this case, the hash value can be used as the index of the bucket. 1, 2, 3, 4, 5 are obtained by f(key) to obtain 1, 2, 3, 4, 0 respectively. These keys can be put into the memory pointed to by the first address of buckets 1, 2, 3, 4, 0, and then the key with a value of 6 is processed to obtain a hash value of 1, which needs to be put into bucket 1, but the first address of bucket 1 already has element 1. Then a piece of memory can be opened for each bucket, and all keys with the same hash value are stored in the memory. The conflicting keys are stored in a one-way linked list, and the hash conflict is resolved. When searching for the corresponding key, you only need to use the key index to find the corresponding bucket, and then start searching from the node corresponding to the first address of the bucket, that is, find it in the linked list order, and compare the key value until the corresponding key information is found. In this embodiment, the elements in the data set are hashed and bucket operations are performed, and the bucket number is calculated and written into the corresponding sub-data set file, and the bucket number is used as the sub-data set label.

[0054] It should be noted that the two preset data segmentation functions provided in the above embodiment can cover most data segmentation scenarios, so when calling the calculation operator provided by this method, there is no need to write and set the data segmentation function additionally. In the multi-party computing process, if there is a need for other data segmentation methods, the data segmentation method can also be edited and called by modifying the part of the data segmentation function in the calculation operator. The above data segmentation function is not used to limit the data segmentation method applicable to this method.

[0055] In step S202, the initiator's computing slave node obtains the calculation result according to the corresponding calculation data provided by each participant's computing slave node, and sends the calculation result to the initiator's computing master node. Since the sub-datasets are divided according to the data cutting function in the computing operator, the sub-datasets that need to interact with the protocol data in the multi-party computing process have the same sub-dataset number. Each computing slave node of the initiator starts to execute the initiator logic in the computing operator at the same time, and each computing slave node of the participant starts to execute the participant logic in the computing operator at the same time. The corresponding algorithm logic will be executed between the initiator and participant nodes with the same sub-dataset id, and data interaction may occur.

[0056] In this embodiment, by labeling the sub-datasets after data segmentation with sub-dataset numbers, the sub-datasets that need to be interacted with by multiple parties have the same sub-dataset numbers in the initiator and each participant, which will further improve the efficiency of multi-party secure computing.

[0057] In some of these embodiments, Figure 3 is a flowchart of another multi-party data interaction method according to an embodiment of the present application. Figure 3 As shown, the calculation results obtained by the initiator's calculation slave node based on the corresponding calculation data provided by the calculation slave nodes of each participant include:

[0058] Step S301, the participant's computing slave node sends the computing data and the corresponding sub-data set number to the participant's computing master node;

[0059] Step S302, the participant's computing master node sends the computing data and the corresponding sub-data set number to the initiator's computing master node, the initiator's computing master node sends the computing data to the corresponding initiator's computing slave node according to the sub-data set number, and the initiator's computing slave node obtains the computing result according to the corresponding computing data provided by each participant's computing slave node.

[0060] In this embodiment, a data interaction method is provided, that is, the data interaction between the initiator and each participating party during the calculation process needs to be transmitted through the computing master node of each party, and cannot be directly carried out between the computing slave nodes of each party, thereby improving the security of the calculation data.

[0061] In some embodiments, Figure 4 is a flowchart of another multi-party data interaction method according to an embodiment of the present application. Figure 4 As shown, the initiator and the participant obtain their own calculation operators respectively, including:

[0062] Step S401, the initiator obtains the calculation operator, parses the configuration file in the calculation operator, reads the algorithm type and version number therein, and performs overwriting if a calculation operator of the same algorithm type already exists. The written calculation operator code is packaged and uploaded to the system of the initiator of the calculation operator call. The system parses the configuration file in the calculation operator package, reads the algorithm type and version number therein, and overwrites and replaces the old calculation operator package if a calculation operator package of the same type and version already exists.

[0063] Step S402: The initiator distributes the computing operator to each participant. The initiator distributes the computing operator to each participant. The distribution method of the computing operator is not limited to the initiator obtaining, updating and distributing.

[0064] This embodiment provides a registration process for computing operators and supports dynamic upgrade and replacement of computing operators. When there are many participants and the number of nodes for each party's calculation is expanded to dozens, the system operation and maintenance can be carried out very conveniently.

[0065] The embodiments of the present application are described and illustrated below through preferred embodiments.

[0066] Figure 5 is a schematic diagram of a multi-party data processing method according to a preferred embodiment of the present application, such as Figure 5 As shown in the figure, the execution process of the computing operator of multi-party secure privacy computing is as follows:

[0067] Step S1: The initiator and each participant of the computing operator uploads the data set to their own system respectively.

[0068] Step S2: The initiator of privacy computing distributes the operator called this time to all participants to ensure that the operators of all parties are up to date. Figure 5 Only institutions 1 and 2 participating in the calculation are marked. When institution 1 initiates privacy computing, institution 1 is regarded as the initiator, institution 2 is regarded as the participant, and so on.

[0069] Step S3: The initiator and all participants split their own data sets into multiple sub-data sets according to the operator's data cutting function, and return the sub-data set number, i.e., the sub-data set id, of each sub-data set. Figure 5 After data set segmentation is performed on the data sets in institution 1 and institution 2, sub-dataset 0, sub-dataset 1, and sub-dataset 2 are obtained respectively.

[0070] The data cutting function provides two default optional data cutting methods with high applicability: if it is a vector data set, it is cut directly in segments, and the sub-dataset id is its corresponding serial number; if it is a set data set, set operations, multi-party intersection, union or complement, etc., read each element in the data set, calculate the bucket number of the sub-dataset to which it belongs after hash bucket operation, and write it to the corresponding sub-dataset file, and the sub-dataset id is the calculated bucket number.

[0071] Step S4: The initiator and each participant distribute the sub-datasets evenly to all privacy computing slave nodes owned by the party. The purpose of even distribution is to further improve the computing efficiency and prevent the computing amount of each computing slave node from being greatly different. There is no restriction on the distribution method of sub-datasets in actual applications. Figure 5 ,Each computing node, including the computing master node and the computing slave node, is allocated a sub-data set.

[0072] In step S5, each computing slave node of the initiator for privacy computing starts to execute the initiator logic of the computing operator at the same time, and each computing slave node of the participant for privacy computing starts to execute the participant logic of the computing operator at the same time. The data used are all sub-data sets allocated by the master node.

[0073] Step S6: The corresponding algorithm logic of privacy computing is executed between the initiator and the participating nodes with the same sub-dataset ID. During the data interaction between the parties, the slave node cannot directly send data to the slave nodes of other institutions. It needs to send it to the computing master node of the party first, and then the computing master node of the party sends the data and the sub-dataset ID to which the data belongs to to the computing master node of the other party. After receiving the data, the computing master node of the other party sends it to the computing slave node corresponding to the receiving data according to the sub-dataset ID to which the data belongs. Figure 5 As shown, there is data transmission between the computing master nodes of institution 1 and institution 2, while there is no actual data transmission between the computing slave nodes assigned with the same sub-dataset number, such as computing slave node 1 of institution 1 and computing slave node 2 of institution 2, but the corresponding algorithm agreed by the computing operator will be executed. Data transmission between computing slave nodes only occurs between the computing master nodes of the same institution.

[0074] Step S7: After all sub-data sets have executed the algorithm logic, all computing slave nodes of the initiator send all computing results to the computing master node. The computing master node of the initiator calls the aggregation function in the operator package to aggregate all sub-results into a complete computing result.

[0075] Through the above preferred embodiments, the system natively supports the expansion of data volume. Privacy computing algorithm developers do not need to pay attention to writing data splitting and aggregation calculation logic and data transmission between nodes in the algorithm process, and can focus more on writing the logic of the algorithm itself. This method supports the dynamic upgrade and replacement of privacy computing operators. When the number of nodes is expanded to dozens, the system operation and maintenance can be carried out very conveniently.

[0076] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here. For example, the order of step S1 and step S2 can be interchanged.

[0077] This embodiment also provides a multi-party data processing system, which is used to implement the above embodiments and preferred implementations. As used below, the terms "module", "unit", "subunit", etc. can implement a combination of software and / or hardware for a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0078] Figure 6 is a structural block diagram of a multi-party data processing system according to an embodiment of the present application, such as Figure 6 As shown, the system includes an initiator 60 and a participant 64. In actual applications, the number of initiators and participants may be more than one. The initiator 60 includes an initiator computing master node 62 and an initiator computing slave node 63, and the participant 64 includes a participant computing master node 66 and a participant computing slave node 67:

[0079] The initiator 60 and the participant 64 respectively obtain their own data sets and computing operators, divide the data sets according to the computing operators to obtain sub-data sets, and distribute the sub-data sets to their own computing slave nodes, including the initiator computing slave node 63 and the participant computing slave node 67. Each computing slave node executes the corresponding computing logic according to the computing operator.

[0080] During the process of executing the calculation logic by the initiator's calculation slave node 63 and the participant's calculation slave node 67, the initiator's calculation slave node 63 obtains the calculation result based on the corresponding calculation data provided by each participant's calculation slave node 67, and sends the calculation result to the initiator's calculation master node 62. The initiator's calculation master node 62 aggregates the calculation result according to the calculation operator to obtain aggregated data.

[0081] In some embodiments, Figure 7 is a structural block diagram of a computing node according to an embodiment of the present application, such as Figure 7As shown, the computing node includes a scheduler 72, and the scheduler includes a virtual machine 74 and a network component 76. The computing node includes a computing master node and a computing slave node, and the virtual machine 74 is used to execute the multi-party data processing method in the above-mentioned various embodiments and the preferred embodiment. The network component 76 is used to perform data communication between computing nodes, realize data interaction and computing operator distribution, etc.

[0082] In some embodiments, the initiator 60 and the participant 64 perform data segmentation on the data set according to the data segmentation function to obtain sub-data sets and label the sub-data set numbers, and distribute the sub-data sets to each of the computing slave nodes of the party, the initiator computing slave node 63 and the participant computing slave node 67, and the initiator computing slave node 63 and the participant computing slave node 67 with the same sub-data set number execute the corresponding computing logic;

[0083] During the execution of the calculation logic, the participating party's calculation slave node 67 sends the calculation data to the initiating party's calculation slave node 63 with the same sub-data set number. The initiating party's calculation slave node 63 obtains the calculation result based on the corresponding calculation data provided by each participating party's calculation slave node 67, and sends the calculation result to the initiating party's calculation master node 62.

[0084] In some embodiments, the participant computing slave node 67 sends the computing data and the corresponding sub-data set number to the participant computing master node 66, the participant computing master node 66 sends the computing data and the corresponding sub-data set number to the initiator computing master node 62, the initiator computing master node 62 sends the computing data to the corresponding initiator computing slave node 63 according to the sub-data set number, and the initiator computing slave node 63 obtains the computing result according to the corresponding computing data provided by each participant computing slave node.

[0085] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination. For specific examples in this embodiment, reference can be made to the examples described in the above embodiments and optional implementations, and no further description is given here.

[0086] This embodiment further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0087] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0088] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0089] The initiator and participants obtain their own data sets and computing operators respectively;

[0090] According to the data cutting function in the calculation operator, the data set is split into sub-data sets, and the sub-data sets are distributed to each calculation slave node of the party. Each calculation slave node executes the corresponding calculation logic according to the calculation operator;

[0091] The initiator's computing slave node obtains the calculation result based on the corresponding calculation data provided by the computing slave nodes of each participant, and sends the calculation result to the initiator's computing master node;

[0092] The initiator's computing master node aggregates the calculation results according to the computing operator to obtain aggregated data.

[0093] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0094] In addition, in combination with the multi-party data processing method in the above embodiment, the embodiment of the present application can provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by a processor, any multi-party data processing method in the above embodiment is implemented.

[0095] Those skilled in the art should understand that the technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0096] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A multi-party data processing method, characterized in that: include: The initiator and participants obtain their own data sets and computing operators respectively; The data set is segmented according to the data cutting function in the computing operator to obtain sub-data sets, and the sub-data sets are distributed to each computing slave node of the party, and each computing slave node executes the corresponding computing logic according to the computing operator; The initiator's computing slave node obtains the computing result according to the corresponding computing data provided by each participant's computing slave node, and sends the computing result to the initiator's computing master node; The initiator computing master node aggregates the computing results according to the computing operator to obtain aggregated data.

2. The multi-party data processing method according to claim 1, characterized in that: The method further comprises: According to the data cutting function, the data set is segmented to obtain the sub-data sets and the sub-data set numbers are marked, and the sub-data sets are distributed to each computing slave node of the party, and the computing slave node of the initiator and the computing slave node of the participant having the same sub-data set number execute the corresponding computing logic; The initiator's computing slave node obtains a computing result based on the corresponding computing data provided by each participant's computing slave node, and sends the computing result to the initiator's computing master node.

3. The multi-party data processing method according to claim 2, characterized in that: The calculation result obtained by the initiator's calculation slave node according to the corresponding calculation data provided by the calculation slave nodes of each participant includes: The participant computing slave node sends the computing data and the corresponding sub-data set number to the participant computing master node. The participating party's computing master node sends the computing data and the corresponding sub-data set number to the initiating party's computing master node. The initiating party's computing master node sends the computing data to the corresponding initiating party's computing slave node according to the sub-data set number. The initiating party's computing slave node obtains the computing result according to the corresponding computing data provided by each participating party's computing slave node.

4. The multi-party data processing method according to claim 2, characterized in that: The step of segmenting the data set to obtain the sub-data sets and labeling the sub-data sets according to the data segmentation function in the calculation operator includes: When the data set is a vector data set, it is cut in segments, and the sub-data set number is the corresponding segment number; When the data set is a set-type data set, each element in the data set is read, and the bucket number is calculated by modulus after hash operation and written into the corresponding sub-data set file. The sub-data set is labeled as the bucket number.

5. The multi-party data processing method according to claim 2, characterized in that: The initiator and the participant respectively obtain their own calculation operators including: The initiator obtains the computing operator, parses the configuration file in the computing operator, reads the algorithm type and version number therein, and performs overwriting if the computing operator of the same algorithm type already exists; The initiator distributes the computing operator to each of the participants.

6. The multi-party data processing method according to claim 1, characterized in that: The calculation operator includes independently set data cutting functions, data aggregation functions and message types of the calculation data.

7. A multi-party data processing system, characterized in that: The initiator includes an initiator computing master node and an initiator computing slave node, and the participant includes a participant computing master node and a participant computing slave node: The initiator and the participant respectively obtain their own data sets and computing operators, perform data segmentation on the data sets according to the computing operators to obtain sub-data sets, and distribute the sub-data sets to their own computing slave nodes, and each computing slave node executes corresponding computing logic according to the computing operators; The initiator's computing slave node obtains the calculation result according to the corresponding calculation data provided by each participant's computing slave node, and sends the calculation result to the initiator's computing master node. The initiator's computing master node aggregates the calculation result according to the calculation operator to obtain aggregated data.

8. The multi-party data processing system according to claim 7, characterized in that: The computing node includes a scheduler, and the scheduler includes a virtual machine and a network component, wherein the computing node includes a computing master node and a computing slave node, The virtual machine executes the multi-party data processing method according to any one of claims 1 to 6; The network component is used for data communication between the computing nodes.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the multi-party data processing method according to any one of claims 1 to 6.

10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the multi-party data processing method according to any one of claims 1 to 6 when running.

Citation Information

Patent Citations

  • Secure multi-party privacy computing method and device, equipment and storage medium

    CN112395642A

  • Key-value pair model safety training and reasoning method based on safety multi-party calculation

    CN113535808A