A multi-source data fusion platform based on privacy computing

By adopting polynomial fitting method and sub-task decomposition technology on the privacy computing platform, the problem of low efficiency of existing privacy computing technology is solved, and an efficient and easy-to-promote privacy computing solution is realized, which is suitable for statistics of industry business data.

CN114048518BActive Publication Date: 2025-06-10ZHEJIANG DIGITAL QIN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111254650.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2025-06-10
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

Existing privacy computing technology is inefficient and difficult to effectively promote and use.

Method used

The privacy calculation method based on polynomial fitting is adopted, and the objective function is decomposed into multiple subtasks through the collaborative work of the service node and the fusion node. Privacy calculation is realized through asynchronous allocation of the data source side and the merging processing of the fusion nodes.

Benefits of technology

It improves the efficiency and scalability of privacy computing, reduces complex encryption and communication processes, supports online privacy computing at different times of data sources, and is suitable for the statistical needs of industry operating data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114048518B_ABST
    Figure CN114048518B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information technology, and specifically relates to a multi-source data fusion platform based on privacy computing, which includes service nodes and fusion nodes. The service nodes display privacy computing tasks, register the source identifiers of the data source parties that want to participate in privacy computing, allocate formal variables to the privacy data of the data source parties involved in the objective function. The service nodes establish a polynomial fitting function of the objective function and expand it into the sum of several product terms. The data source parties split the privacy data into several multipliers and allocate them to several data source parties. The data source parties regard the values of the multipliers as the values of the formal variables, calculate the values of each product term respectively, which are recorded as intermediate values, and send them to the fusion nodes. The fusion nodes multiply the intermediate values with the same sub-task numbers, and use the correction values to correct the coefficients of the product terms to obtain the values of the product terms. Adding up the values of all product terms, the value of the polynomial fitting function can be obtained. The substantial effect of the present invention is that it can meet a wide range of privacy computing requirements and has higher execution efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and particularly to a multi-source data fusion platform based on privacy computing. Background Art

[0002] The subdivision of industries can improve the utilization rate of resources and the quality of final products, and has significant economic and social benefits. With the development of the economy, a large number of industrial structures with fine division of labor and cooperation have emerged in the market. In the subdivided industrial structure, there are a large number of market participants. During their production and operation processes, a large amount of data is generated, mainly production data, operation data, and transportation data. Due to a large amount of collaborative production, the relevant data of the upstream and downstream of the participants has an important impact on the production and operation of the participants themselves. Moreover, the overall coordination degree of the participants in the industry ultimately affects the efficiency of the industry. Data is generated during the production process, and production and operation scheduling based on data can also improve the resource utilization rate and production efficiency of the industry. Therefore, in the industry, data is playing an increasingly important role. Currently, data has become the fourth major production factor and has an important impact on production and operation. The value of data is concentrated in the sharing and flow of data. However, due to competitive relationships and confidentiality protection, current market participants have many concerns about sharing data, which severely restricts the realization of data value. Although privacy computing technologies have been proposed to fuse data from various parties, the current privacy computing technologies have the problem of low efficiency.

[0003] For example, Chinese Patent CN112699391A, with a publication date of April 23, 2021, discloses a method and device for sending target data, a storage medium, and an electronic device. The method includes: receiving a computing routine sent by a privacy computing scheduling management unit, where the computing routine is a computing routine generated by the privacy computing scheduling management unit according to a data request sent by a data consumer; running the computing routine to obtain the target data requested by the data request from multiple encrypted data stored in a trusted computing unit, where the trusted computing unit stores multiple encrypted data sent by multiple data senders; sending the target data to an intermediate result computing unit, so that the intermediate result computing unit merges the received multiple target data and sends the merged target data to the data consumer. The privacy computing scheduling management unit, the trusted computing unit, and the intermediate result computing unit are all set in the privacy computing platform. Its technical solution requires the assistance of trusted hardware, with high computing costs, and its popularization and use are severely restricted. Therefore, it is necessary to study privacy computing technologies with high efficiency and easy popularization and use. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: the current lack of high-efficiency privacy computing technology. A multi-source data fusion platform based on privacy computing is proposed.

[0005] To solve the above technical problem, the technical solution adopted by the present invention is: a multi-source data fusion platform based on privacy computing, which is used to count industry operation data, including service nodes and fusion nodes. The data source party connects to the service node, and the service node assigns a source identifier. Any data source party initiates a privacy computing task and submits an objective function. The service node displays the privacy computing task, registers the source identifiers of the data source parties that want to participate in the privacy computing. After reaching the predetermined deadline, the service node assigns formal variables to the privacy numbers of the data source parties involved in the objective function. The service node establishes a polynomial fitting function of the objective function, expands the polynomial fitting function into the sum of several product terms, creates subtasks for each product term, assigns subtask numbers to the subtasks, and sends the subtasks and subtask numbers to all data source parties. The data source parties split their respective privacy numbers into several multipliers, and the several multipliers are assigned to several data source parties. Each data source party obtains one multiplier of all the formal variables, and the data source party regards the value of the multiplier as the value of the formal variable, calculates the value of each product term respectively and records it as an intermediate value, associates the intermediate value with the subtask number and sends it to the fusion node. The fusion node multiplies the intermediate values with the same subtask number, the fusion node requests a correction value from the service node, uses the correction value to correct the product term coefficient and then obtains the value of the product term, and adds up the values of all product terms to obtain the value of the polynomial fitting function, which is the result of the privacy computing.

[0006] Preferably, there are multiple fusion nodes, and the multiple fusion nodes have sequential numbers. The remainder of the subtask number divided by the number of fusion nodes is used as the designated fusion node for the subtask. Several data source parties send the intermediate values of the subtasks to the designated fusion node. The designated fusion node sums up the received intermediate values of the subtasks and then signs and broadcasts them. Any fusion node collects all the broadcast sums and sums them up again to obtain the value of the polynomial fitting function.

[0007] Preferably, the service node obtains the number of data source parties participating in the subtask, subtracts 1 from the number and uses it as the coefficient of the product term corresponding to the subtask as the correction value, associates the correction value with the subtask number and sends it to the fusion node. The fusion node multiplies the intermediate values of the subtasks to obtain the value of the product term, and divides the value of the product term by the correction value to complete the correction of the product term coefficient.

[0008] Preferably, the service node assigns a correction coefficient to each formal variable of each subtask, regards the correction coefficient as the value of the formal variable multiplier, calculates the value of the subtask, so that the value of the subtask is equal to m times the coefficient of the corresponding product term. The service node takes m as the correction value, sends the correction value to the fusion node, and sends the correction coefficient associated with the formal variable and the subtask number to the data source party. The data source party multiplies the multiplier by the correction coefficient and then calculates the intermediate value of the product term. After the fusion node multiplies the intermediate values of the subtasks, it obtains the value of the product term, and divides the value of the product term by the correction value to complete the correction of the product term coefficient.

[0009] Preferably, the service node establishes a multiplier table. The rows of the multiplier table are formal variables, and the columns of the multiplier table are source identifiers. The multiplier table is initially initialized with null values. A number of data source parties all apply for CA certificates, and a number of data source parties allocate multipliers asynchronously. Specifically, it includes: the data source party connects to the service node, queries the row of the formal variable corresponding to its own private number in the multiplier table. If there is a non-null value in the row, it uses the private key to decrypt and obtains the filled multiplier. The data source party generates a number of multipliers to fill the null values in the row. When filling in the multiplier, it encrypts with the public key of the data source party corresponding to the column; the data source party queries the column corresponding to itself. If there are null values in the rows of the column, it generates a random number as the guessed value of the formal variable corresponding to the row of the column, encrypts the guessed value with the public key of the data source party corresponding to the row, and fills it into the multiplier table; the data source party uses the multiplier of each formal variable recorded in the column corresponding to itself to calculate the intermediate value of the subtask; when all data source parties are connected to the service node, the multiplier table will be filled, and the fusion node will receive all the intermediate values of all subtasks, so it can calculate the value of each subtask.

[0010] Preferably, the fusion node stores the intermediate values of the received subtasks. The multiplier of the private number retained by the data source party is recorded as the retained number. After a privacy calculation, if the private number of the data source party is updated and a privacy calculation needs to be performed again, the data source party generates a new retained number so that the product of the retained number and the other multipliers is equal to the updated private number. The data source party calculates the intermediate value of the subtask by itself, sends the intermediate value of the subtask to the fusion node. The fusion node replaces the saved intermediate value with the newly received intermediate value, recalculates the value of each product term, and uses the new value of the product term to calculate the value of the polynomial fitting function again, which is the result of the privacy calculation after the private number is updated.

[0011] Preferably, the fusion node stores the intermediate values of the received subtasks. After one round of privacy calculation, if there is a new data source party that wants to incorporate its private data into the privacy calculation, the new data source party splits its private data into several addends, randomly sends the several addends to several data source parties that have participated in the privacy calculation before. The several data source parties that receive the addends add the addends to their own private data to obtain the updated private data, generate a new reserved number, and make the product of the reserved number and the other multiplicands equal to the updated private data. The data source parties calculate the intermediate values of the subtasks by themselves, send the intermediate values of the subtasks to the fusion node, and the fusion node replaces the saved intermediate values with the newly received intermediate values, recalculates the values of each product term, and uses the values of the new product terms to calculate the value of the polynomial fitting function again, which is the result of the privacy calculation after the private data of the new data source party is added.

[0012] Preferably, when the service node establishes the polynomial fitting function of the objective function, it performs the following steps: The service node requests the value range of each private data in the objective function from the data source parties; within the value range of each private data, uniformly generate example numbers to form multiple values of the private data, randomly combine the example numbers of the private data into a group, denoted as a value group; substitute the value group into the objective function to obtain the output of the objective function, and use the output as a label to mark the value group as sample data.

[0013] Preferably, several data source parties respectively divide the value ranges of their own private data into several intervals, respectively count the probabilities of their own private data falling into each interval as interval probabilities, and send the interval probabilities to the service node; the service node randomly generates example numbers within the value range of the private data so that the distribution probability of the example numbers in the intervals is equal to the interval probabilities; randomly combine the example numbers of the private data into a group, denoted as a value group; substitute the value group into the objective function to obtain the output of the objective function, and use the output as a label to mark the value group as sample data.

[0014] The substantial effects of the present invention are as follows: By using polynomial fitting for the objective function to form a new form expression of the objective function, there is a way to achieve privacy computing under the new form expression. In theory, polynomials can fit any continuous function, so this platform can meet a wide range of privacy computing requirements. The privacy computing is implemented in the mode of subtasks, without the need for complex encryption and communication processes, and only ordinary encryption methods are used, which has higher execution efficiency. Through the privacy computing solution provided by this platform, the data source parties participating in the privacy computing do not need to be online simultaneously and can still complete the privacy computing, so the execution is more convenient. Moreover, through the privacy computing solution provided by this platform, after one privacy computing, if new data source parties want to join the privacy computing, there is no need to start over, and only a few participating data source parties need to assist in the update, which is very suitable for the statistics of industry operation data and provides a reference for market entities to judge industry trends and reasonably arrange production and operation activities. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic structural diagram of the multi-source data fusion platform for Example 1.

[0016] Figure 2 It is a schematic diagram of the privacy computing method for Example 1.

[0017] Figure 3 It is a schematic diagram of the method for correcting the coefficients of product terms for Example 1.

[0018] Figure 4 It is a schematic diagram of the alternative method for correcting the coefficients of product terms for Example 1.

[0019] Figure 5 It is a schematic diagram of the method for asynchronously allocating multipliers to data source parties for Example 1.

[0020] Figure 6 It is a schematic diagram of the secondary privacy computing method for Example 1.

[0021] Figure 7 It is a schematic diagram of the method for a new data source party to join the privacy computing for Example 1.

[0022] Figure 8 It is a schematic diagram of the method for establishing a polynomial fitting function for Example 1.

[0023] Wherein: 10, service node; 20, data source party; 30, fusion node. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The following will further specifically describe the specific embodiments of the present invention through specific examples in conjunction with the drawings.

[0025] A multi-source data fusion platform based on privacy computing is used for statistically analyzing industry operation data. Please refer to the attached Figure 1, this platform includes service node 10 and fusion nodes 30. There are multiple fusion nodes 30, and the multiple fusion nodes 30 have sequential numbers. Please refer to the appendix Figure 2 , the method for performing privacy computing on this platform includes: Step A01) The data source party 20 connects to the service node 10, and the service node 10 assigns a source identifier. Any data source party 20 initiates a privacy computing task and submits an objective function; Step A02) The service node 10 displays the privacy computing task and registers the source identifiers of the data source parties 20 that intend to participate in the privacy computing; Step A03) After reaching the predetermined deadline, the service node 10 assigns formal variables to the privacy numbers of the data source parties 20 involved in the objective function; Step A04) The service node 10 establishes a polynomial fitting function of the objective function and expands the polynomial fitting function into the sum of several product terms; Step A05) Establish sub-tasks for each product term, assign sub-task numbers to the sub-tasks, and send the sub-tasks and sub-task numbers to all data source parties 20; Step A06) The data source parties 20 split their respective privacy numbers into several multipliers, and the several multipliers are assigned to several data source parties 20. Each data source party 20 obtains one multiplier of all the formal variables; Step A07) The data source parties 20 regard the values of the multipliers as the values of the formal variables, calculate the values of each product term respectively and record them as intermediate values, and send the intermediate values associated with the sub-task numbers to the fusion nodes 30; Step A08) The fusion nodes 30 multiply the intermediate values with the same sub-task numbers, and the fusion nodes 30 request correction values from the service node 10; Step A09) Use the correction values to correct the coefficients of the product terms to obtain the values of the product terms, and add up the values of all the product terms to obtain the value of the polynomial fitting function, which is the result of the privacy computing. Among them, the remainder of the sub-task number divided by the number of fusion nodes 30 is used as the designated fusion node 30 for the sub-task. Several data source parties 20 send the intermediate values of the sub-task to the designated fusion node 30. The designated fusion node 30 sums up the received intermediate values of the sub-task, signs and broadcasts them. Any fusion node 30 collects the sum of all the broadcasts and sums them up again to obtain the value of the polynomial fitting function.

[0026] Table 1 Privacy Computing Tasks Publicly Announced by Service Node 10

[0027] Name Objective Function Registered Data Source Party 20 Based on the sales volumes of three mobile phone brands A / B / C, calculate the mobile phone consumption heat in the first quarter Consumption Heat = w1 * Total Sales of the Three Brands + w2 * Variance of the Sales Volumes of the Three Brands' Mobile Phones Dealer A, Dealer B, Dealer C, Dealer D

[0028] As shown in Table 1, it is a privacy calculation initiated by mobile phone dealers in a certain city, aiming to calculate the mobile phone consumption heat in this city this quarter. The calculation method is to select the sales volumes of mobile phones of three representative brands, namely mobile phone brands A / B / C. Brands A / B / C are used to represent high / medium / low-end mobile phones respectively, representing the mobile phone market in a certain city. Since mobile phones of each brand are sold by multiple dealers, the more dealers who sign up to participate in this privacy calculation, the more representative the calculation result will be. In actual implementation, the privacy calculation requires dealers to participate voluntarily. In this embodiment, it is assumed that dealers A, B, C, and D are the main mobile phone sellers in a certain city, and it is assumed that all four dealers are willing to participate in this privacy calculation. That is, dealers A, B, C, and D are respectively used as the four data source parties in this embodiment.

[0029] Dealer A sells mobile phones of brand A, dealer B sells mobile phones of brand B, dealer C sells mobile phones of brand C, and dealer D sells mobile phones of brand B and brand C. When the registration deadline arrives, service node 10 generates formal variables. Generating formal variables is divided into two steps. First, generate the mathematical expression of the objective function, that is, express the objective function: consumption heat = w1 * total sales volume of three brands + w2 * variance of the sales volumes of mobile phones of three brands as: f(x,y,z)=w1*Sn+w2*((x-μ)^2+(y-μ)^2+(z-μ)^2) / 3, where Sn=x+y+z, μ=Sn / 3, and x, y, z respectively represent the quarterly sales volumes of mobile phone brands A / B / C. x, y, z are the initial formal variables. Since the quarterly sales volume y of brand B is the sum of the sales volumes of dealer B and dealer D, and the quarterly sales volume z of brand C is the sum of the sales volumes of dealer C and dealer D, it is necessary to further divide y and z into two formal variables respectively, that is, y=y1+y2, where y1 represents the quarterly sales volume of brand B mobile phones of dealer B, and y2 represents the quarterly sales volume of brand B mobile phones of dealer D, z=z1+z2, where z1 represents the quarterly sales volume of brand C mobile phones of dealer C, and z2 represents the quarterly sales volume of brand C mobile phones of dealer D. Thus, the expression of the objective function is finally:

[0030] f(x,y1,y2,z1,z2)=w1*Sn+w2*((x-μ)^2+(y1+y2-μ)^2+(z1+z2-μ)^2) / 3, where Sn=x+y1+y2+z1+z2, μ=Sn / 3. Allocate x to dealer A, y1 to dealer B, z1 to dealer C, and y2 and z2 to dealer D. Expand the objective function to obtain:

[0031] f(x, y1, y2, z1, z2) = w1 * x + w1 * y1 + w1 * y2 + w1 * z1 + w1 * z2 + w2 / 3 * x^2 - 2 * w2 / 3 * x * μ + w2 * μ^2 + w2 / 3 * y1^2 + w2 / 3 * y2^2 + 2 * w2 / 3 * y1 * y2 - 2 * w2 / 3 * y1 * μ - 2 * w2 / 3 * y2 * μ + w2 / 3 * z1^2 + w2 / 3 * z2^2 + 2 * w2 / 3 * z1 * z2 - 2 * w2 / 3 * z1 * μ - 2 * w2 / 3 * z2 * μ。

[0032] Replacing μ in the above formula with (x + y1 + y2 + z1 + z2) / 3, an expression of the polynomial fitting function containing only x, y1, y2, z1, and z2 can be obtained. For the convenience of narration in this embodiment, the expanded product terms are expressed using a general term. The general term is: W * x^i1 * y1^i2 * y2^i3 * z1^i4 * z2^i5, where W represents the coefficient of the product term, and each product term has its own coefficient W. For the terms not in the objective function, that is, the coefficient W of the general term takes the value 0. For the convenience of record in this embodiment, the subscripts of the coefficients are not specially marked. For example, for the product term w1 * x, it is the product term where W in the general term takes the value w1, i1 takes the value 1, and i2 to i5 take the value 0. For the product term 2 * w2 / 3 * y1 * y2, it is the product term where W in the general term takes the value 2 * w2 / 3, i1, i4, and i5 take the value 0, and i2 and i3 take the value 1. For the product term w2 / 3 * z2^2, it is the product term where W in the general term takes the value w2 / 3, i2 to i4 take the value 0, and i5 takes the value 2. Thus, the expression of the objective function is f(x, y1, y2, z1, z2) = ∑W * x^i1 * y1^i2 * y2^i3 * z1^i4 * z2^i5, where i1 to i5 all take values from [0, 2]. The objective function given in this embodiment can be directly written in the form of a polynomial, so the polynomial expression is also in this form. If the polynomial fitting function is obtained through the sample data fitting method given in this embodiment, then when the fitting is appropriate, an expression of the polynomial fitting function that is basically the same as the objective function expression should be obtained. However, regardless of the polynomial fitting function, the polynomial fitting function also has the above-mentioned general term form.

[0033] The number of product terms in the expression of the final polynomial fitting function is denoted as N. Then, the service node 10 generates N subtasks, and the product term corresponding to the subtask is denoted as F_j = W_j * x^i1 * y1^i2 * y2^i3 * z1^i4 * z2^i5, where j ∈ [1, N]. Here, j is the subtask number. The subtask numbers have an order, but this order does not mean the order of the product terms in the polynomial fitting function. That is, after the terms of the polynomial fitting function are scrambled, the subtasks are generated and the subtask numbers are assigned to the subtasks.

[0034] Send the subtasks and subtask numbers to the data source party 20, that is, send all (j, W_j*x^i1*y1^i2*y2^i3*z1^i4*z2^i5) in the polynomial fitting function to distributors A, B, C, and D. The subtasks and subtask numbers are in the form of (1, w1*x), (2, 2*w2 / 3*y1*y2), (3, w2 / 3*z2^2), ….

[0035] Then, distributors A, B, C, and D respectively generate multipliers for their respective private numbers. That is, distributor A generates x = x1*x2*x3*x4, and distributes x1 to x4 to A, B, C, and D respectively. Distributor B generates y1 = y11*y12*y13*y14, distributor C generates z1 = z11*z12*z13*z14, and distributor D generates y2 = y21*y22*y23*y24 and z2 = z21*z22*z23*z24 respectively. The multipliers are distributed among A, B, C, and D. The multipliers received by distributor A are x1, y11, y21, z11, and z21. Similarly, distributors B, C, and D also receive one multiplier for each formal variable.

[0036] Each distributor obtains one multiplier for each formal variable. Regarding the multipliers as the values of the formal variables, each distributor can thus calculate the values of the product terms in each subtask. For example, for the subtask (1, w1*x), distributors A, B, C, and D can all calculate a result, which is the intermediate value. Denote them as (1, T_1A), (1, T_1B), (1, T_1C), (1, T_1D) respectively, and send these four vectors to the fusion node 30. Similarly, for (2, 2*w2 / 3*y1*y2) and (3, w2 / 3*z2^2), there are also corresponding four intermediate values.

[0037] The fusion node 30 receives (1, T_1A), (1, T_1B), (1, T_1C), (1, T_1D). Since the fusion node 30 does not know the specific form of the product terms in the subtasks and only knows that these four intermediate values all belong to the subtask numbered 1, it is unable to reverse the values of the private numbers based on the intermediate values. The fusion node 30 multiplies T_1A to T_1D to obtain a result R_1. R_1 is not yet the result of the first product term because when calculating the intermediate values, each distributor multiplied the coefficient W_1 of the product term into the intermediate values, and the calculations of the four distributors caused the intermediate values to be multiplied four times. Therefore, the service node 10 needs to generate the cube of the coefficient W_1 as the correction value. Specifically, please refer to the appendix Figure 3, the method for the service node 10 and the fusion node 30 to correct the coefficient of the product term of the subtask includes: Step B01) The service node 10 obtains the number of data source parties 20 participating in the subtask, takes the value obtained by subtracting 1 from the number as the exponent of the coefficient of the product term corresponding to the subtask for exponentiation, takes it as the correction value, and sends the correction value associated with the subtask number to the fusion node 30; Step B02) The fusion node 30 multiplies the intermediate values of the subtasks to obtain the value of the product term, and divides the value of the product term by the correction value to complete the correction of the coefficient of the product term.

[0038] The fusion node 30 divides R_1 by the correction value again to obtain the correct value of the product term. The fusion node 30 can thus reverse-calculate the value of the coefficient, but still cannot reverse-calculate the degree of each formal variable in the product term, and thus still cannot reverse-calculate the value of the private number. After the values of all subtasks are obtained, summation is performed to obtain the result of the polynomial fitting function. That is the result of the private calculation.

[0039] This embodiment also provides an alternative implementation for correcting the coefficient of the product term of the subtask, so that the fusion node 30 cannot obtain the accurate coefficient of the product term. Please refer to the appendix Figure 4, specifically including: Step B11) The service node 10 assigns a correction coefficient to each formal variable of each subtask; Step B12) Regarding the correction coefficient as the value of the formal variable multiplier, calculate the value of the subtask, so that the value of the subtask is equal to m times the coefficient of the corresponding product term; Step B13) The service node takes m as the correction value; Step B14) Send the correction value to the fusion node 30, and send the correction coefficient associated with the formal variable and the subtask number to the data source party 20; Step B15) The data source party 20 multiplies the multiplier by the correction coefficient and then calculates the intermediate value of the product term; Step B16) The fusion node 30 multiplies the intermediate values of the subtasks to obtain the value of the product term, and divides the value of the product term by the correction value to complete the correction of the coefficient of the product term. For example, for the subtask (1, w1 * x), the correction coefficient generated by the service node 10 is such that when regarding the correction coefficient as the value of the formal variable of the product term, the calculation result is m. Then, the value of the correction coefficient can be deduced as (m * w1^-3)^0.25, and m is sent to the fusion node 30 as the correction value. In this way, the fusion node 30 cannot obtain the accurate value of the coefficient of the subtask term. For the corresponding subtask (2, 2 * w2 / 3 * y1 * y2), a correction coefficient q is generated such that (2 * w2 / 3 * q * q)^4 = m * (2 * w2 / 3), then q = (m * (2 * w2 / 3)^-3)^0.125. When the four data dealers calculate the intermediate value of the subtask (2, 2 * w2 / 3 * y1 * y2), they first multiply the multipliers of y1 and y2 by q and then calculate the intermediate value of the subtask. Finally, the result of multiplying the intermediate values calculated by the four data dealers is: (2 * w2 / 3)^4 * y1 * y2 * q^8 = (2 * w2 / 3)^4 * y1 * y2 * m * (2 * w2 / 3)^-3 = (2 * w2 / 3) * y1 * y2 * m, that is, exactly one more m. After receiving the correction value, the fusion node 30 divides the result of multiplying the intermediate values by the correction value, which is exactly the value of the product term. By using m to confuse the coefficient value of the product term, the privacy number can be better protected.

[0040] The service node 10 establishes a multiplier table. The rows of the multiplier table are formal variables, the columns of the multiplier table are source identifiers, and the multiplier table is initialized to null values. Several data source parties 20 all apply for CA certificates, and several data source parties 20 allocate multipliers in an asynchronous manner. Please refer to the appendix Figure 5, specifically including: Step C01) The data source party 20 connects to the service node 10 and queries the row of the formal variable corresponding to its own privacy number in the multiplier table; Step C02) If there are already non-null values in the row, decrypt using the private key to obtain the filled multipliers. The data source party 20 generates several multipliers to fill the null values in the row. When filling in the multipliers, encrypt using the public key of the data source party 20 corresponding to the column; Step C03) The data source party 20 queries the column corresponding to itself. If there are null values in the rows of the column, generate random numbers as the guessed values of the formal variables corresponding to the rows of the column, and fill them into the multiplier table after encrypting the guessed values using the public key of the data source party 20 corresponding to the row; Step C04) The data source party 20 uses the multipliers of each formal variable recorded in the column corresponding to itself to calculate the intermediate value of the subtask; Step C05) When all data source parties 20 are connected to the service node 10, the multiplier table will be filled, and the fusion node 30 will receive all the intermediate values of all subtasks, so it can calculate the value of each subtask.

[0041] As shown in Table 2, it is the initial product table established in this embodiment, and the values in the table are all null values NULL. Suppose dealer A first connects to the service node 10 and queries the row of the formal variable x corresponding to its own privacy number, and finds that all are null values, then generates 4 multipliers, x = x1 * x2 * x3 * x4, and fills x1, x2, x3, and x4 into the product table after encrypting them using the public keys of dealers A, B, C, and D respectively, as shown in Table 3. Then dealer A queries the column corresponding to itself and finds that there is only a non-null value at the formal variable x. Then generate random numbers as the guessed values of y1, y2, z1, and z2 respectively, and fill them into the product table after encrypting them using the public keys of the data source parties 20 corresponding to the respective rows. The same is shown in Table 3. Among them, Key A() represents encryption using the public key of dealer A. Dealer A obtains all the required multipliers, calculates the intermediate value of the subtask, and sends the intermediate value to the fusion node 30.

[0042] Then dealer B connects to the service node 10, queries the row corresponding to its own privacy number, and finds that there is already a non-null value. Decrypt using its own private key to obtain the value of y11, and then randomly generate other multipliers so that all multipliers meet the requirements. Fill the newly generated multipliers into the multiplier table. Subsequently, query the column corresponding to itself and find that there are three null values, corresponding to the multipliers of y2, z1, and z2 respectively. Then generate 3 random numbers as the guessed values and fill them into the multiplier table after encrypting them using the public keys of the corresponding dealers respectively. At this time, dealer B also has enough multipliers to calculate the intermediate value of the subtask. Similarly, subsequently, dealers C and D connect to the service node 10 successively and execute the same steps, and finally the calculation of the polynomial fitting function can be completed.

[0043] Table 2 Initial multiplier table established in this embodiment

[0044] A B C D x NULL NULL NULL NULL y1 NULL NULL NULL NULL z1 NULL NULL NULL NULL y2 NULL NULL NULL NULL z2 NULL NULL NULL NULL

[0045] Table 3 Multiplier Table after Dealer A's Visit

[0046] A B C D x Key A(x1) Key B(x2) Key C(x3) Key D(x4) y1 Key B(y11) NULL NULL NULL z1 Key C(z11) NULL NULL NULL y2 Key D(y21) NULL NULL NULL z2 Key D(z21) NULL NULL NULL

[0047] The fusion node 30 stores the intermediate value of the received subtask. The multiplier of the privacy number retained by the data source party 20 is denoted as the retained number. After a privacy calculation, if the privacy number of the data source party 20 needs to be updated and a privacy calculation needs to be performed again, the following steps are executed. Please refer to the appendix Figure 6 , including: Step D01) The data source party 20 generates a new retained number so that the product of the retained number and the other multipliers is equal to the updated privacy number; Step D02) The data source party 20 calculates the intermediate value of the subtask by itself and sends the intermediate value of the subtask to the fusion node 30; Step D03) The fusion node 30 replaces the saved intermediate value with the newly received intermediate value and recalculates the value of each product term; Step D04) Use the value of the new product term to calculate the value of the polynomial fitting function again, which is the result of the privacy calculation after the privacy number is updated.

[0048] After the privacy calculation is completed, due to the negligence of Dealer B, there is a statistical omission in the quarterly sales volume used for the privacy calculation. After re-statistics, the updated value y1 + △ is obtained. The retained number of Dealer B is y12. The value of y12 is regenerated so that y12’ * y11 * y13 * y141 = y1 + △. Use y12’ to recalculate the intermediate value of the relevant subtask and send the intermediate value associated with the subtask number to the fusion node 30. The fusion node 30 replaces the saved intermediate value with the recalculated intermediate value of the subtask, multiplies the intermediate values with the same subtask number, and corrects the coefficient to obtain the value of the subtask, which is the result of the privacy calculation involving the updated value y1 + △.

[0049] The fusion node 30 stores the intermediate value of the received subtask. After a privacy calculation, if there is a new data source party 20 that wants to include its own privacy number in the privacy calculation, the following steps are executed for the privacy calculation. Please refer to the appendix Figure 7, including: Step E01) The new data source party 20 splits its own privacy data into several addends; Step E02) Randomly send several addends to several data source parties 20 that have participated in privacy computing; Step E03) Several data source parties 20 that receive the addends add the addends to their own privacy data to obtain the updated privacy data; Step E04) Generate a new reserved number so that the product of the reserved number and the remaining multipliers is equal to the updated privacy data; Step E05) The data source party 20 calculates the intermediate value of the subtask by itself and sends the intermediate value of the subtask to the fusion node 30; Step E06) The fusion node 30 replaces the saved intermediate value with the newly received intermediate value and recalculates the value of each product term; Step E06) Use the value of the new product term to calculate the value of the polynomial fitting function again, which is the result of privacy computing after the privacy data of the new data source party 20 is added. For example, dealer Wu sells mobile phones of brand B. Since dealer Wu missed the deadline for privacy computing and thus could not be assigned a formal variable, dealer Wu split its sales volume into two addends and sent them to dealer Yi and dealer Ding respectively. Both dealer Yi and dealer Ding regarded it as an update of their own privacy data and executed the process of secondary privacy computing, that is, Steps D01) to D04), to achieve a more valuable privacy computing result.

[0050] When the service node 10 establishes a polynomial fitting function of the objective function, please refer to the appendix Figure 8 , perform the following steps: Step F01) The service node 10 requests the value range of each privacy data in the objective function from the data source party 20; Step F02) Uniformly generate example numbers within the value range of each privacy data to form multiple values of the privacy data, and randomly combine the example numbers of the privacy data into a group, denoted as a value group; Step F03) Substitute the value group into the objective function to obtain the output of the objective function, and use the output as a label to mark the value group as sample data.

[0051] Furthermore, this embodiment uses a sample data generation scheme with distribution probability, which specifically includes the following steps: Several data source parties 20 respectively divide the value range of their own privacy data into several intervals, respectively count the probability that their own privacy data falls into each interval as the interval probability, and send the interval probability to the service node 10; The service node 10 randomly generates example numbers within the value range of the privacy data so that the distribution probability of the example numbers in the interval is equal to the interval probability; Randomly combine the example numbers of the privacy data into a group, denoted as a value group; Substitute the value group into the objective function to obtain the output of the objective function, and use the output as a label to mark the value group as sample data.

[0052] The beneficial technical effects of this embodiment are as follows: By using polynomial fitting for the objective function, a new form of expression of the objective function is formed. Under this new form of expression, there is a way to achieve privacy computing. In theory, polynomials can fit any continuous function, so this platform can meet a wide range of privacy computing requirements. The privacy computing is implemented in the mode of subtasks, without the need for complex encryption and communication processes, and only ordinary encryption methods are used, which has higher execution efficiency. Through the privacy computing solution provided by this platform, the data source party 20 participating in the privacy computing does not need to be online simultaneously and can still complete the privacy computing, so the execution is more convenient. Moreover, through the privacy computing solution provided by this platform, after one privacy computing, if a new data source party 20 wants to join the privacy computing, there is no need to start over, and only a few participating data source parties 20 need to assist in the update, which is very suitable for the statistics of industry operation data and provides a reference for market entities to judge industry trends and reasonably arrange production and business activities.

[0053] The above-described embodiment is only a preferred solution of the present invention and does not impose any form of limitation on the present invention. There are other variations and modifications without exceeding the technical solutions described in the claims.

Claims

1. A multi-source data fusion platform based on privacy computing, used to collect statistics on industry operating data. It is characterized in that It includes a service node and a fusion node. The data source connects to the service node. The service node allocates a source identifier. Any data source initiates a privacy computing task and submits an objective function. The service node displays the privacy computing task and registers the source identifier of the data source that wants to join the privacy computing. After the predetermined deadline is reached, the service node allocates formal variables to the privacy numbers of the data sources involved in the objective function. The service node establishes a polynomial fitting function of the objective function, expands the polynomial fitting function into the sum of several product terms, establishes a subtask for each product term, assigns a subtask number to the subtask, and sends the subtask and the subtask number to the server. The data sources are sent to all data sources. The data sources split their privacy numbers into several multipliers, which are distributed to several data sources. Each data source will obtain a multiplier of all formal variables. The data source regards the value of the multiplier as the value of the formal variable, and calculates the value of each product term separately and records it as the intermediate value. The intermediate value is associated with the subtask number and sent to the fusion node. The fusion node multiplies the intermediate value with the same subtask number. The fusion node asks the service node for the correction value, and uses the correction value to correct the product term coefficient to obtain the value of the product term. The value of all product terms is added to obtain the value of the polynomial fitting function, which is the result of the privacy calculation.

2. According to the multi-source data fusion platform based on privacy computing as described in claim 1, It is characterized in that The service node obtains the number of data sources participating in the subtask, and uses the value after the number is subtracted by 1 as the exponential power of the coefficient of the product term corresponding to the subtask, as the correction value, and sends the correction value associated with the subtask number to the fusion node. The fusion node multiplies the intermediate values ​​of the subtasks to obtain the value of the product term, and divides the value of the product term by the correction value to complete the correction of the coefficient of the product term.

3. According to the multi-source data fusion platform based on privacy computing according to claim 1 or 2, It is characterized in that The service node assigns a correction coefficient to each formal variable of each subtask, regards the correction coefficient as the value of the formal variable multiplier, calculates the value of the subtask, and makes the value of the subtask equal to m times the coefficient of the corresponding product term. The service node uses m as the correction value, sends the correction value to the fusion node, associates the correction coefficient with the formal variable and the subtask number and sends it to the data source. The data source multiplies the multiplier by the correction coefficient and then calculates the intermediate value of the product term. The fusion node multiplies the intermediate values ​​of the subtasks to obtain the value of the product term, and divides the value of the product term by the correction value to complete the correction of the coefficient of the product term.

4. According to the multi-source data fusion platform based on privacy computing according to claim 1 or 2, It is characterized in that The service node establishes a multiplier table, the behavior of the multiplier table is a variable, the column of the multiplier table is a source identifier, the multiplier table is initialized to a null value, several data sources all apply for CA certificates, and several data sources allocate multipliers in an asynchronous manner, specifically including: The data source party connects to the service node, queries the rows of the multiplier table corresponding to the formal variables of its own private numbers. If there are already non-null values in the rows, it decrypts them using the private key to obtain the filled multipliers. The data source party generates several multipliers to fill the null values in the rows. When filling in the multipliers, it encrypts them using the public key of the data source party corresponding to the column. The data source party queries the columns corresponding to itself. If there are null values in the rows of the columns, it generates random numbers as the guessed values of the formal variables corresponding to the rows of the columns, encrypts the guessed values using the public key of the data source party corresponding to the rows, and fills them into the multiplier table. The data source party uses the multipliers of each formal variable recorded in the columns corresponding to itself to calculate the intermediate values of the subtasks. When all data source parties are connected to the service node, the multiplier table will be filled, and the fusion node will receive all the intermediate values of all subtasks, so it can calculate the values of each subtask.

5. A multi-source data fusion platform based on privacy computing according to claim 1 or 2, characterized in that, the fusion node stores the intermediate values of the received subtasks. The multipliers of the private numbers retained by the data source party are recorded as retained numbers. After one privacy calculation, if the private numbers of the data source party are updated and privacy calculation needs to be performed again, the data source party generates new retained numbers so that the product of the retained numbers and the other multipliers is equal to the updated private numbers. The data source party calculates the intermediate values of the subtasks by itself, sends the intermediate values of the subtasks to the fusion node, the fusion node replaces the saved intermediate values with the newly received intermediate values, recalculates the values of each product term, and uses the new values of the product terms to recalculate the value of the polynomial fitting function, which is the result of the privacy calculation after the update of the private numbers.

6. A multi-source data fusion platform based on privacy computing according to claim 1 or 2, characterized in that, the fusion node stores the intermediate values of the received subtasks. After one privacy calculation, if there are new data source parties that want to incorporate their own private numbers into the privacy calculation, the new data source parties split their own private numbers into several addends, randomly send the several addends to several data source parties that have participated in the privacy calculation before. The several data source parties that receive the addends add the addends to their own private numbers to obtain the updated private numbers, generate new retained numbers so that the product of the retained numbers and the other multipliers is equal to the updated private numbers. The data source party calculates the intermediate values of the subtasks by itself, sends the intermediate values of the subtasks to the fusion node, the fusion node replaces the saved intermediate values with the newly received intermediate values, recalculates the values of each product term, and uses the new values of the product terms to recalculate the value of the polynomial fitting function, which is the result of the privacy calculation after the private numbers of the new data source parties are added.

7. A multi-source data fusion platform based on privacy computing according to claim 1 or 2, characterized in that, when the service node establishes the polynomial fitting function of the objective function, it performs the following steps: the service node requests the data source party for the value ranges of each private number in the objective function; within the value ranges of each private number, uniformly generate sample numbers to form multiple values of the private numbers, and randomly combine the sample numbers of the private numbers into a group, which is recorded as a value group; Substitute the value group into the objective function to obtain the output of the objective function, and use the said output as a label to mark the said value group as sample data.

Citation Information

Patent Citations

  • Target data sending method and privacy computing platform

    CN112699391A

  • Task processing method and device

    CN105893497A

  • Data processing method, system and device based on node group and medium

    CN111931253A