Method and system for multi-party joint data filtering based on privacy protection
By breaking down the overall filtering conditions into sub-filtering conditions and performing logical operations between participants and collaborators, the problems of communication costs and privacy leaks in multi-party data filtering are solved, and secure and efficient determination of data filtering results is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2022-12-02
- Publication Date
- 2026-04-17
AI Technical Summary
In the process of multi-party joint data filtering, existing technologies require data from all parties to be aggregated and filtered by the collaborating party, which increases communication costs and poses a risk of privacy leakage.
The collaborating party breaks down the overall filtering conditions into multiple sub-filtering conditions and distributes them to the corresponding participants. The participants determine the logical result set based on the sub-filtering conditions. The collaborating party then summarizes these results to obtain the filtering result for the common sample set. Alternatively, the collaborating party performs de-identification processing and distributes the sub-filtering conditions and de-identification expressions. The participants then summarize the logical result set and perform calculations to determine the filtering result.
It enables multi-party data filtering without leaking participant data, reducing communication costs and ensuring data security.
Smart Images

Figure CN116089997B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data security through one or more embodiments, and more particularly to a method and system for multi-party collaborative data filtering based on privacy protection. Background Technology
[0002] In most cases, data filtering requires collaboration among multiple parties to perform joint calculations.
[0003] In traditional technologies, when filtering data from multiple parties, the data needs to be aggregated and submitted to a collaborating party, which then performs the filtering based on the specified criteria. However, this method requires all participating parties to upload data, which not only increases communication costs but also raises privacy concerns.
[0004] Therefore, there is an urgent need for a solution that enables multiple parties to jointly filter data while ensuring that the data of all parties is not leaked. Summary of the Invention
[0005] This specification describes one or more embodiments of a method and system for multi-party collaborative data filtering based on privacy protection, which can effectively ensure the security of data of each participating party.
[0006] Firstly, a method for data filtering based on privacy protection through multi-party collaboration is provided. The multi-party collaboration includes a collaborating party and n participating parties, each possessing a shared sample set and distinct features of each shared sample. The method includes:
[0007] The collaborating party breaks down the total filtering conditions into m sub-filtering conditions and the logical operations between them, where a single sub-filtering condition is used to limit the value of a feature item.
[0008] The collaborating party sends each sub-filter condition to the participating party that has the feature items defined by that sub-filter condition;
[0009] A participant that receives at least one sub-filter condition determines at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and provides the logical result set to the collaborating party;
[0010] The collaborating party performs the logical operation based on the m logical result sets returned by the participating parties to obtain the target logical value of each shared sample corresponding to the total filtering condition;
[0011] The collaborating party obtains the filtering result for the shared sample set based on the target logical value, and provides the filtering result to the n participating parties.
[0012] Secondly, a privacy-preserving multi-party collaborative data filtering method is provided, wherein the multi-party includes a collaborating party and n participating parties, the n participating parties having a shared sample set and each having different feature terms of each shared sample; the method is executed by the collaborating party and includes:
[0013] The total filtering condition is broken down into m sub-filtering conditions and the logical operations between them. Each sub-filtering condition is used to limit the value of a feature term.
[0014] Each sub-filter condition is sent to the participant that has the feature items defined by the sub-filter condition, so that the participant determines its logical result set corresponding to the received sub-filter condition for the common sample set and provides it to the collaborating party.
[0015] Based on the m logical result sets returned by the participants, the logical operation is performed to obtain the target logical value of each common sample corresponding to the total filtering condition;
[0016] Based on the target logical value, obtain the filtering result for the shared sample set, and provide the filtering result to the n participants.
[0017] Thirdly, a method for data filtering based on privacy protection through multi-party collaboration is provided. The multi-party collaboration includes a collaborating party and n participating parties, each possessing a shared sample set and distinct features of each shared sample. The method includes:
[0018] The collaborating party breaks down the total filtering conditions into m sub-filtering conditions and the logical operations between them, where a single sub-filtering condition is used to limit the value of a feature item.
[0019] The collaborating party performs desensitization processing on the total filtering conditions to obtain a desensitized expression, which at least indicates the logical operations between the m sub-filtering conditions;
[0020] The collaborating party sends each sub-filter condition to the participating party that has the feature items defined by the sub-filter condition, and provides the de-identification expression to each participating party;
[0021] Upon receiving at least one sub-filter condition, participant i determines at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and provides the logical result set to other participants;
[0022] The participant i summarizes its determined logical result set and the logical result sets received from other participants into m logical result sets, and performs logical operations in the desensitization expression based on them to obtain the target logical value of each common sample corresponding to the total filtering condition.
[0023] The participant i obtains the filtering result for the shared sample set based on the target logical value.
[0024] Fourthly, a system for multi-party collaborative data filtering based on privacy protection is provided, including a collaborating party and n participating parties, wherein the n participating parties have a common sample set and each has different features of each common sample;
[0025] The collaborating party is used to decompose the total filtering conditions into m sub-filtering conditions and the logical operations between them, wherein a single sub-filtering condition is used to limit the value of a feature item.
[0026] The collaborating party is also used to send each sub-filter condition to the participating party that has the feature items defined by the sub-filter condition;
[0027] A participant that receives at least one sub-filter condition is configured to determine at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and provide the logical result set to the collaborating party;
[0028] The collaborating party is also used to perform the logical operation based on the m logical result sets returned by the participating parties to obtain the target logical value of each shared sample corresponding to the total filtering condition;
[0029] The collaborating party is also configured to obtain the filtering result for the shared sample set based on the target logical value, and provide the filtering result to the n participating parties.
[0030] Fifthly, a device for multi-party collaborative data filtering based on privacy protection is provided. The multi-party system includes a collaborating party and n participating parties, each of which has a shared sample set and possesses different features of each shared sample. The device is disposed within the collaborating party and includes:
[0031] The decomposition unit is used to decompose the total filtering conditions into m sub-filtering conditions and the logical operations between them, wherein a single sub-filtering condition is used to limit the value of a feature term.
[0032] The sending unit is used to send each sub-filter condition to the participant that has the feature items defined by the sub-filter condition, so that the participant determines its logical result set corresponding to the received sub-filter condition for the common sample set and provides it to the collaborating party.
[0033] The operation unit is used to perform the logical operation based on the m logical result sets returned by the participants to obtain the target logical value of each common sample corresponding to the total filtering condition;
[0034] The acquisition unit is used to acquire the filtering result for the common sample set based on the target logical value, and to provide the filtering result to the n participants.
[0035] Sixthly, a system for data filtering based on privacy protection through multi-party collaboration is provided, wherein the multi-party collaboration includes a collaborating party and n participating parties, wherein the n participating parties have a common sample set and each has different features of each common sample;
[0036] The collaborating party is used to decompose the total filtering conditions into m sub-filtering conditions and the logical operations between them, wherein a single sub-filtering condition is used to limit the value of a feature item.
[0037] The collaborating party is also used to perform desensitization processing on the total filtering conditions to obtain a desensitized expression, wherein at least the logical operations between the m sub-filtering conditions are indicated.
[0038] The collaborating party is also used to send each sub-filter condition to the participating party that has the feature item defined by the sub-filter condition, and to provide the desensitization expression to each participating party;
[0039] A participant i that receives at least one sub-filter condition is configured to determine at least one logical result set corresponding to the at least one sub-filter condition for the common sample set, and provide the logical result set to other participants;
[0040] The participant i is also used to summarize the logical result set it has determined and the logical result set received from other participants into m logical result sets, and perform logical operations in the desensitization expression based on them to obtain the target logical value of each common sample corresponding to the total filtering condition.
[0041] The participant i is also used to obtain the filtering result for the common sample set based on the target logical value.
[0042] In a seventh aspect, a computer storage medium is provided having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the methods of the first, second, or third aspect.
[0043] Eighthly, a computing device is provided, including a memory and a processor, wherein executable code is stored in the memory, and the processor, when executing the executable code, implements the method of the first, second, or third aspect.
[0044] This specification provides a method and system for multi-party collaborative data filtering based on privacy protection, through one or more embodiments. The collaborating parties decompose the overall filtering conditions into multiple sub-filtering conditions and distribute these sub-filtering conditions to the corresponding participating parties. Then, each participating party uploads its logical result set determined for the received sub-filtering conditions to the collaborating party. The collaborating party then aggregates and calculates the logical result sets uploaded by the participating parties to obtain the filtering result for a shared sample set. Therefore, in the embodiments of this specification, participating parties only need to upload the logical result set to the collaborating party, without uploading the original data, to achieve multi-party data filtering. This ensures the security of the participating parties' data and also saves communication costs. Attached Figure Description
[0045] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a schematic diagram illustrating an implementation scenario of one embodiment disclosed in this specification;
[0047] Figure 2 This is a schematic diagram illustrating an implementation scenario of another embodiment disclosed in this specification;
[0048] Figure 3 This diagram illustrates an interaction method for privacy-preserving multi-party collaborative data filtering according to one embodiment.
[0049] Figure 4 This diagram illustrates an interaction method for privacy-preserving multi-party collaborative data filtering according to one embodiment.
[0050] Figure 5 This diagram illustrates a system for privacy-preserving multi-party collaborative data filtering according to one embodiment.
[0051] Figure 6 A schematic diagram of an apparatus for privacy-preserving multi-party collaborative data filtering according to one embodiment is shown;
[0052] Figure 7 This diagram illustrates a system for privacy-preserving multi-party collaborative data filtering according to one embodiment. Detailed Implementation
[0053] The solution provided in this specification will now be described with reference to the accompanying drawings.
[0054] Figure 1This is a schematic diagram illustrating an implementation scenario of one of the embodiments disclosed in this specification. Figure 1 In this context, the collaborating party and the n participating parties can be implemented as any device, platform, server, or device cluster with computing and processing capabilities. The n participating parties share a common sample set and each possesses distinct features from the shared samples. Here, n is a positive integer.
[0055] In one example, the common sample set of n participants can be obtained by the n participants through performing Privacy Set Intersection (PSI).
[0056] Figure 1 In this process, the collaborating party can break down the total filtering condition into m sub-filtering conditions, where m is a positive integer. Then, the collaborating party can send each of the m sub-filtering conditions to the participants who possess the features defined by that sub-filtering condition. Next, each participant receiving a sub-filtering condition determines its corresponding logical result set for the shared sample set and provides this logical result set to the collaborating party. Finally, the collaborating party aggregates the logical result sets returned by the participants and, based on the aggregated result, determines the filtering result for the shared sample set and provides this filtering result to the n participants.
[0057] As can be seen, in the above scheme, the collaborating party first breaks down the sub-filtering conditions and distributes them to the participating parties. Then, the collaborating party summarizes the logical result set determined by the participating parties and then determines the filtering result of the common sample set.
[0058] Of course, in practical applications, the aforementioned collaborators can also be selected from among the participating parties; that is, the selected participating party (also called the target participating party) serves as both a collaborator and a participating party. It should be understood that when the collaborator is selected from among the participating parties, this collaborator can also possess some feature items from each shared sample. Therefore, after decomposing and obtaining m sub-filtering conditions, it can retain the defined feature items as sub-filtering conditions for its own feature items and determine the corresponding logical result set. Then, based on the logical result set determined by the collaborator and the received logical result set, the filtering result for the shared sample set is determined.
[0059] Figure 2 This is a schematic diagram illustrating an implementation scenario of another embodiment disclosed in this specification. Figure 2 In this context, the collaborating party and the n participating parties can be implemented as any device, platform, server, or device cluster with computing and processing capabilities. The n participating parties share a common sample set and each possesses different characteristics of each common sample within that set.
[0060] Figure 2In this process, collaborating parties can break down the overall filtering condition into m sub-filtering conditions and perform de-identification processing on the overall filtering condition to obtain the corresponding de-identified expression. Then, the collaborating party can send each of the m sub-filtering conditions to the participants possessing the features defined by that sub-filtering condition, and provide the de-identified expression to n participants. Next, the participants receiving the sub-filtering conditions determine their corresponding logical result set for the shared sample set, and provide this logical result set to the other participants. Finally, based on their determined logical result set and the received logical result set, the participants calculate the value of the de-identified expression, thereby determining the filtering result for the shared sample set.
[0061] As can be seen, in this scheme, the collaborating party breaks down the sub-filtering conditions and distributes them to the corresponding participants. Then, the participants aggregate the m logical result sets to determine the filtering results for the common sample set. The following uses... Figure 1 The following is a detailed explanation of the solution, using the illustrated implementation scenario as an example.
[0062] Similarly, in this embodiment, the collaborator can also be selected from among the participants, meaning the selected participant (also called the target participant) serves as both a collaborator and a participant. It should be understood that, in the case where the collaborator is selected from among the participants, this collaborator may also possess some feature items from each shared sample. Therefore, after decomposing and obtaining m sub-filtering conditions, it can retain the defined feature items as sub-filtering conditions for its own feature items, determine the corresponding logical result set, and provide this logical result set to other participants. Then, based on the logical result set determined by the collaborator and the received logical result set, the value of the desensitization expression is calculated, thereby determining the filtering result for the shared sample set.
[0063] Figure 3 This diagram illustrates an interaction diagram of a privacy-preserving multi-party collaborative data filtering method according to one embodiment. It should be noted that because the interaction processes of each participating party and collaborating party are similar, therefore... Figure 3 The diagram mainly illustrates the interaction steps between any participant (referred to as the first participant for ease of description) and the collaborator. The interaction steps between other participants and the collaborator can be found in the interaction steps between the first participant and the collaborator.
[0064] like Figure 3 As shown, the first participant and the collaborating party interact through the following steps:
[0065] In step S302, the collaborating party decomposes the total filtering condition into m sub-filtering conditions and the logical operations between them, where a single sub-filtering condition is used to limit the value of a feature term.
[0066] In one embodiment, the above-mentioned total filtering condition can be expressed as an expression, the result of which is a definite value.
[0067] For example, suppose there are n participants, including party A and party B, who share a common sample set. Party A possesses the feature terms for each common sample: col_1, col_2, and col_3, while party B possesses the feature terms for each common sample: col_a, col_b, and col_c. Furthermore, assume that the values of the feature terms possessed by party A are as shown in Table 1.
[0068] Table 1
[0069]
[0070]
[0071] Furthermore, it is assumed that the values of each feature possessed by Party B are as shown in Table 2.
[0072] Table 2
[0073] Sample identification col_a col_b col_c k3 30 bc e3 k4 40 bd e4 k5 50 be e5 k6 60 bf e6 k7 70 bg e7 k8 80 bh e3 k9 90 bi e4 k10 100 bj e5 k11 110 bk e6 k12 120 bl e7
[0074] It should be understood that a complete shared sample can be obtained based on the values of the features corresponding to a certain sample identifier possessed by Party A and the values of the features corresponding to that sample identifier possessed by Party B. For example, by concatenating the second row of Table 1 and Table 2, a shared sample with sample identifier k3 can be obtained.
[0075] In the above example, the total filter condition could be, for example, the following expression: ((col_2>='e'and col_2<='k')or(col_b>='bi'and col_b<='bq'))and(col_3>=18)and(col_c>='e3').
[0076] Given the overall filtering condition as shown in the expression above, breaking down the overall filtering condition into m sub-filtering conditions means decomposing the expression into several atomic predicates. Here, a predicate refers to an expression that returns TRUE (hereinafter represented as 1), FALSE (hereinafter represented as 0), or UNKNOWN.
[0077] For example, the expression above can be broken down into the following atomic predicates:
[0078] co l_2>='e'
[0079] co l_2<='k'
[0080] co l_b>='bi'
[0081] co l_b<='bq'
[0082] co l_3>=18
[0083] co l_c>='e3'
[0084] Furthermore, the above expression can be broken down into logical operators: "and" and "or".
[0085] It should be noted that the above general filtering conditions are only an illustrative example. In practical applications, the complex expressions described above may also include aggregate functions such as sum() or count(), etc., which is not limited in this specification.
[0086] In one embodiment, after obtaining m sub-filter conditions, a corresponding unique identifier (hereinafter also referred to as a predefined identifier) can be assigned to each of them. For example, the unique identifier assigned to each of the above atomic predicates can be as follows:
[0087]
[0088] If a unique identifier is also assigned to the atomic predicate, the above expression can be rewritten. For example, the rewritten expression can be:
[0089] ((__pred_0__and__pred_1__)or(__pred_2__and__pred_3__))and__pred_4__and__pred_5__
[0090] In step S304, the collaborating party sends each sub-filter condition to the participating party that has the feature items defined by that sub-filter condition.
[0091] In the aforementioned example, since Party A possesses the features: col_1, col_2, and col_3, and Party B possesses the features: col_a, col_b, and col_c, the collaborating party can send the atomic predicates: "col_2>='e'", "col_2<='k'", and "col_3>=18" to Party A, and send the atomic predicates: "col_b>='bi'", "col_b<='bq'", and "col_c>='e3'" to Party B.
[0092] It should be understood that when collaborating parties set corresponding unique identifiers for each sub-filter condition, they can also send the unique identifiers corresponding to the sub-filter conditions to the participating parties. For example, “__pred_0__”, “__pred_1__”, and “__pred_4__” can be sent to party A, and “__pred_2__”, “__pred_3__”, and “__pred_5__” can be sent to party B.
[0093] Step S306: The participant who receives at least one sub-filter condition determines at least one logical result set corresponding to at least one sub-filter condition for the common sample set, and provides the logical result set to the collaborating party.
[0094] It should be understood that the participants in the receiving sub-filtering conditions described in this specification can be some or all of the n participants. For example, in the aforementioned example, when the n participants only include party A and party B, then the participants in the receiving sub-filtering conditions are all participants. However, when the n participants also include party C, and party C has a common feature term: co l_x, then the participants in the receiving sub-filtering conditions are only some of the participants, namely parties A and B. This is because the feature term not limited by the sub-filtering conditions is co l_x.
[0095] Step 306 specifically involves, for any first sub-filtering condition among the at least one sub-filtering condition mentioned above, setting the logical value corresponding to each common sample to 1 or 0 based on whether the corresponding feature of each common sample in the common sample set satisfies the first sub-filtering condition, thereby obtaining a first logical result set corresponding to the first sub-filtering condition. In other words, the first logical result set includes each logical value of each common sample corresponding to the first sub-filtering condition.
[0096] Taking party A as an example, and the first sub-filter condition as "co l_2>='e'", since only the shared samples identified by k3 and k4 in Table 1 do not satisfy the first sub-filter condition, the logical value of the shared samples identified by k3 and k4 corresponding to the first sub-filter condition is set to 0, and all other logical values are set to 1. That is to say, the first logical result set corresponding to the first sub-filter condition includes two 0s and eight 1s.
[0097] Similarly, the logical result sets corresponding to "col_2<='k'" and "col_3>=18" can be determined. Finally, the three logical result sets determined for the three sub-filtering conditions received by Party A can be shown in Table 3.
[0098] Table 3
[0099] Sample identification Logical Result Set 1 Logical Result Set 2 Logical Result Set 3 k3 0 1 1 k4 0 1 1 k5 1 1 1 k6 1 1 1 k7 1 1 1 k8 1 1 1 k9 1 1 1 k10 1 1 1 k11 1 1 1 k12 1 0 1
[0100] In Table 3, logical result set 1 corresponds to the sub-filter condition: "co l_2>='e'" (i.e., __pred_0__), logical result set 2 corresponds to the sub-filter condition: "co l_2<='k'" (i.e., __pred_1__), and logical result set 3 corresponds to the sub-filter condition: "co l_3>=18" (i.e., __pred_4__).
[0101] In addition, for the three sub-filtering conditions received by Party B, the corresponding three logical result sets can also be determined, as detailed in Table 4.
[0102] Table 4
[0103] Sample identification Logical Result Set 1 Logical Result Set 2 Logical Result Set 3 k3 0 1 1 k4 0 1 1 k5 0 1 1 k6 0 1 1 k7 0 1 1 k8 0 1 1 k9 1 1 1 k10 1 1 1 k11 1 1 1 k12 1 1 1
[0104] In Table 4, logical result set 1 corresponds to the sub-filter condition: "co l_b>='bi'" (i.e., __pred_2__), logical result set 2 corresponds to the sub-filter condition: "co l_b<='bq'" (i.e., __pred_3__), and logical result set 3 corresponds to the sub-filter condition: "co l_c>='e3'" (i.e., __pred_5__).
[0105] After obtaining at least one logical result set corresponding to at least one sub-filter condition received by a participating party, the participating party can provide each logical result set and its corresponding sub-filter condition to the collaborating party. It should be understood that if the collaborating party also sets a unique identifier for each sub-filter condition and sends this unique identifier to the corresponding participating party, the sub-filter condition can also be replaced with the corresponding unique identifier. Alternatively, both the sub-filter condition and the unique identifier can be provided to the collaborating party; this specification does not limit this approach.
[0106] Of course, in practical applications, to further ensure data security, participants can perform symmetric or homomorphic encryption on each logical value in the logical result set to obtain encrypted logical values. Based on these encrypted logical values, an encrypted result set can be formed, which the participants can then provide to their collaborators.
[0107] In addition, participating parties can also provide the collaborating parties with the anonymous identifiers of the shared samples corresponding to each logical value in the logical result set. These anonymous identifiers can be obtained by anonymizing the sample identifiers of the corresponding shared samples using a target algorithm (e.g., a hash algorithm) agreed upon by the n participating parties.
[0108] It should be understood that when the target algorithm is an irreversible algorithm, after obtaining the anonymous identifier corresponding to the sample identifier, the participants can record the correspondence between the sample identifier and the anonymous identifier for subsequent use in restoring the anonymous identifier.
[0109] In step S308, the collaborating party performs logical operations based on the m logical result sets returned by the participating parties to obtain the target logical value of each shared sample corresponding to the overall filtering condition.
[0110] It should be understood that since each logical result set corresponds one-to-one with a sub-filter condition, the collaborating party can receive m logical result sets.
[0111] In one embodiment, the collaborating parties can organize the m logical results that correspond to the m logical values of the same common sample together, and perform logical operations between the corresponding m sub-filtering conditions on the m logical values to obtain the target logical value of each common sample corresponding to the overall filtering condition.
[0112] In another embodiment, the collaborating parties arrange the m logical result sets into a numerical array. Each row of this array corresponds to one of the m logical values of a common sample, and each column corresponds to a sub-filtering condition. For each column of the numerical array, the logical values in that column are concatenated sequentially to obtain numerical vectors corresponding to each sub-filtering condition. Logical operations are then performed on each numerical vector between the corresponding sub-filtering conditions to obtain the target logical value for each common sample corresponding to the overall filtering condition.
[0113] It should be understood that if the collaborating party receives an encrypted result set sent by the participating party, for example, an encrypted result set obtained through symmetric encryption, then the collaborating party can first decrypt each encrypted logical value in the encrypted result set. For example, it can decrypt using the decryption key corresponding to the symmetric encryption. Then, it can perform logical operations on the m decrypted logical result sets.
[0114] If the above encrypted result set is obtained through homomorphic encryption, then the same logical operation as the corresponding m logical result sets is directly performed on the m encrypted result sets, and then the target logical value is obtained through the decryption operation result.
[0115] In the two embodiments described above, the collaborating parties can organize and obtain the m logical values of each common sample based on the anonymous identifier (or sample identifier) corresponding to each logical value in the m logical result sets.
[0116] It should be noted that the method for obtaining the target logic value in the other embodiment described above can also be called a vectorized calculation method. This vectorized calculation method can greatly improve the calculation efficiency of the target logic value.
[0117] In conjunction with the foregoing examples, the numerical array obtained through the other embodiment described above can be shown in Table 5.
[0118] Table 5
[0119]
[0120]
[0121] Here, k3'-k12' are the anonymization identifiers of each shared sample. Furthermore, the order of the logical result sets can be determined based on the order of the corresponding sub-filtering conditions in the expression.
[0122] For example, in the example above, the sub-filter condition "co l_2>='e'" (that is, __pred_0__) is arranged in the first position, so the corresponding logical result set is arranged in the first position; the sub-filter condition "co l_c>='e3'" (that is, __pred_5__) is arranged in the last position, so the corresponding logical result set is arranged in the last position, and so on.
[0123] For each logical value in Table 5, after concatenating the logical values column by column, we can obtain the following 6 numerical vectors: [0011111111], [1111111110], [0000001111], [1111111111], [1111111111], and [1111111111]. Replacing the unique identifiers in the modified expression with these 6 numerical vectors, we get: ([0011111111]and[111111110])or([0000001111]and[1111111111])and[1111111111]and[1111111111]. After performing the above logical operation on these 6 numerical vectors, we get the result vector: 0011111111.
[0124] It should be understood that one element value in the above result vector represents the target logical value of a common sample corresponding to the overall filtering condition.
[0125] As can be seen from the above, the vectorized calculation method described above can combine 10 logical operations into one logical operation, thereby greatly improving the computational efficiency.
[0126] For the aforementioned example, the target logical values of the total filtering conditions for each shared sample can be shown in Table 6.
[0127] Table 6
[0128]
[0129]
[0130] It should be understood that the target logical values in Table 6 can be obtained by performing logical operations on the m logical values of each common sample organized together. That is, by performing logical operations on the 6 logical values in each row of Table 5. Alternatively, they can be calculated using a vectorized calculation method, which involves concatenating the logical values in each column of Table 5 into a numerical vector and then performing logical operations on the 6 numerical vectors.
[0131] In step S310, the collaborating party obtains the filtering results for the shared sample set based on the target logical value, and provides the filtering results to n participating parties.
[0132] Specifically, target samples with the first logical value can be selected from the shared sample set. These target samples are then used as the filtering result of the shared sample set.
[0133] In one embodiment, the first value is 1.
[0134] For example, in the example above, the filtering result can be each target sample corresponding to k5'-k12'.
[0135] Furthermore, the collaborating party's provision of the filtering results to the n participants may include: the collaborating party providing the anonymized identifiers corresponding to each target sample to the n participants. Thus, each participant, based on the anonymized identifiers, determines the target samples from the shared sample set and, based on these target samples, performs target computation jointly with other participants. This target computation could be, for example, longitudinal joint learning, PSI, etc.
[0136] Specifically, the participants can first restore the received anonymous identifier to a sample identifier, and then determine each target sample from the shared sample set based on the sample identifier.
[0137] Here, restoring the anonymous identifier to the sample identifier can be done by querying the correspondence between the sample identifier and the anonymous identifier recorded in step 306.
[0138] In summary, the privacy-preserving multi-party collaborative data filtering method provided in this specification mainly consists of the following two stages: First, participating parties determine the logical result set of the shared sample set corresponding to the sub-filtering conditions they receive, and provide the logical result set to collaborating parties. Second, collaborating parties perform logical operations based on the received logical result set to obtain the target logical value of the shared sample set corresponding to the overall filtering condition, thereby determining the filtering result of the shared sample set. Through these two stages, it can be ensured that the data of each participating party does not leave the domain, and the amount of communication can be reduced. That is, it can balance the factors of communication cost and data security.
[0139] The following is another example Figure 2 The following is a detailed explanation of the solution, using the illustrated implementation scenario as an example.
[0140] Figure 4 This diagram illustrates an interaction diagram of a privacy-preserving multi-party collaborative data filtering method according to one embodiment. It should be noted that because the interaction processes of each participating party and collaborating party are similar, therefore... Figure 4 The diagram mainly illustrates the interaction steps between any participant (referred to as the first participant for ease of description) and the collaborator. The interaction steps between other participants and the collaborator can be found in the interaction steps between the first participant and the collaborator.
[0141] like Figure 4 As shown, the first participant and the collaborating party interact through the following steps:
[0142] In step S402, the collaborating party decomposes the total filtering condition into m sub-filtering conditions and the logical operations between them, where a single sub-filtering condition is used to limit the value of a feature term.
[0143] For a detailed description of step S402, please refer to step S302, which will not be repeated here.
[0144] In step S404, the collaborating party performs desensitization processing on the total filtering conditions to obtain a desensitized expression, which indicates at least the logical operations between the m sub-filtering conditions.
[0145] In one embodiment, the collaborating party may perform desensitization processing on the overall filtering conditions by replacing m sub-filter conditions with corresponding m predetermined identifiers, and forming a desensitization expression based on the logical operations between the m predetermined identifiers and their corresponding m sub-filter conditions.
[0146] For example, suppose the overall filter condition is the following expression: ((col_2>='e'and col_2<='k')or(col_b>='bi'and col_b<='bq'))and(col_3>=18)and(col_c>='e3'), then the decomposed sub-filter conditions can be as follows:
[0147] co l_2>='e'
[0148] co l_2<='k'
[0149] co l_b>='bi'
[0150] co l_b<='bq'
[0151] co l_3>=18
[0152] co l_c>='e3'
[0153] Furthermore, it is assumed that the predefined identifiers corresponding to each sub-filter condition are as follows:
[0154]
[0155]
[0156] The resulting desensitization expression can then be: ((__pred_0__and__pred_1__)or(__pred_2__and__pred_3__))and__pred_4__and__pred_5__
[0157] It should be understood that this is called a desensitized expression because it does not contain any information related to the individual participants.
[0158] In step S406, the collaborating party sends each sub-filter condition to the participating party that has the feature items defined by the sub-filter condition, and provides the de-identification expression to each participating party.
[0159] It should be understood that "each participant" here refers to all participants. Therefore, participants who receive the sub-filter conditions will also receive the de-identification expression. Participants who do not receive the sub-filter conditions (such as Party C) will only receive the de-identification expression.
[0160] For a detailed explanation of step S406, please refer to step S304; this specification will not repeat it here.
[0161] Step S408: The first participant, having received at least one sub-filter condition, determines at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and provides the logical result set to the other participants.
[0162] The other participants here refer to the n-1 participants other than the first participant.
[0163] Specifically, for any first sub-filtering condition among the above at least one sub-filtering conditions, based on whether the corresponding features of each common sample in the common sample set satisfy the first sub-filtering condition, the logical value corresponding to each common sample is set to 1 or 0, resulting in a first logical result set corresponding to the first sub-filtering condition. In other words, the first logical result set includes each logical value of each common sample corresponding to the first sub-filtering condition.
[0164] After obtaining at least one logical result set corresponding to at least one sub-filter condition received by the first participant, the first participant may provide the unique identifier (also known as the predetermined identifier) of the at least one logical result set and its corresponding sub-filter condition to other participants.
[0165] Of course, in practical applications, to further ensure data security, the first participant can perform symmetric or homomorphic encryption on each logical value in the logical result set to obtain encrypted logical values. Based on these encrypted logical values, an encrypted result set can be formed, which the first participant can then provide to other participants.
[0166] Furthermore, the first participant can also provide the other participants with the anonymous identifiers of the shared samples corresponding to each logical value in the logical result set. These anonymous identifiers can be obtained by the first participant using a target algorithm (e.g., a hash algorithm) agreed upon with the other participants to anonymize the sample identifiers.
[0167] It should be understood that when the target algorithm is an irreversible algorithm, after obtaining the anonymous identifier corresponding to the sample identifier, the first participant can record the correspondence between the sample identifier and the anonymous identifier for subsequent use in restoring the anonymous identifier.
[0168] In step S410, the first participant summarizes its determined logical result set and the logical result sets received from other participants into m logical result sets, and performs logical operations in the desensitization expression based on them to obtain the target logical value of each common sample corresponding to the overall filtering condition.
[0169] Other participants here refer to those participants other than the first participant among the n participants who receive the sub-filtering conditions.
[0170] Here, the first participant can receive not only the logical result sets sent by other participants, but also the predetermined identifiers of the sub-filtering conditions corresponding to those logical result sets. Furthermore, since each logical result set corresponds one-to-one with a sub-filtering condition, the first participant can receive m logical result sets.
[0171] The above-mentioned aggregation into m logical result sets may specifically include: the first participant, based on the predetermined identifier of its determined logical result set and the predetermined identifier of the received logical result set, aggregates its determined logical result set and the received logical result set to obtain m logical result sets.
[0172] Here, the method by which the first participant obtains the target logical value of each shared sample corresponding to the overall filtering condition can refer to the two embodiments described in step S308, which will not be repeated here.
[0173] Furthermore, if the first participant receives an encrypted result set from other participants, such as an encrypted result set obtained through symmetric encryption, then the first participant can first decrypt each encrypted logical value in the encrypted result set. For example, it can decrypt using the decryption key corresponding to the symmetric encryption. Then, it can perform the logical operations in the de-identified expression on the m decrypted logical result sets.
[0174] If the above encrypted result set is obtained through homomorphic encryption, then the same logical operation as the corresponding m logical result sets is directly performed on the m encrypted result sets, and then the target logical value is obtained through the decryption operation result.
[0175] In step S412, the first participant obtains the filtering results for the shared sample set based on the target logical value.
[0176] Specifically, the first participant can select target samples from the shared sample set whose corresponding target logical value is the first value. These target samples are then used as the filtering result of the shared sample set.
[0177] After the n participants obtain their respective filtering results, they can perform vertical joint learning or PSI based on these filtering results.
[0178] In summary, the privacy-preserving multi-party collaborative data filtering method provided in this specification involves a collaborating party breaking down sub-filtering conditions and distributing them to the corresponding participating parties. The participating parties then aggregate m logical result sets to determine the filtering result for the shared sample set. This ensures that the data of each participating party remains within its domain, thereby achieving data privacy protection.
[0179] Corresponding to the above-described method for multi-party collaborative data filtering based on privacy protection, one embodiment of this specification also provides a system for multi-party collaborative data filtering based on privacy protection, such as... Figure 5 As shown, the system includes a collaborator 502 and n participants 504. The n participants 504 have a common sample set and each has different features of each common sample.
[0180] Collaborator 502 is used to decompose the total filtering condition into m sub-filtering conditions and the logical operations between them, where a single sub-filtering condition is used to limit the value of a feature term.
[0181] Collaborator 502 is also used to send each sub-filter condition to participant 504 who has the feature items defined by the sub-filter condition.
[0182] Participant 504, having received at least one sub-filter condition, determines at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and provides the logical result set to collaborator 502.
[0183] Participant 504 is specifically used for:
[0184] For any first sub-filtering condition among at least one sub-filtering condition, based on whether the corresponding features of each common sample in the common sample set satisfy the first sub-filtering condition, the logical value corresponding to each common sample is set to 1 or 0, thus obtaining the first logical result set corresponding to the first sub-filtering condition.
[0185] Collaborator 502 is also used to perform logical operations based on the m logical result sets returned by participant 504 to obtain the target logical value of each shared sample corresponding to the total filtering condition.
[0186] Collaborator 502 is specifically used for:
[0187] Arrange the m logical result sets into a numerical array, where each row of the numerical array corresponds to the m logical values of a common sample, and each column corresponds to a sub-filter condition.
[0188] For each column in the numerical array, the logical values located in that column are concatenated in sequence to obtain the numerical vectors corresponding to each sub-filter condition;
[0189] Perform logical operations on each numerical vector to obtain the target logical value of each common sample corresponding to the overall filtering condition.
[0190] Collaborator 502 is also used to obtain the filtering results for the shared sample set based on the target logical value, and to provide the filtering results to n participants 504.
[0191] Collaborator 502 is specifically used for:
[0192] From the common sample set, select the target samples whose corresponding target logical value is the first value;
[0193] The filtering results determine each target sample as a shared sample set.
[0194] In one embodiment, each of the above-mentioned shared samples has a corresponding anonymous identifier;
[0195] Collaborator 502 is specifically used to: provide the anonymity identifiers corresponding to each target sample to n participating parties 504;
[0196] Participant 504 is also used to identify each target sample from the shared sample set based on each anonymous identifier, and to perform target calculations in conjunction with other participants.
[0197] In one embodiment, each shared sample has a corresponding anonymity identifier, and participant 504 is further specifically used for:
[0198] Provide the collaborating party 502 with the anonymous identifiers of the logical result set and the corresponding common samples for each logical value;
[0199] Collaborator 502 is specifically used to: arrange m logical result sets into a numerical array based on each anonymous identifier.
[0200] In one embodiment, participant 504 is further specifically used for:
[0201] Symmetric encryption is performed on each logical value in the logical result set, and the resulting encrypted result set is provided to the collaborating party 502.
[0202] Collaborator 502 is also used to decrypt each encrypted logical value in the encrypted result set to obtain the logical result set.
[0203] In one embodiment, each shared sample has a corresponding anonymous identifier, and collaborator 502 is specifically used for:
[0204] Provide the anonymized identifiers corresponding to each target sample to n participants (504);
[0205] Participant 504 is also used to identify each target sample from the shared sample set based on each anonymous identifier, and to perform target calculations in conjunction with other participants.
[0206] The functions of each functional module of the device in the above embodiments of this specification can be implemented through the steps of the above method embodiments. Therefore, the specific working process of the device provided in one embodiment of this specification will not be repeated here.
[0207] This specification provides an embodiment of a privacy-preserving multi-party collaborative data filtering system that can ensure the security of the data of the participating parties and save communication costs.
[0208] Corresponding to the aforementioned privacy-preserving multi-party collaborative data filtering method, one embodiment of this specification also provides a privacy-preserving multi-party collaborative data filtering apparatus. The multi-party includes a collaborating party and n participating parties, each of the n participating parties having a shared sample set and possessing different features of each shared sample. The apparatus is disposed within the collaborating party, such as... Figure 6 As shown, the device may include:
[0209] The decomposition unit 602 is used to decompose the total filtering conditions into m sub-filtering conditions and the logical operations between them, wherein a single sub-filtering condition is used to limit the value of a feature item.
[0210] The sending unit 604 is used to send each sub-filter condition to the participant that has the feature items defined by the sub-filter condition, so that the participant can determine its logical result set corresponding to the received sub-filter condition for the common sample set and provide it to the collaborating party.
[0211] The operation unit 606 is used to perform logical operations based on the m logical result sets returned by the participants to obtain the target logical value of each common sample corresponding to the total filtering condition.
[0212] The acquisition unit 608 is used to obtain the filtering results for the common sample set based on the target logical value, and provide the filtering results to n participants.
[0213] The acquisition unit 608 is specifically used for:
[0214] From the common sample set, select the target samples whose corresponding target logical value is the first value;
[0215] The filtering results determine each target sample as a shared sample set.
[0216] In one embodiment, the arithmetic unit 606 includes:
[0217] The arrangement submodule 6062 is used to arrange m logical result sets into a numerical array, where each row of the numerical array corresponds to the m logical values of a common sample, and each column corresponds to a sub-filter condition.
[0218] The splicing submodule 6064 is used to sequentially splice the logical values located in each column of the numerical array to obtain the numerical vectors corresponding to each sub-filter condition.
[0219] The logic operation submodule 6066 is used to perform logic operations on each numerical vector to obtain the target logic value of each common sample corresponding to the overall filtering condition.
[0220] In one embodiment, the above-described apparatus further includes:
[0221] The receiving unit 610 is used to receive from the participants the anonymous identifiers of the common samples corresponding to each logical value in the logical result set.
[0222] The 6062 layout submodule is specifically used for:
[0223] Based on each anonymous identifier, arrange the m logical result sets into a numerical array.
[0224] The functions of each functional module of the device in the above embodiments of this specification can be implemented through the steps of the above method embodiments. Therefore, the specific working process of the device provided in one embodiment of this specification will not be repeated here.
[0225] This specification provides an embodiment of a privacy-preserving multi-party collaborative data filtering apparatus that can ensure the security of the data of the participating parties and save communication costs.
[0226] Corresponding to the above-described method for multi-party collaborative data filtering based on privacy protection, one embodiment of this specification also provides a system for multi-party collaborative data filtering based on privacy protection, such as... Figure 7 As shown, the system includes a collaborator 702 and n participants 704. The n participants 704 have a common sample set and each has different features of each common sample.
[0227] Collaborator 702 is used to decompose the total filtering condition into m sub-filtering conditions and the logical operations between them, where a single sub-filtering condition is used to limit the value of a feature term.
[0228] Collaborator 702 is also used to desensitize the total filtering conditions to obtain a desensitized expression, which indicates at least the logical operations between the m sub-filtering conditions.
[0229] Collaborator 702 is specifically used to: replace the m sub-filter conditions with the corresponding m predetermined identifiers, and form a desensitization expression based on the logical operation between the m predetermined identifiers and their corresponding m sub-filter conditions.
[0230] Collaborator 702 is also used to send each sub-filter condition to participant 704 who has the feature items defined by the sub-filter condition, and to provide the desensitization expression to each participant 704.
[0231] Participant 704, having received at least one sub-filter condition, is used to determine at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and to provide the logical result set to other participants.
[0232] Participant 704 is also used to summarize the logical result set it has determined and the logical result set received from other participants into m logical result sets, and to perform logical operations in the above desensitization expression based on them to obtain the target logical value of each common sample corresponding to the total filtering condition.
[0233] Participant 704 is also used to obtain filtering results for the shared sample set based on the target logical value.
[0234] In one embodiment, the collaborating party 702 is further configured to provide each predetermined identifier corresponding to each sub-filter condition to the participating party 704.
[0235] In one embodiment, participant 704 is further configured to receive a predetermined identifier corresponding to the logical result set determined by other participants.
[0236] Participant 704 is also specifically used for:
[0237] Based on the predetermined identifier of the determined logical result set and the predetermined identifier of the received logical result set, the determined logical result set and the received logical result set are summarized to obtain m logical result sets.
[0238] The functions of each functional module of the device in the above embodiments of this specification can be implemented through the steps of the above method embodiments. Therefore, the specific working process of the device provided in one embodiment of this specification will not be repeated here.
[0239] This specification provides an embodiment of a privacy-preserving multi-party collaborative data filtering system that can ensure the security of the data of the participating parties and save communication costs.
[0240] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform a combination Figure 3 or Figure 4 The method described.
[0241] According to another embodiment, a computing device is also provided, including a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, it implements a combination... Figure 3 or Figure 4 The method described.
[0242] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system, apparatus, and device embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0243] The steps of the methods or algorithms described in conjunction with the disclosure in this specification can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, EPROM, EEPROM, registers, hard disk, external hard disk, CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a server. Of course, the processor and storage medium can also exist as discrete components in the server.
[0244] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in this invention can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium accessible to a general-purpose or special-purpose computer.
[0245] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0246] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this specification. It should be understood that the above description is only a specific embodiment of this specification and is not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this specification should be included within the scope of protection of this specification.
Claims
1. A method for data filtering based on privacy protection through multi-party collaboration, wherein the multi-party collaboration includes a collaborating party and n participating parties, the n participating parties having a shared sample set and each having different features of each shared sample; the method includes: The collaborating party breaks down the total filtering conditions into m sub-filtering conditions and the logical operations between them, where a single sub-filtering condition is used to limit the value of a feature item. The collaborating party sends each sub-filter condition to the participating party that has the feature items defined by that sub-filter condition; A participant that receives at least one sub-filter condition determines at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and provides the logical result set to the collaborating party; The collaborating party performs the logical operation based on the m logical result sets returned by the participating parties to obtain the target logical value of each shared sample corresponding to the total filtering condition; The collaborating party obtains the filtering result for the shared sample set based on the target logical value, and provides the filtering result to the n participating parties.
2. The method of claim 1, wherein, Determining at least one logical result set corresponding to the at least one sub-filter condition for the common sample set includes: For any first sub-filtering condition among the at least one sub-filtering conditions, based on whether the corresponding features of each common sample in the common sample set satisfy the first sub-filtering condition, the logical value corresponding to each common sample is set to 1 or 0, thereby obtaining a first logical result set corresponding to the first sub-filtering condition.
3. The method of claim 1, wherein, The collaborating party performs the logical operation based on the m logical result sets returned by the participating parties, including: The m logical result sets are arranged into a numerical array, where each row of the numerical array corresponds to the m logical values of a common sample, and each column corresponds to a sub-filtering condition. For each column in the numerical array, the logical values located in that column are concatenated in sequence to obtain the numerical vectors corresponding to each sub-filter condition; The logical operation is performed on each of the numerical vectors to obtain the target logical value of each common sample corresponding to the overall filtering condition.
4. The method of claim 3, wherein, Each shared sample has a corresponding anonymous identifier; providing the logical result set to the collaborating party includes: Provide the collaborating party with the anonymous identifiers of the common samples corresponding to each logical value in the logical result set; The step of arranging the m logical result sets into a numerical array includes: Based on the aforementioned anonymous identifiers, the m logical result sets are arranged into a numerical array.
5. The method according to claim 1, wherein, Providing the logical result set to the collaborating party includes: Symmetric encryption is performed on each logical value in the logical result set, and the resulting encrypted result set is provided to the collaborating party. The method further includes: The collaborating party decrypts each encrypted logical value in the encrypted result set to obtain the logical result set.
6. The method according to claim 1, wherein, The step of obtaining the filtering results for the shared sample set includes: From the shared sample set, select each target sample whose corresponding target logical value is the first value; The target samples are determined as the filtering results of the common sample set.
7. The method according to claim 6, wherein, Each shared sample has a corresponding anonymous identifier; Providing the filtering results to the n participants includes: Provide the anonymity identifiers corresponding to each target sample to the n participants; The method further includes: The participating parties determine the target samples from the shared sample set based on the anonymous identifiers, and perform target calculations in conjunction with other participating parties.
8. A method for data filtering based on privacy protection through multi-party collaboration, wherein the multi-party collaboration includes a collaborating party and n participating parties, wherein the n participating parties have a common sample set and each has different features of each common sample; The method is executed by the collaborating party and includes: The total filtering condition is broken down into m sub-filtering conditions and the logical operations between them. Each sub-filtering condition is used to limit the value of a feature term. Each sub-filter condition is sent to the participant that has the feature items defined by the sub-filter condition, so that the participant determines its logical result set corresponding to the received sub-filter condition for the common sample set and provides it to the collaborating party. Based on the m logical result sets returned by the participants, the logical operation is performed to obtain the target logical value of each common sample corresponding to the total filtering condition; Based on the target logical value, obtain the filtering result for the shared sample set, and provide the filtering result to the n participants.
9. The method according to claim 8, wherein, The logical operation based on the m logical result sets returned by the participants includes: The m logical result sets are arranged into a numerical array, where each row of the numerical array corresponds to the m logical values of a common sample, and each column corresponds to a sub-filtering condition. For each column in the numerical array, the logical values located in that column are concatenated in sequence to obtain the numerical vectors corresponding to each sub-filter condition; The logical operation is performed on each of the numerical vectors to obtain the target logical value of each common sample corresponding to the overall filtering condition.
10. The method of claim 9, further comprising: The participants receive anonymous identifiers for each shared sample corresponding to each logical value in the logical result set. The step of arranging the m logical result sets into a numerical array includes: Based on the aforementioned anonymous identifiers, the m logical result sets are arranged into a numerical array.
11. The method according to claim 8, wherein, The step of obtaining the filtering results for the shared sample set includes: From the shared sample set, select each target sample whose corresponding target logical value is the first value; The target samples are determined as the filtering results of the common sample set.
12. A method for data filtering based on privacy protection through multi-party collaboration, wherein the multi-party collaboration includes a collaborating party and n participating parties, the n participating parties having a shared sample set and each having different feature terms of each shared sample; the method includes: The collaborating party breaks down the total filtering conditions into m sub-filtering conditions and the logical operations between them, where a single sub-filtering condition is used to limit the value of a feature item. The collaborating party performs desensitization processing on the total filtering conditions to obtain a desensitized expression, which at least indicates the logical operations between the m sub-filtering conditions; The collaborating party sends each sub-filter condition to the participating party that has the feature items defined by the sub-filter condition, and provides the de-identification expression to each participating party; Upon receiving at least one sub-filter condition, participant i determines at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and provides the logical result set to other participants; The participant i summarizes its determined logical result set and the logical result sets received from other participants into m logical result sets, and performs logical operations in the desensitization expression based on them to obtain the target logical value of each common sample corresponding to the total filtering condition. The participant i obtains the filtering result for the shared sample set based on the target logical value.
13. The method according to claim 12, wherein, The collaborating party performs desensitization processing on the overall filtering conditions, including: The collaborating party replaces the m sub-filter conditions with the corresponding m predetermined identifiers, and forms the desensitization expression based on the logical operations between the m predetermined identifiers and their corresponding m sub-filter conditions; The method further includes: The collaborating party will provide each predetermined identifier corresponding to each sub-filter condition to the corresponding participating party.
14. A method for data filtering based on privacy protection through multi-party collaboration, wherein the multi-party collaboration includes n participants, the n participants have a common sample set, and each of them has different features of each common sample; The method is executed through a target participant among the n participants, including: The total filtering condition is broken down into m sub-filtering conditions and the logical operations between them. Each sub-filtering condition is used to limit the value of a feature term. Each sub-filter condition is sent to the participant that has the feature items defined by the sub-filter condition, so that the participant determines its logical result set corresponding to the received sub-filter condition for the common sample set and provides it to the target participant. Based on the logical result set determined by the target participant and / or the logical result set returned by other participants, m logical result sets are determined, and the logical operation is performed on them to obtain the target logical value of each common sample corresponding to the total filtering condition. Based on the target logical value, obtain the filtering result for the common sample set, and provide the filtering result to n-1 participants other than the target participant among the n participants.
15. A privacy-preserving multi-party collaborative data filtering system, comprising a collaborating party and n participating parties, wherein the n participating parties have a shared sample set and each has different features of each shared sample; The collaborating party is used to decompose the total filtering conditions into m sub-filtering conditions and the logical operations between them, wherein a single sub-filtering condition is used to limit the value of a feature item. The collaborating party is also used to send each sub-filter condition to the participating party that has the feature items defined by the sub-filter condition; A participant that receives at least one sub-filter condition is configured to determine at least one logical result set corresponding to the at least one sub-filter condition for the shared sample set, and provide the logical result set to the collaborating party; The collaborating party is also used to perform the logical operation based on the m logical result sets returned by the participating parties to obtain the target logical value of each shared sample corresponding to the total filtering condition; The collaborating party is also configured to obtain the filtering result for the shared sample set based on the target logical value, and provide the filtering result to the n participating parties.
16. The system according to claim 15, wherein, The participating parties are specifically used for: For any first sub-filtering condition among the at least one sub-filtering conditions, based on whether the corresponding features of each common sample in the common sample set satisfy the first sub-filtering condition, the logical value corresponding to each common sample is set to 1 or 0, thereby obtaining a first logical result set corresponding to the first sub-filtering condition.
17. The system according to claim 15, wherein, The participating party is specifically responsible for: performing symmetric encryption on each logical value in the logical result set, and providing the resulting encrypted result set to the collaborating party; The collaborating party is also used to decrypt each encrypted logical value in the encrypted result set to obtain the logical result set.
18. The system according to claim 15, wherein, Each shared sample has a corresponding anonymous identifier; Specifically, the collaborating party is used to: provide the anonymity identifiers corresponding to each target sample to the n participating parties; The participating parties are also used to determine the target samples from the shared sample set based on the anonymous identifiers, and to perform target calculations in conjunction with other participating parties.
19. An apparatus for data filtering based on privacy protection through multi-party collaboration, wherein the multi-party collaboration includes a collaborating party and n participating parties, wherein the n participating parties have a common sample set and each has different features of each common sample; The device is disposed in the collaborating party and includes: The decomposition unit is used to decompose the total filtering conditions into m sub-filtering conditions and the logical operations between them, wherein a single sub-filtering condition is used to limit the value of a feature term. The sending unit is used to send each sub-filter condition to the participant that has the feature items defined by the sub-filter condition, so that the participant determines its logical result set corresponding to the received sub-filter condition for the common sample set and provides it to the collaborating party. The operation unit is used to perform the logical operation based on the m logical result sets returned by the participants to obtain the target logical value of each common sample corresponding to the total filtering condition; The acquisition unit is used to acquire the filtering result for the common sample set based on the target logical value, and to provide the filtering result to the n participants.
20. The apparatus according to claim 19, wherein, The arithmetic unit includes: The arrangement submodule is used to arrange the m logical result sets into a numerical array, where one row of the numerical array corresponds to the m logical values of a common sample, and one column corresponds to a sub-filtering condition. The splicing submodule is used to sequentially splice the logical values located in each column of the numerical array to obtain the numerical vectors corresponding to each sub-filter condition. The logic operation submodule is used to perform the logic operation on each of the numerical vectors to obtain the target logic value of each common sample corresponding to the overall filtering condition.
21. The apparatus of claim 20, further comprising: The receiving unit is used to receive from the participants the anonymous identifiers of the common samples corresponding to each logical value in the logical result set. The arrangement submodule is specifically used for: Based on the aforementioned anonymous identifiers, the m logical result sets are arranged into a numerical array.
22. The apparatus according to claim 19, wherein, The acquisition unit is specifically used for: From the shared sample set, select each target sample whose corresponding target logical value is the first value; The target samples are determined as the filtering results of the common sample set.
23. A system for data filtering based on privacy protection through multi-party collaboration, wherein the multi-party collaboration includes a collaborating party and n participating parties, wherein the n participating parties have a shared sample set and each has different features of each shared sample; The collaborating party is used to decompose the total filtering conditions into m sub-filtering conditions and the logical operations between them, wherein a single sub-filtering condition is used to limit the value of a feature item. The collaborating party is also used to perform desensitization processing on the total filtering conditions to obtain a desensitized expression, wherein at least the logical operations between the m sub-filtering conditions are indicated. The collaborating party is also used to send each sub-filter condition to the participating party that has the feature item defined by the sub-filter condition, and to provide the desensitization expression to each participating party; A participant i that receives at least one sub-filter condition is configured to determine at least one logical result set corresponding to the at least one sub-filter condition for the common sample set, and provide the logical result set to other participants; The participant i is also used to summarize the logical result set it has determined and the logical result set received from other participants into m logical result sets, and perform logical operations in the desensitization expression based on them to obtain the target logical value of each common sample corresponding to the total filtering condition. The participant i is also used to obtain the filtering result for the common sample set based on the target logical value.
24. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed in a computer, it causes the computer to perform the method of any one of claims 1-14.
25. A computing device comprising a memory and a processor, wherein, The memory stores executable code, and when the processor executes the executable code, it implements the method of any one of claims 1-14.
Citation Information
Patent Citations
Data statistics method and device
CN111611618A
Multi-party joint security statistics method and device
CN112084530A