Method for filtering a dataset from a plurality of dataset nodes

CN117493630BActive Publication Date: 2026-09-29BEIJING FUCUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311595088.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2026-09-29
Estimated Expiration
2043-11-27

AI Technical Summary

Technical Problem

在这种场景下要执行安全的高价值数据探查,往往只能通过点对点的质量评估技术来进行单点识别,效率低且不能捕捉相关节点的数据价值影响力信息和共性关系,而且无法发挥网络拓扑能力,难以提供高效的批量化的高价值数据安全筛选方法

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117493630B_ABST
    Figure CN117493630B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method for screening datasets from a plurality of dataset nodes. The method comprises: calculating intersection sizes between each two dataset nodes by a secure intersection approach; calculating intersection ratios between each two dataset nodes according to the intersection sizes between the two dataset nodes; constructing a data linkage graph according to the calculated intersection ratios; iteratively calculating data values of each dataset node based on the data linkage graph; and selecting a target dataset from the plurality of dataset nodes according to the data values of each dataset node. Iteratively calculating the data values of each dataset node based on the data linkage graph comprises: initializing the data values of each dataset node as a preset constant value; and traversing the data linkage graph in each iteration to distribute the data value of each dataset node to the dataset nodes linked thereto according to the weights of the edges until the data values of each dataset node converge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically, to a method for filtering datasets from multiple dataset nodes. Background Technology

[0002] With the development of the internet, various government entities, industry entities, companies, and institutions can be connected via the internet. Each entity can be viewed as a node. In this scenario, secure high-value data exploration often relies on point-to-point quality assessment techniques for single-point identification. This approach is inefficient, fails to capture the data value and influence information and common relationships among relevant nodes, and cannot leverage network topology capabilities, making it difficult to provide an efficient, batch-based method for securely screening high-value data. Summary of the Invention

[0003] The embodiments described herein provide a method, apparatus, and computer-readable storage medium storing a computer program for filtering a dataset from multiple dataset nodes.

[0004] According to a first aspect of this disclosure, a method for filtering datasets from multiple dataset nodes is provided. The method includes: calculating the intersection size between every two dataset nodes using a secure intersection method; calculating the intersection ratio between two dataset nodes based on the intersection size; constructing a data link graph of the multiple dataset nodes based on the calculated intersection ratio, wherein in the data link graph, each dataset node is represented as a vertex, and the intersection ratio between every two dataset nodes is used as the weight of the edge between the two dataset nodes; iteratively calculating the data value of each dataset node based on the data link graph; and selecting a target dataset from the multiple dataset nodes based on the data value of each dataset node. The iterative calculation of the data value of each dataset node based on the data link graph includes: initializing the data value of each dataset node to a preset constant value; and traversing the data link graph in each iteration to distribute the data value of each dataset node to its linked dataset nodes according to the edge weights, until the data value of each dataset node converges.

[0005] In some embodiments of this disclosure, the data link graph is constructed as an undirected graph. The intersection ratio between nodes of the i-th dataset and nodes of the j-th dataset is calculated as follows:

[0006] P = (2 × C) / (A + B)

[0007] Where P represents the intersection ratio between the i-th dataset node and the j-th dataset node, A represents the dataset size of the i-th dataset node, B represents the dataset size of the j-th dataset node, and C represents the intersection size between the i-th dataset node and the j-th dataset node.

[0008] In some embodiments of this disclosure, the data value of the i-th dataset node is calculated in each iteration as follows:

[0009]

[0010] Where PR(i) represents the data value of the i-th dataset node, d is a constant greater than 0 and less than 1, In(i) represents all dataset nodes linked to the i-th dataset node, PR(j) represents the current data value of the j-th dataset node linked to the i-th dataset node, Out(j) represents all dataset nodes linked to the j-th dataset node, Weight(j,i) represents the weight of the edge between the j-th dataset node and the i-th dataset node, and Weight(j,k) represents the weight of the edge between the j-th dataset node and the k-th dataset node.

[0011] In some embodiments of this disclosure, the data link graph is constructed as a directed graph. The intersection ratio from node i to node j of dataset i is calculated as follows:

[0012] Pa = C / A

[0013] Where Pa represents the intersection ratio from the i-th dataset node to the j-th dataset node, A represents the dataset size of the i-th dataset node, and C represents the intersection size between the i-th dataset node and the j-th dataset node.

[0014] In some embodiments of this disclosure, the data value of the i-th dataset node is calculated in each iteration as follows:

[0015]

[0016] Where PR(i) represents the data value of the i-th dataset node, d is a constant greater than 0 and less than 1, N represents the total number of nodes in the data link graph, In(i) represents all dataset nodes linked to the i-th dataset node, PR(j) represents the current data value of the j-th dataset node linked to the i-th dataset node, Weight(j,i) represents the weight of the edge from the j-th dataset node to the i-th dataset node, and ΣWeight(j) represents the sum of the weights of all outgoing links from the j-th dataset node.

[0017] In some embodiments of this disclosure, calculating the intersection size between every two dataset nodes in a plurality of dataset nodes using a secure intersection method includes: obtaining a unique identifier vector in a first original data matrix of a first dataset node; converting the unique identifier vector in the first original data matrix into a first hash vector; obtaining a unique identifier vector in a second original data matrix of a second dataset node; converting the unique identifier vector in the second original data matrix into a second hash vector; comparing each first hash value in the first hash vector with each second hash value in the second hash vector to determine the number of first hash values ​​in the first hash vector that are equal to the second hash values ​​as the intersection size between the first dataset node and the second dataset node.The comparison of the first hash value and the second hash value includes: jointly determining whether the first hash value is less than the second hash value by the first dataset node and the second dataset node; in response to the first hash value not being less than the second hash value, jointly determining whether the second hash value is less than the first hash value by the first dataset node and the second dataset node; in response to the second hash value not being less than the first hash value, determining that the first hash value is equal to the second hash value; wherein, jointly determining whether the first hash value is less than the second hash value by the first dataset node and the second dataset node includes: the first dataset node fragmenting the first hash value into a first fragment value and a second fragment value and sending the second fragment value to the second dataset node; the second dataset node fragmenting the second hash value into a third fragment value and a fourth fragment value and sending the third fragment value to the first dataset node; the first dataset node subtracting the third fragment value from the first fragment value to obtain a fifth fragment value; the second dataset node subtracting the fourth fragment value from the second fragment value to obtain a sixth fragment value; generating a first Boolean zero fragment, a second Boolean zero fragment, a first arithmetic zero fragment, and a second arithmetic zero fragment, wherein the XOR result of the first Boolean zero fragment and the second Boolean zero fragment is 0, and the XOR result of the first arithmetic zero fragment and the second arithmetic zero fragment is equal. The result of the addition is 0; the first Boolean zero fragment and the first arithmetic zero fragment are assigned to the first dataset node; the second Boolean zero fragment and the second arithmetic zero fragment are assigned to the second dataset node; the first dataset node calculates the sum of the fifth fragment value and the first arithmetic zero fragment, XORed with the first Boolean zero fragment to obtain the first operational fragment; the second dataset node calculates the sum of the sixth fragment value and the second arithmetic zero fragment to obtain the second operational fragment; the first dataset node and the second dataset node jointly use the first parallel prefix adder at the first dataset node and the second parallel prefix adder at the second dataset node to obtain the first sign bit fragment at the first dataset node and the second sign bit fragment at the second dataset node, wherein the input of the first parallel prefix adder is the first operational fragment and the third operational fragment, the input of the second parallel prefix adder is the second operational fragment and the fourth operational fragment, the third operational fragment is equal to 0, and the fourth operational fragment is equal to the second Boolean zero fragment; an XOR operation is performed on the first sign bit fragment and the second sign bit fragment to obtain a comparison value; in response to the comparison value being true, it is determined that the first hash value is less than the second hash value; and in response to the comparison value being false, it is determined that the first hash value is not less than the second hash value.

[0018] In some embodiments of this disclosure, constructing a data link graph of multiple dataset nodes based on a calculated intersection ratio includes: evaluating the data quality of the multiple dataset nodes according to a target evaluation metric; and deleting dataset nodes with data quality below a preset value from the data link graph.

[0019] In some embodiments of this disclosure, constructing a data link graph of multiple dataset nodes based on a calculated intersection ratio includes: obtaining the data source domain preference of a target dataset node among the multiple dataset nodes; determining candidate dataset nodes that conform to the data source domain preference among all dataset nodes linked to the target dataset node; and adjusting the weight of the edge between each candidate dataset node and the target dataset node in the data link graph according to an enhancement factor.

[0020] In some embodiments of this disclosure, adjusting the weight of the edge between each candidate dataset node and the target dataset node according to the enhancement factor includes multiplying the weight of the edge between each candidate dataset node and the target dataset node by the enhancement factor. Wherein, when the data source domain preference is similar data, the enhancement factor is a first constant. When the data source domain preference is complementary data, the enhancement factor is a second constant.

[0021] In some embodiments of this disclosure, adjusting the weights of the edges between each candidate dataset node and the target dataset node according to the enhancement factor includes adjusting the weights according to the following formula:

[0022]

[0023] Among them, w t+1 w represents the adjusted weights. t Let represent the weights before adjustment, and p represent the enhancement factor. When the data source domain preference is similar data, the enhancement factor is the third constant. When the data source domain preference is complementary data, the enhancement factor is the fourth constant.

[0024] According to a second aspect of this disclosure, an apparatus for filtering a dataset from a plurality of dataset nodes is provided. The apparatus includes at least one processor and at least one memory storing a computer program. When the computer program is executed by the at least one processor, the apparatus causes to perform the steps of the method according to the first aspect of this disclosure.

[0025] According to a third aspect of this disclosure, a computer-readable storage medium is provided storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described according to a first aspect of this disclosure. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein:

[0027] Figure 1 This is a schematic topology diagram of the data network;

[0028] Figure 2 This is a schematic flowchart of a method for filtering a dataset from multiple dataset nodes according to embodiments of the present disclosure;

[0029] Figure 3 This is an illustrative flowchart and signaling scheme for determining whether a first hash value is less than a second hash value according to embodiments of the present disclosure;

[0030] Figure 4 yes Figure 3 A schematic flowchart and signaling scheme for action 311 in the diagram;

[0031] Figure 5 yes Figure 4 A schematic flowchart and signaling scheme for action 403 in the process;

[0032] Figure 6 yes Figure 4 Schematic flowcharts and signaling schemes for actions 404 and 405 in the code;

[0033] Figure 7 This is an example diagram of a data link diagram according to an embodiment of the present disclosure;

[0034] Figure 8 This is another example diagram of a data link diagram according to an embodiment of the present disclosure;

[0035] Figure 9 This is yet another example diagram of a data link diagram according to an embodiment of the present disclosure;

[0036] Figure 10 This is a schematic block diagram of an apparatus for filtering a dataset from multiple dataset nodes according to embodiments of the present disclosure.

[0037] It should be noted that the elements in the attached diagram are schematic and not drawn to scale. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.

[0039] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having meanings consistent with their meanings in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. Furthermore, terms such as “first” and “second” are used only to distinguish one component (or part of a component) from another component (or another part of a component).

[0040] This disclosure proposes a method for filtering datasets from multiple dataset nodes, aiming to achieve efficient and batch secure filtering of high-value data. These multiple dataset nodes can be distributed across a data network. Figure 1 This diagram illustrates a schematic topology of a data network. The data network may include multiple subnets 10. Each subnet 10 includes a hub node 11 and multiple participating nodes 12 directly connected to the hub node. The hub nodes 11 in these subnets 10 are directly interconnected. Hub nodes 11 can interconnect with each other via a dedicated network. The hub node 11 performs functions such as information aggregation and address navigation for the participating nodes 12. Participating nodes 12 can be various government entities, industry entities, company entities, institutional entities, etc. Participating nodes 12 directly connected to the same hub node 11 communicate through that hub node 11. Participating nodes 12 directly connected to different hub nodes 11 communicate through their respective directly connected hub nodes 11. In other words, participating nodes 12 only communicate directly with the hub nodes 11 they are directly connected to; hub nodes 11 can communicate directly with each other, while participating nodes 12 must communicate with each other through their respective hub nodes 11.

[0041] In practice, a data network may contain a massive number of subnets 10. Each subnet 10 may contain a massive number of participating nodes 12. If each participating node 12 is considered a dataset node, then the number of dataset nodes in the data network can be enormous. In some application scenarios, it is necessary to quickly filter out high-value datasets from a massive number of dataset nodes.

[0042] The method for filtering datasets from multiple dataset nodes according to embodiments of the present disclosure is based on secure computing technology to automate the evaluation of dataset value, and can complete the batch dataset value evaluation of the entire Internet of Things without exposing the privacy information of the data. Figure 2 A schematic flowchart illustrating a method for filtering a dataset from multiple dataset nodes according to an embodiment of the present disclosure is shown.

[0043] exist Figure 2At box S202, the intersection size between every two dataset nodes in the plurality of dataset nodes is calculated using a secure intersection method. A secure intersection method refers to a method that does not disclose the original data information (privacy information) of the dataset nodes and does not disclose sensitive information such as intersection and non-intersection. The intersection size between any two dataset nodes in the plurality of dataset nodes can be calculated in the same way. The following uses the calculation of the intersection size between the first dataset node and the second dataset node as an example to illustrate the secure intersection calculation process. Here, the first dataset node and the second dataset node are any two different dataset nodes in the plurality of dataset nodes. In the embodiments of this disclosure, it is not necessary to calculate the intersection size between a dataset node and itself.

[0044] In one example, a unique identifier vector is obtained from the first raw data matrix of the first dataset node. This unique identifier vector includes the unique identifier (ID) of each raw data in the first raw data matrix. Then, a hash function is used to transform the unique identifier vector in the first raw data matrix into a first hash vector. The first hash vector includes multiple first hash values, each corresponding to a unique identifier in the first raw data matrix.

[0045] In parallel, a unique identifier vector from the second original data matrix of the second dataset nodes can be obtained. This unique identifier vector includes the ID of each original data in the second original data matrix. Then, a hash function is used to transform the unique identifier vector in the second original data matrix into a second hash vector. The second hash vector includes multiple second hash values, each corresponding to a unique identifier in the second original data matrix.

[0046] Then, each first hash value in the first hash vector is compared with each second hash value in the second hash vector to determine the number of first hash values ​​in the first hash vector that are equal to the second hash values ​​as the intersection size between the first dataset nodes and the second dataset nodes.

[0047] The first hash value and the second hash value can be securely compared in the following ways: the first dataset node and the second dataset node jointly determine whether the first hash value is less than the second hash value; if the first hash value is not less than the second hash value, the first dataset node and the second dataset node jointly determine whether the second hash value is less than the first hash value; if the second hash value is not less than the first hash value, the first hash value is determined to be equal to the second hash value.

[0048] Figure 3A schematic flowchart and signaling scheme illustrating the steps for determining whether a first hash value is less than a second hash value according to an embodiment of the present disclosure are shown. In action 301, the first dataset node P1 fragments the first hash value x into a first fragment value x1 and a second fragment value x2 (x = x1 + x2), and in action 303, sends the second fragment value x2 to the second dataset node P2. In action 302, the second dataset node P2 fragments the second hash value y into a third fragment value y1 and a fourth fragment value y2 (y = y1 + y2), and in action 304, sends the third fragment value y1 to the first dataset node P1.

[0049] In action 305, the first dataset node P1 subtracts the third fragment value y1 from the first fragment value x1 to obtain the fifth fragment value z1, i.e., z1 = x1 - y1. In action 306, the second dataset node P2 subtracts the fourth fragment value y2 from the second fragment value x2 to obtain the sixth fragment value z2, i.e., z2 = x2 - y2.

[0050] The first Boolean fragment a1, the second Boolean fragment a2, the first arithmetic fragment b1, and the second arithmetic fragment b2 can be generated from either the first dataset node P1 or the second dataset node P2. The result of XORing the first Boolean fragment a1 and the second Boolean fragment a2 is 0 (a1⊕a2=0), and the result of adding the first arithmetic fragment b1 and the second arithmetic fragment b2 is 0 (b1+b2=0).

[0051] In action 307, the first Boolean fragment a1 and the first arithmetic fragment b1 are assigned to the first dataset node P1. In action 308, the second Boolean fragment a2 and the second arithmetic fragment b2 are assigned to the second dataset node P2.

[0052] In action 309, the first dataset node P1 calculates the sum of the fifth fragment value z1 and the first arithmetic zero fragment b1, XORed with the first Boolean zero fragment a1 to obtain the first operational fragment op11, i.e., op11 = (z1 + b1) ⊕ a1. The first dataset node P1 also holds the third operational fragment op21, where op21 = 0.

[0053] In action 310, the second dataset node P2 calculates the sum of the sixth fragment value z2 and the second arithmetic zero fragment b2 to obtain the second operational fragment op22, i.e., op22 = z2 + b2. The second dataset node P2 also holds the fourth operational fragment op12, where op12 = a2.

[0054] In action 311, the first dataset node P1 and the second dataset node P2 jointly utilize the first parallel prefix adder at the first dataset node P1 and the second parallel prefix adder at the second dataset node P2 to obtain the first sign bit fragment B1 at the first dataset node P1 and the second sign bit fragment B2 at the second dataset node P2. The inputs of the first parallel prefix adder are the first operation fragment op11 and the third operation fragment op21, and the inputs of the second parallel prefix adder are the second operation fragment op22 and the fourth operation fragment op12.

[0055] In the example where the comparison result is determined by the first dataset node P1, the second dataset node P2 sends the second sign bit fragment B2 to the first dataset node P1 in action 313. The first dataset node P1 then performs an XOR operation on the first sign bit fragment B1 and the second sign bit fragment B2 in action 314 to obtain a comparison value. If the comparison value is true, it is determined that the first hash value is less than the second hash value. If the comparison value is false, it is determined that the first hash value is not less than the second hash value. Similarly, the comparison result can also be determined by the second dataset node P2.

[0056] Figure 4 Show Figure 3 The specific process of action 311. In action 403, the first dataset node P1 generates the first intermediate fragment G1 and the second intermediate fragment G2 based on the first operation fragment op11 and the third operation fragment op21, and the second dataset node P2 generates the second operation fragment op22 and the fourth operation fragment op12.

[0057] Figure 5 This diagram illustrates a schematic flowchart and signaling scheme for an AND operation jointly performed by the first dataset node P1 and the second dataset node P2. Figure 5 The following example illustrates this concept: The first dataset node P1 has the first input fragment W1 and the second input fragment V1, and the second dataset node P2 has the third input fragment W2 and the fourth input fragment V2. Figure 4 Action 403 in Figure 5 In the scheme shown, the first operation fragment op11 is equivalent to the first input fragment W1, the third operation fragment op21 is equivalent to the second input fragment V1, the second operation fragment op22 is equivalent to the third input fragment W2, and the fourth operation fragment op12 is equivalent to the fourth input fragment V2.

[0058] The following description Figure 5 The process is shown.

[0059] The first dataset node P1 obtains a triplet fragment matrix in action 501.<R1,S1,T1> In the second dataset, node P2 obtains the triplet fragment matrix in action 502.<R2,S2,T2> Where (R1⊕R2)&(S1⊕S2)=(T1⊕T2).

[0060] In action 503, node P1 of the first dataset performs an XOR operation on W1 and R1 to obtain the third intermediate fragment D1, and performs an XOR operation on V1 and S1 to obtain the fourth intermediate fragment E1. In action 504, node P2 of the second dataset performs an XOR operation on W2 and R2 to obtain the fifth intermediate fragment D2, and performs an XOR operation on V2 and S2 to obtain the sixth intermediate fragment E2.

[0061] In action 505, the second dataset node P2 sends D2 and E2 to the first dataset node P1. In action 506, the first dataset node P1 sends D1 and E1 to the second dataset node P2. In action 507, the first dataset node P1 performs an XOR operation on D1 and D2 to obtain the first composite fragment D, and performs an XOR operation on E1 and E2 to obtain the second composite fragment E. Similarly, in action 508, the second dataset node P2 performs an XOR operation on D1 and D2 to obtain the first composite fragment D, and performs an XOR operation on E1 and E2 to obtain the second composite fragment E.

[0062] In action 509, node P1 of the first dataset calculates the first output fragment O1 = T1⊕(R1&E)⊕(S1&D)⊕(E&D). In action 510, node P2 of the second dataset calculates the second output fragment O2 = T2⊕(R2&E)⊕(S2&D). Figure 4 Action 403 in Figure 5 In the scheme shown, the first intermediate fragment G1 is equivalent to the first output fragment O1, and the second intermediate fragment G2 is equivalent to the second output fragment O2.

[0063] Back Figure 4 In action 404, node P1 of the first dataset performs a bitwise cyclic calculation on each bit of G1 based on the fifth intermediate fragment p1 (p1 = op11 ⊕ op21). In action 405, node P2 of the second dataset performs a bitwise cyclic calculation on each bit of G2 based on the sixth intermediate fragment p2 (p2 = op12 ⊕ op22). Figure 6 Show Figure 4 The schematic flowchart and signaling scheme for actions 404 and 405 in the process.

[0064] In the first dataset node P1, action 601 performs a left shift of 2 on G1. i Bitwise operations, i.e., G1 = G1 << 2 i Where i represents the index of the current loop. The first dataset node P1 is still in action 601, performing a left shift of 2 on p1.i Bit operations (i.e., p1 = p1 << 2) i Then update p1 to the result of XORing p1 with kmask (i.e., p1 = p1 ⊕ kmask). Here, kmask is a matrix of the same size as op11, and each element has a value of 2. i -1.

[0065] In action 602, node P2 in the second dataset performs a left shift of 2 on G2. i Bitwise operations, i.e., G2 = G2 << 2 i In the second dataset, node P2 is still in action 602, which performs a left shift of 2. i Bit operations (i.e., p2 = p2 << 2) i Then update p2 to the result of the XOR of p2 and kmask (i.e., p2 = p2 ⊕ kmask).

[0066] At action 603, the first dataset node P1 and the second dataset node P2 jointly execute the action. Figure 5 The AND operation is shown. G1 is equivalent to the first input fragment W1, p1 is equivalent to the second input fragment V1, G2 is equivalent to the third input fragment W2, and p2 is equivalent to the fourth input fragment V2. After the operation of action 603, the first dataset node P1 obtains the seventh intermediate fragment F1, and the second dataset node P2 obtains the eighth intermediate fragment F2. F1 is equivalent to the first output fragment O1, and F2 is equivalent to the second output fragment O2.

[0067] At action 604, the first dataset node P1 and the second dataset node P2 jointly execute the action. Figure 5 The AND operation is shown. p1 corresponds to the first input fragment W1 and the second input fragment V1, and p2 corresponds to the third input fragment W2 and the fourth input fragment V2. After operation 604, the first dataset node P1 obtains the updated p1, and the second dataset node P2 obtains the updated p2. The updated p1 corresponds to the first output fragment O1, and the updated p2 corresponds to the second output fragment O2. The updated p1 will be substituted into operation 601 of the next loop. The updated p2 will be substituted into operation 602 of the next loop.

[0068] In action 606, node P1 of the first dataset performs an XOR operation on G1 and F1 to obtain an updated G1 (i.e., G1 = G1⊕F1). The updated G1 will be substituted into action 601 of the next loop. In action 607, node P2 of the second dataset performs an XOR operation on G2 and F2 to obtain an updated G2 (i.e., G2 = G2⊕F2). The updated G2 will be substituted into action 602 of the next loop.

[0069] Back to Figure 4, the first data set node P1 left-shifts G1 by 1 bit at act 406 to obtain a ninth intermediate fragment C1 (i.e., C1=G1<<1). The second data set node P2 left-shifts G2 by 1 bit at act 407 to obtain a tenth intermediate fragment C2 (i.e., C2=G2<<1).

[0070] The first data set node P1 performs an XOR operation on p1 and C1 at act 408 to obtain an eleventh intermediate fragment Z1 (i.e., Z1=p1⊕C1). The second data set node P2 performs an XOR operation on p2 and C2 at act 409 to obtain a twelfth intermediate fragment Z2 (i.e., Z2=p2⊕C2).

[0071] The first data set node P1 performs a bitwise AND operation on Z1 and mask at act 410 to obtain an updated Z1 (i.e., Z1=Z1 & mask), where mask=0x1<<n-1, and n represents the number of bits of the first hash value x. The second data set node P2 performs a bitwise AND operation on Z2 and mask at act 411 to obtain an updated Z2 (i.e., Z2=Z2 & mask), where mask=0x1<<n-1, and n represents the number of bits of the second hash value y.

[0072] The first data set node P1 converts Z1 into a boolean type at act 412 to obtain a first sign bit fragment B1. The second data set node P2 converts Z2 into a boolean type at act 413 to obtain a second sign bit fragment B2.

[0073] In the above process, since the first data set node P1 does not obtain complete information of the second hash value, and the second data set node P2 does not obtain complete information of the first hash value, the calculation process is secure, and no other information will be leaked except the intersection size.

[0074] The process of determining whether the second hash value is less than the first hash value can be Figure 3 similar to the process of, and will not be repeated herein.

[0075] Returning to Figure 2 , at block S204, the intersection ratio between every two data set nodes among the plurality of data set nodes is calculated according to the intersection size of the two data set nodes.

[0076] In some embodiments of the present disclosure, the directionality between two data set nodes is not considered. The intersection ratio between the i-th data set node and the j-th data set node is calculated as:

[0077] P =(2×C) / (A+B) (1)

[0078] Where P represents the intersection ratio between the i-th dataset node and the j-th dataset node, A represents the dataset size of the i-th dataset node, B represents the dataset size of the j-th dataset node, and C represents the intersection size between the i-th dataset node and the j-th dataset node. In this context, the i-th dataset node and the j-th dataset node refer to any two distinct dataset nodes among the multiple dataset nodes.

[0079] In other embodiments of this disclosure, the directionality between two dataset nodes is considered. The intersection ratio from the i-th dataset node to the j-th dataset node is calculated as follows:

[0080] Pa = C / A (2)

[0081] Where Pa represents the intersection ratio from the i-th dataset node to the j-th dataset node, A represents the dataset size of the i-th dataset node, and C represents the intersection size between the i-th dataset node and the j-th dataset node.

[0082] The proportion of the intersection from node j in the dataset to node i in the dataset is calculated as follows:

[0083] Pb = C / B (3)

[0084] Where Pb represents the intersection ratio from the j-th dataset node to the i-th dataset node, B represents the dataset size of the j-th dataset node, and C represents the intersection size between the i-th dataset node and the j-th dataset node.

[0085] At box S206, a data link graph of multiple dataset nodes is constructed based on the calculated intersection ratio. In this graph, each dataset node is represented as a vertex, and the intersection ratio between any two dataset nodes serves as the weight of the edge between them. If the intersection ratio between two dataset nodes is 0, then the two dataset nodes are not connected.

[0086] In some embodiments of this disclosure, the data link graph can be constructed as an undirected graph. The intersection ratio between any two different dataset nodes is calculated according to equation (1). Figure 7 This is an example graph of data links constructed as an undirected graph. Nodes 1 through 20 of the dataset are each represented as a vertex. The numbers labeled on the edges between two vertices indicate the proportion of intersection between the two dataset nodes.

[0087] In other embodiments of this disclosure, the data link graph can be constructed as a directed graph. The intersection ratio between any two different dataset nodes is calculated according to equations (2) and (3). Figure 8This is an example graph of data links constructed as a directed graph. Dataset nodes 1 through 20 are each represented as a vertex. Edges between two vertices are labeled with arrows, indicating the direction. For example, the number 0.34 on the edge from dataset node 5 to dataset node 6 indicates that the intersection ratio of dataset nodes 5 and 6 is 0.34. The number 0.78 on the edge from dataset node 6 to dataset node 5 indicates that the intersection ratio of dataset nodes 6 and 5 is 0.78.

[0088] In some embodiments of this disclosure, after constructing Figure 7 or Figure 8 Following the data linking graph shown, the data quality of multiple dataset nodes in the graph can be evaluated based on target evaluation metrics. Then, dataset nodes with data quality below a preset value (referred to as "low-quality dataset nodes") are removed from the data linking graph. This is equivalent to an initial screening. Target evaluation metrics can be calculated using a federated learning-based preprocessing algorithm. Target evaluation metrics may include: missing data rate (if the missing information is too large to meet the computational requirements), anomaly rate (range detection, numerical validity detection, numerical logic detection), feature collinearity variance inflation factor (VIF) detection, information value (IV) detection (the contribution of features to the model's predictive ability), and data duplication detection (the proportion of duplicate data), etc. Figure 9 Show Figure 7 An example diagram of the data link graph after low-quality dataset nodes have been deleted.

[0089] Furthermore, embodiments of this disclosure propose enhancing the edges in the data link graph based on the similarity or complementarity of the dataset nodes' domains. This data link graph can be the one constructed at box S206, or it can be the data link graph after deleting low-quality dataset nodes. Data set nodes in the data network generally belong to different domains, such as the telecommunications, insurance, and social security domains. Data set nodes in the data network can explicitly state their domain preferences for the required data sources; for example, an insurance company node might request social security government data, or it might request data sources for conducting similar insurance business. Based on these explicit preference requirements, the strength of the preference is reflected by enhancing the edge weights.

[0090] In some embodiments of this disclosure, the data source domain preference of the target dataset node among multiple dataset nodes can be obtained. Then, candidate dataset nodes that match the data source domain preference are determined from all dataset nodes linked to the target dataset node. Next, in the data link graph, the weights of the edges between each candidate dataset node and the target dataset node are adjusted according to an enhancement factor. The edge weights can be enhanced linearly or non-linearly.

[0091] In an embodiment where edge weights are enhanced linearly, the weight of each edge from a candidate dataset node to a target dataset node is multiplied by an enhancement factor. The enhancement factor is a first constant when the data source domain preference is similar data, and a second constant when the data source domain preference is complementary data. Both the first and second constants are greater than 1. The first and second constants can be adjusted as needed to achieve a suitable edge weight enhancement effect. In one example, the first constant equals 1.1, and the second constant equals 1.3.

[0092] In embodiments where edge weights are enhanced in a non-linear manner, the weights can be adjusted according to the following formula:

[0093]

[0094] Among them, w t+1 w represents the adjusted weights. t Let represent the weights before adjustment, and p represent the enhancement factor. When the data source domain preference is similar data, the enhancement factor is a third constant. When the data source domain preference is complementary data, the enhancement factor is a fourth constant. Both the third and fourth constants are less than 1. In one example, the third constant equals 0.2, and the fourth constant equals 0.5.

[0095] Back Figure 2 At box S208, the data value of each dataset node is iteratively calculated based on the data link graph. This data link graph can be the one constructed at box S206, the one after deleting low-quality dataset nodes, or the one after enhancing edge weights.

[0096] In some embodiments of this disclosure, the data value of each dataset node can be initialized to a preset constant value. Then, in each iteration, the data link graph is traversed to distribute the data value of each dataset node to the dataset nodes it is linked to according to the edge weights, until the data value of each dataset node converges.

[0097] In the embodiment where the data link graph is constructed as an undirected graph, the data value of the i-th dataset node is calculated in each iteration as follows:

[0098]

[0099] Where PR(i) represents the data value of the i-th dataset node, d is a constant greater than 0 and less than 1, In(i) represents all dataset nodes linked to the i-th dataset node, PR(j) represents the current data value of the j-th dataset node linked to the i-th dataset node, Out(j) represents all dataset nodes linked to the j-th dataset node, Weight(j,i) represents the weight of the edge between the j-th dataset node and the i-th dataset node, and Weight(j,k) represents the weight of the edge between the j-th dataset node and the k-th dataset node.

[0100] In the embodiment where the data link graph is constructed as a directed graph, the data value of the i-th dataset node is calculated in each iteration as follows:

[0101]

[0102] Where PR(i) represents the data value of the i-th dataset node, d is a constant greater than 0 and less than 1, N represents the total number of nodes in the data link graph, In(i) represents all dataset nodes linked to the i-th dataset node, PR(j) represents the current data value of the j-th dataset node linked to the i-th dataset node, Weight(j,i) represents the weight of the edge from the j-th dataset node to the i-th dataset node, and ΣWeight(j) represents the sum of the weights of all outgoing links from the j-th dataset node.

[0103] In the iterative calculation of the data value of each dataset node, a dataset node receives more data value if it is linked to a dataset node with high data value. If a dataset node is linked to many dataset nodes, the transmitted data value is more dispersed. The weights of edges in the data link graph reflect the quality of the link between two dataset nodes, making link quality a factor influencing data value. If a dataset node has a higher weight linking to a target dataset node, this link will have a greater influence on data value calculation, thus increasing the data value of the target dataset node.

[0104] By using edge weights, the importance of dataset nodes can be controlled more precisely, as link weights can be adjusted based on various factors, such as the authority of the data source and the relevance of the link. This helps to more accurately assess the importance of datasets in data networks and improve the quality of dataset selection results.

[0105] Furthermore, to improve the stability and convergence of the algorithm, a constant d is introduced in both equations (5) and (6). d can be considered as a damping factor, which helps to converge the data value of each dataset node.

[0106] At box S210, the target dataset is selected from multiple dataset nodes based on the data value of each dataset node. In one example, the data values ​​of the dataset nodes can be sorted, and the target dataset is selected according to the sorting result. In another example, the data value of each dataset node can be compared with a data value threshold, and the target dataset is selected from the dataset nodes whose data values ​​exceed the data value threshold.

[0107] Figure 10 A schematic block diagram of an apparatus for filtering a dataset from multiple dataset nodes according to an embodiment of the present disclosure is shown. Figure 10 As shown, the device 700 may include a processor 710 and a memory 720 storing a computer program. When the computer program is executed by the processor 710, the device 700 is able to perform actions such as... Figure 2 The steps of method 200 are shown below. In one example, device 700 may be a computer device or a cloud computing node. Device 700 can calculate the intersection size between every two dataset nodes in the plurality of dataset nodes using a secure intersection method. Device 700 can calculate the intersection ratio between two dataset nodes based on the intersection size between every two dataset nodes in the plurality of dataset nodes. Device 700 can construct a data link graph of the plurality of dataset nodes based on the calculated intersection ratio. In the data link graph, each dataset node is represented as a vertex. The intersection ratio between every two dataset nodes is used as the weight of the edge between the two dataset nodes. Device 700 can iteratively calculate the data value of each dataset node based on the data link graph. Device 700 can select a target dataset from the plurality of dataset nodes based on the data value of each dataset node. The iterative calculation of the data value of each dataset node based on the data link graph includes: initializing the data value of each dataset node to a preset constant value; and traversing the data link graph in each iteration to distribute the data value of each dataset node to the dataset nodes it is linked to according to the edge weights, until the data value of each dataset node converges.

[0108] In embodiments of this disclosure, processor 710 may be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. Memory 720 may be any type of memory implemented using data storage technologies, including but not limited to random access memory, read-only memory, semiconductor-based memory, flash memory, disk storage, etc.

[0109] Furthermore, in embodiments of this disclosure, the device 700 may also include an input device 730, such as a keyboard or mouse, for inputting input data for each dataset node. Additionally, the device 700 may also include an output device 740, such as a display, for outputting the filtering results.

[0110] In other embodiments of this disclosure, a computer-readable storage medium storing a computer program is also provided, wherein the computer program, when executed by a processor, is capable of performing the following functions: Figure 2 The steps of the method shown.

[0111] In summary, the method for filtering datasets from multiple dataset nodes according to embodiments of this disclosure can automate the evaluation of dataset value, and can complete batch dataset value evaluations without exposing data privacy information. By leveraging the weights of edges in the data link graph, the importance of dataset nodes can be controlled more precisely, improving the quality of dataset filtering results.

[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses and methods according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0113] Unless otherwise expressly indicated by the context, the singular form of words used herein and in the appended claims includes the plural form, and vice versa. Thus, when referring to the singular, the plural form of the corresponding term is generally included. Similarly, the terms “comprising” and “including” shall be interpreted as including rather than exclusively. Likewise, the terms “including” and “or” shall be interpreted as including unless such interpretation is expressly prohibited herein. Where the term “example” is used herein, particularly when it follows a set of terms, the “example” is merely exemplary and illustrative and should not be considered exclusive or extensive.

[0114] Further aspects and scope of adaptation become apparent from the description provided herein. It should be understood that various aspects of this application may be implemented individually or in combination with one or more other aspects. It should also be understood that the descriptions and specific embodiments herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0115] Several embodiments of this disclosure have been described in detail above. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of this disclosure without departing from the spirit and scope of this disclosure. The scope of protection of this disclosure is defined by the appended claims.

Claims

1. A method for filtering datasets from multiple dataset nodes, characterized in that, The method includes: The intersection size between any two dataset nodes in the plurality of dataset nodes is calculated using a secure intersection method, including: Obtain the unique identifier vector in the first original data matrix of the first dataset node; Convert the unique identifier vector in the first original data matrix into a first hash vector; Obtain the unique identifier vector in the second original data matrix of the second dataset node; Convert the unique identifier vector in the second original data matrix into a second hash vector; Each first hash value in the first hash vector is compared with each second hash value in the second hash vector to determine the number of first hash values ​​in the first hash vector that are equal to the second hash values ​​as the size of the intersection between the first dataset node and the second dataset node; The intersection ratio between two dataset nodes is calculated based on the size of the intersection between each pair of dataset nodes in the plurality of dataset nodes; A data link graph of the multiple dataset nodes is constructed based on the calculated intersection ratio, wherein each dataset node is represented as a vertex in the data link graph, and the intersection ratio between any two dataset nodes is used as the weight of the edge between the two dataset nodes. The data value of each dataset node is iteratively calculated based on the data link graph; and The target dataset is selected from the plurality of dataset nodes based on the data value of each dataset node. The iterative calculation of the data value of each dataset node based on the data link graph includes: Initialize the data value of each dataset node to a preset constant value; and In each iteration, the data link graph is traversed to distribute the data value of each dataset node to the dataset nodes it is linked to according to the weight of the edges, until the data value of each dataset node converges.

2. The method according to claim 1, characterized in that, The data link graph is constructed as an undirected graph, and the intersection ratio between nodes of the i-th dataset and nodes of the j-th dataset is calculated as follows: P = (2 × C) / (A + B) Where P represents the intersection ratio between the i-th dataset node and the j-th dataset node, A represents the dataset size of the i-th dataset node, B represents the dataset size of the j-th dataset node, and C represents the intersection size between the i-th dataset node and the j-th dataset node.

3. The method according to claim 2, characterized in that, In each iteration, the data value of the i-th dataset node is calculated as follows: Wherein, PR(i) represents the data value of the i-th dataset node, d is a constant greater than 0 and less than 1, In(i) represents all dataset nodes linked to the i-th dataset node, PR(j) represents the current data value of the j-th dataset node linked to the i-th dataset node, Out(j) represents all dataset nodes linked to the j-th dataset node, Weight(j, i) represents the weight of the edge between the j-th dataset node and the i-th dataset node, and Weight(j, k) represents the weight of the edge between the j-th dataset node and the k-th dataset node.

4. The method according to claim 1, characterized in that, The data link graph is constructed as a directed graph, and the intersection ratio from the node of the i-th dataset to the node of the j-th dataset is calculated as follows: Pa = C / A Where Pa represents the intersection ratio from the i-th dataset node to the j-th dataset node, A represents the dataset size of the i-th dataset node, and C represents the intersection size between the i-th dataset node and the j-th dataset node.

5. The method according to claim 4, characterized in that, In each iteration, the data value of the i-th dataset node is calculated as follows: Wherein, PR(i) represents the data value of the i-th dataset node, d is a constant greater than 0 and less than 1, N represents the total number of nodes in the data link graph, In(i) represents all dataset nodes linked to the i-th dataset node, PR(j) represents the current data value of the j-th dataset node linked to the i-th dataset node, Weight(j, i) represents the weight of the edge from the j-th dataset node to the i-th dataset node, and ΣWeight(j) represents the sum of the weights of all outgoing links from the j-th dataset node.

6. The method according to any one of claims 1 to 5, characterized in that, Comparing the first hash value with the second hash value includes: The first dataset node and the second dataset node jointly determine whether the first hash value is less than the second hash value; In response to the first hash value being not less than the second hash value, the first dataset node and the second dataset node jointly determine whether the second hash value is less than the first hash value; In response to the second hash value being not less than the first hash value, the first hash value is determined to be equal to the second hash value; The determination of whether the first hash value is less than the second hash value by the first dataset node and the second dataset node jointly includes: The first dataset node fragments the first hash value into a first fragment value and a second fragment value, and sends the second fragment value to the second dataset node. The second dataset node fragments the second hash value into a third fragment value and a fourth fragment value, and sends the third fragment value to the first dataset node; The first fragment value is obtained by subtracting the third fragment value from the first fragment value using the first dataset node; The second dataset node subtracts the fourth fragment value from the second fragment value to obtain the sixth fragment value; Generate a first Boolean fragment, a second Boolean fragment, a first arithmetic fragment, and a second arithmetic fragment, wherein the XOR result of the first Boolean fragment and the second Boolean fragment is 0, and the sum of the first arithmetic fragment and the second arithmetic fragment is 0; Assign the first Boolean fragment and the first arithmetic fragment to the first dataset node; Assign the second Boolean fragment and the second arithmetic fragment to the second dataset node; The first operational fragment is obtained by XORing the sum of the fifth fragment value and the first arithmetic zero fragment with the first Boolean zero fragment, calculated from the first dataset node. The second dataset node calculates the sum of the sixth fragment value and the second arithmetic zero fragment to obtain the second operational fragment; The first dataset node and the second dataset node jointly utilize the first parallel prefix adder at the first dataset node and the second parallel prefix adder at the second dataset node to obtain a first sign bit fragment at the first dataset node and a second sign bit fragment at the second dataset node. The inputs of the first parallel prefix adder are the first operation fragment and the third operation fragment, and the inputs of the second parallel prefix adder are the second operation fragment and the fourth operation fragment. The third operation fragment is equal to 0, and the fourth operation fragment is equal to the second Boolean zero fragment. Perform an XOR operation on the first sign bit fragment and the second sign bit fragment to obtain a comparison value; In response to the comparison value being true, it is determined that the first hash value is less than the second hash value; and In response to the comparison value being false, it is determined that the first hash value is not less than the second hash value.

7. The method according to any one of claims 1 to 5, characterized in that, Constructing the data link graph of the multiple dataset nodes based on the calculated intersection ratio includes: The data quality of the multiple dataset nodes is evaluated based on the target evaluation metrics; and Nodes in datasets with data quality below a preset value will be removed from the data link graph.

8. The method according to any one of claims 1 to 5, characterized in that, Constructing the data link graph of the multiple dataset nodes based on the calculated intersection ratio includes: Obtain the data source domain preference of the target dataset node among the multiple dataset nodes; Identify candidate dataset nodes that match the data source domain preferences from all dataset nodes linked to the target dataset node; and In the data link graph, the weights of the edges between each candidate dataset node and the target dataset node are adjusted according to an enhancement factor.

9. The method according to claim 8, characterized in that, Adjusting the weight of the edge between each candidate dataset node and the target dataset node according to the enhancement factor includes: multiplying the weight of the edge between each candidate dataset node and the target dataset node by the enhancement factor; Wherein, when the data source domain preference is similar data, the enhancement factor is a first constant; When the data source domain preference is complementary data, the enhancement factor is a second constant.

10. The method according to claim 8, characterized in that, Adjusting the weights of the edges between each candidate dataset node and the target dataset node based on the enhancement factor includes adjusting the weights according to the following formula: Among them, w t+1 w represents the adjusted weights. t represents the weights before adjustment, and p represents the enhancement factor; When the data source domain preference is similar data, the enhancement factor is a third constant; When the data source domain preference is complementary data, the enhancement factor is a fourth constant.

Citation Information

Patent Citations

  • Multi-party security joint training processing method, device, equipment, medium and program product

    CN115577383A

  • Privacy set intersection and data analysis method based on trusted hardware

    CN116484389A