A method, apparatus, and computer-readable storage medium for label propagation in an association network

By building a federal affiliated network in the affiliated network and performing multiple rounds of tag propagation, the problem of privacy protection in cross-institutional data joint application is solved, and effective tag propagation and efficient utilization of data value of cross-platform networks are achieved.

CN115733763BActive Publication Date: 2025-06-20CHINA UNIONPAY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211492068.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2025-06-20
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

The existing technology cannot realize the joint application of cross-organization data under the premise of privacy protection, resulting in the tag updates in the associated network being unable to utilize the data of the associated network of both parties at the same time, and the data value cannot be efficiently utilized.

Method used

By associating the first and second association networks based on the security interception protocol, a federated association network is built, and multiple rounds of label propagation are performed on its nodes iteratively, the label propagation probability between adjacent nodes is determined, and the label of each node is updated.

Benefits of technology

On the premise of ensuring that private data is not leaked, tag dissemination on cross-platform networks is realized and the efficiency of data value utilization is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115733763B_ABST
    Figure CN115733763B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus and computer-readable storage medium for label propagation in an associated network. The method includes: constructing a first associated network based on first-party data and a second associated network based on second-party data; associating the first associated network and the second associated network based on a secure intersection protocol to obtain a federated associated network; iteratively performing multiple rounds of label propagation on the nodes of the federated associated network; wherein each round of label propagation includes: determining the label propagation probability between adjacent nodes in the federated associated graph; for each node, determining the current-round label of each node according to the current-round labels of the neighbor nodes and the label propagation probability of the neighbor nodes for the node. By using the above method, it is possible to achieve label propagation in a cross-platform network while ensuring the privacy of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computers, and particularly relates to a method and device for label propagation in an association network and a computer-readable storage medium. Background Art

[0002] This section aims to provide background or context for the embodiments of the present invention described in the claims. The description herein is not admitted to be prior art merely because it is included in this section.

[0003] With the increasing strictness of privacy protection laws, data cooperation between institutions increasingly needs to consider data privacy protection issues. Currently, privacy computing technologies mainly focus on scenarios such as federated learning, secure intersection, and anonymous query, all of which are for joint computing of single-point data. The data sources calculated by the current label propagation algorithm in the association network are all local. It is impossible to realize cross-institutional data joint application under the premise of privacy protection. The update of labels cannot utilize the association network data of both parties at the same time, and the data value is not efficiently utilized.

[0004] Therefore, how to achieve label propagation in a federated network under the premise of privacy protection is an urgent problem to be solved. Summary of the Invention

[0005] In view of the problems existing in the above-mentioned prior art, a method and device for label propagation in an association network and a computer-readable storage medium are proposed. Using this method, device and computer-readable storage medium, the above problems can be solved.

[0006] The present invention provides the following solutions.

[0007] In a first aspect, a method for label propagation in an association network is provided, including: constructing a first association network based on first-party data and a second association network based on second-party data; associating the first association network and the second association network based on a secure intersection protocol to obtain a federated association network; iteratively performing multiple rounds of label propagation on the nodes of the federated association network; where each round of label propagation includes: determining the label propagation probability between adjacent nodes in the federated association graph; for each node, determining the current round label of each node according to the current round label of the neighbor nodes and the label propagation probability of the neighbor nodes for the node.

[0008] In an embodiment, associating the first association network and the second association network based on a secure intersection protocol to obtain a federated association network further includes: performing encrypted intersection on the first-party data and the second-party data, determining the common nodes in the first association network and the second association network, and associating the first association network and the second association network according to the common nodes to obtain a federated association network;

[0009] In one implementation, determining the label propagation probability between adjacent nodes in the federated association graph further includes: determining the edge weight w of the edge ij between node i and its neighbor node j in the federated association network ij ; determining the sum of edge weights ∑ j w ij ; according to the ratio of the edge weight w ij and the sum of edge weights ∑ j w ij , determining the label propagation probability P ij of neighbor node j for node i.

[0010] In one implementation, if node i is a non-shared node, all neighbor nodes J represent all neighbor nodes of the graph where node i is located.

[0011] In one implementation, if node i is a shared node, all neighbor nodes J represent the set of all neighbor nodes a of node i in the first association network and all neighbor nodes b of node i in the second association network.

[0012] In one implementation, if node i is a shared node, the first party and the second party interact with the sum of edge weights between node i and all neighbor nodes a of the first association network and the sum of edge weights between node i and all neighbor nodes b of the second association network.

[0013] In one implementation, it further includes: if node i is a shared node, using the following formula to determine the sum of edge weights ∑ j w ij : ∑ j w ij = ∑ a w ia + ∑ b w ib ; where ∑ a w ia is the sum of edge weights between node i and all neighbor nodes a of the first association network, and ∑ b w ib is the sum of edge weights between node i and all neighbor nodes b of the second association network.

[0014] In one implementation, iteratively performing multiple rounds of label propagation on the nodes of the federated association network further includes: determining the labeled nodes and unlabeled nodes of the federated association network; updating the labels of the unlabeled nodes round by round until the labels of the unlabeled nodes no longer change and / or exceed the update round threshold; and keeping the labels of the labeled nodes unchanged.

[0015] In one implementation, determining the current round label of each node according to the current round label of the neighbor nodes and the label propagation probability of the neighbor nodes for the node includes: for each node, determining the current round label of each neighbor node of the node and the label propagation probability of each neighbor node for the node; calculating the sum of the label propagation probabilities corresponding to each label among all the neighbor nodes of the node to obtain the aggregated label propagation probability corresponding to each label; and updating the current round label of the node according to the label with the maximum aggregated label propagation probability.

[0016] In one implementation, if the node is a non - shared node, the method further includes: the party where the node is located calculates the aggregated label propagation probability corresponding to each label of the neighbor nodes of the node.

[0017] In one implementation, if the node is a shared node, the method further includes: the first party calculates the first - party aggregated label propagation probability corresponding to the labels of all the neighbor nodes of the node in the first associated network; the second party calculates the second aggregated label propagation probability corresponding to the labels of all the neighbor nodes of the node in the second associated network; the first party and the second party exchange the first aggregated label propagation probability and the second aggregated label propagation probability; and the first party and the second party each perform label propagation probability aggregation again based on the exchanged information to obtain the aggregated label propagation probability corresponding to each label.

[0018] In one implementation, it further includes: determining the graph weights of the first associated network and the second associated network according to the closeness of the node relationships in the first associated network and the second associated network; and introducing the graph weights during the interaction process between the first associated network and the second associated network.

[0019] In one implementation, introducing the graph weights during the interaction process between the first associated network and the second associated network includes: if node i is a shared node, using the following formula to determine the sum of edge weights ∑ j w ij :

[0020] ∑ j w ij =θ a ∑ a w ia +θ b ∑ b w ib ; where ∑ a w ia is the sum of the edge weights between node i and all the neighbor nodes a in the first associated network, ∑ b w ib is the sum of the edge weights between node i and all the neighbor nodes b in the second associated network, θ a is the graph weight of the first associated network, and θ b is the graph weight of the second associated network.

[0021] In one embodiment, when introducing graph weights during the interaction between the first associated network and the second associated network, it further includes: after the first party and the second party exchange the first label propagation aggregation probability and the second label propagation aggregation probability, performing label propagation probability aggregation again based on the graph weights of the first associated network and the second associated network to obtain the label propagation aggregation probability corresponding to each label.

[0022] In one embodiment, it further includes: if the first associated network and the second associated network are directed graph networks, only taking the incoming neighbor nodes of each node as neighbor nodes.

[0023] In a second aspect, a label propagation device for an associated network is provided, including: a graph construction module for constructing a first associated network based on first-party data and constructing a second associated network based on second-party data; a federated network module for associating the first associated network and the second associated network based on a secure intersection protocol to obtain a federated associated network; a label propagation module for iteratively performing multiple rounds of label propagation on the nodes of the federated associated network; wherein, each round of label propagation includes: determining the label propagation probability between adjacent nodes in the federated associated graph; for each node, determining the current-round label of each node according to the current-round labels of the neighbor nodes and the label propagation probability of the neighbor nodes for the node.

[0024] In a third aspect, a label propagation device for an associated network is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute: the method as in the first aspect.

[0025] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores a program, which when executed by a multi-core processor, enables the multi-core processor to execute the method as in the first aspect.

[0026] One of the advantages of the above embodiments is that it can achieve label propagation in a cross-platform network while ensuring private data.

[0027] Other advantages of the present invention will be explained in more detail in conjunction with the following description and drawings.

[0028] It should be understood that the above description is only an overview of the technical solution of the present invention, so as to be able to understand the technical means of the present invention more clearly, and thus it can be implemented according to the content of the description. In order to make the above and other objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically illustrated below. Description of the Drawings

[0029] By reading the following detailed description of the exemplary embodiments, those of ordinary skill in the art will understand the advantages and benefits described herein, as well as other advantages and benefits. The drawings are only for the purpose of illustrating the exemplary embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0030] Figure 1 FIG. is a schematic structural diagram of a label propagation device for an association network according to an embodiment of the present invention;

[0031] Figure 2 FIG. is a schematic flowchart of a label propagation method for an association network according to an embodiment of the present invention;

[0032] Figure 3 FIG.

[0032] is a schematic diagram of a first association network and a second association network according to an embodiment of the present invention;

[0033] Figure 4 FIG. Figure 4 is a schematic diagram of a federated association network according to an embodiment of the present invention;

[0034] Figure 5 FIG. Figure 5 is a schematic diagram of determining the label propagation probabilities of a first association network and a second association network according to an embodiment of the present invention;

[0035] Figure 6 FIG. Figure 6 is a schematic diagram of determining the label propagation probability of a federated association network according to an embodiment of the present invention;

[0036] Figure 7 FIG. Figure 7 is a schematic diagram of label propagation of an association network according to an embodiment of the present invention;

[0037] Figure 8 FIG. Figure 8 is a schematic diagram of label propagation of an association network according to an embodiment of the present invention;

[0038] Figure 9 FIG. Figure 9 is a schematic structural diagram of a label propagation device for an association network according to an embodiment of the present invention.

[0039] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. Detailed Embodiments

[0040] The following will describe in more detail the exemplary embodiments of the present disclosure with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0041] In the description of the embodiments of the present application, it should be understood that terms such as "including" or "having" are intended to indicate the presence of features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0042] Unless otherwise specified, " / " means "or". For example, A / B may mean A or B. "And / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone.

[0043] The terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0044] To clearly illustrate the embodiments of the present application, some concepts that may appear in subsequent embodiments will be introduced first.

[0045] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0046] First, refer to Figure 1 , which schematically shows a diagram of an environment 100 in which an exemplary implementation according to the present disclosure can be used.

[0047] Figure 1 A diagram showing an example of a computing device 100 according to an embodiment of the present disclosure is shown. It should be noted that Figure 1 This can be a schematic diagram of the hardware operating environment structure of the label propagation method for an association network. The device based on the label propagation method for an association network in the embodiments of the present invention can be a terminal device such as a PC or a portable computer.

[0048] As Figure 1As shown in the figure, the apparatus for the label propagation method of the association network may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to implement the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0049] Those skilled in the art can understand that Figure 1 the structure of the label propagation device of the association network shown in the figure does not limit the apparatus for the label propagation method of the association network, and may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0050] Such as Figure 1 As shown in the figure, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a program for the label propagation method of the association network. Among them, the operating system is a program for managing and controlling the hardware and software resources of the label propagation device of the association network, and supports the operation of the label propagation program of the association network and other software or programs.

[0051] In Figure 1 the label propagation device of the association network shown in the figure, the user interface 1003 is mainly used to receive requests, data, etc. sent by the first terminal, the second terminal, and the supervision terminal; the network interface 1004 is mainly used to connect to the background server and perform data communication with the background server; and the processor 1001 may be used to call the label propagation program stored in the memory 1005 and perform the following operations:

[0052] Construct a first association network based on the first-party data, and construct a second association network based on the second-party data; associate the first association network and the second association network based on the secure intersection protocol to obtain a federated association network; iteratively perform multiple rounds of label propagation on the nodes of the federated association network; where each round of label propagation includes: determining the label propagation probability between adjacent nodes in the federated association graph; for each node, determining the current round label of each node according to the current round label of the neighbor nodes and the label propagation probability of the neighbor nodes for the node.

[0053] Thus, the two parties only need to exchange non-private data such as label propagation probabilities, and they may be able to perform cross-institutional data association network joint calculation and application solutions without the original data of both parties leaving the database.

[0054] Figure 2 FIG. shows a flowchart of a method for performing label propagation of an association network according to an embodiment of the present disclosure. This method can be executed, for example, by a computing device 100 as shown in Figure 1 It should be understood that method 200 may further include additional blocks not shown and / or may omit the shown blocks, and the scope of the present disclosure is not limited in this regard.

[0055] Step 210, constructing a first association network based on the first-party data and a second association network based on the second-party data;

[0056] For example, referring to Figure 3 , Party A and Party B respectively form nodes and edges in the association network based on their own data. Assume that Party A is a bank, and forms an association network for Party A as a transfer association network through the transfer data between users. Among them, the user mobile phone number is the node in the association network, and the nodes with transfer relationships are connected by edges, and the transfer amount is the edge weight value between the nodes. Party B is an operator, and forms an association network for Party B as a call association network through the call data between users. The user mobile phone number is the node in the association network, and the nodes with call records are connected by edges, and the call times are the edge weight values between the nodes. Optionally, the edge weight values of their respective association networks can be normalized.

[0057] Step 220, associating the first association network and the second association network based on a secure intersection protocol to obtain a federated association network;

[0058] In one embodiment, the above step 220 further includes: performing encrypted intersection on the first-party data and the second-party data, determining the common nodes in the first association network and the second association network, and associating the first association network and the second association network according to the common nodes to obtain a federated association network.

[0059] For example, referring to Figure 4 , a secure intersection algorithm (such as a privacy intersection algorithm based on RSA+HASH) can be used to perform a secure intersection on the two-party node data, discover the common nodes without exposing the original data, and thus form a virtual federated association network. As shown in Figure 4 , it can be Figure 3The first associated network and the second associated network in it are associated to obtain the federated associated network. Among them, Va represents the nodes of Party A, Vb represents the nodes of Party B, Vab represents the common nodes of both parties, and Vab1, Vab2, and Vab3 are the common nodes of both parties. Taking the node Vab1 as an example, looking at it from the associated network of Party A alone, there are 2 neighbor nodes of Vab1. From the perspective of the global data, there are 4 neighbor nodes of Vab1.

[0060] Step 230, iteratively execute multiple rounds of label propagation on the nodes of the federated associated network;

[0061] Specifically, the nodes of the federated associated network can include labeled nodes and unlabeled nodes. For example, there can be a part of unlabeled nodes in the first associated network and the second associated network respectively. Another example is that all the nodes in the first associated network are labeled nodes, and all the nodes in the second associated network are unlabeled nodes. And so on.

[0062] In one implementation, for the case where the above-mentioned federated associated network includes both labeled nodes and unlabeled nodes at the same time, the above step 230 can further include the following steps: First, determine the labeled nodes and unlabeled nodes of the federated associated network; update the labels of the unlabeled nodes round by round until the labels of the unlabeled nodes no longer change and / or exceed the update round threshold; and keep the labels of the labeled nodes unchanged. In this way, the labels of the original samples can be ensured to remain unchanged, and the accuracy of label propagation can be guaranteed.

[0063] Optionally, the labels of the labeled nodes can also be updated dynamically round by round, that is, update the labels of all the nodes in the federated associated network round by round until the labels of the nodes no longer change and / or exceed the update round threshold. In this way, the original labels can be corrected and the hidden risk labels can be mined.

[0064] In the above step 230, each round of label propagation specifically includes the following steps 231-232:

[0065] Step 231, determine the label propagation probability between adjacent nodes in the federated associated graph;

[0066] In one implementation, the above step 231 can specifically include:

[0067] (1) Determine the edge weight w of the edge ij between the node i and its neighbor node j in the federated associated network ij ;

[0068] (2) Determine the sum of the edge weights ∑ of the edges between the node i and all its neighbor nodes J j w ij ;

[0069] In one implementation, if node i is a non-shared node, all neighbor nodes J represent all neighbor nodes of the graph where node i is located.

[0070] In one implementation, if node i is a shared node, all neighbor nodes J represent the set of all neighbor nodes a of node i in the first associated network and all neighbor nodes b of node i in the second associated network.

[0071] Further, if node i is a shared node, the first party and the second party interact on the sum of edge weights ∑ a w ia between node i and all neighbor nodes a of the first associated network, and b the sum of edge weights ∑ ib w j between node i and all neighbor nodes b of the second associated network. In this way, both the first party and the second party can calculate the sum of edge weights ∑ ij between node i and all its neighbor nodes J based on the sum of two-party edge weights interacted by both parties.

[0072] Further, in one implementation, if node i is a shared node, based on the sum of two-party edge weights interacted by both parties, both the first party and the second party can use the following formula to determine the sum of edge weights ∑ j w ij :

[0073] ∑ j w ij = ∑ a w ia + ∑ b w ib ;

[0074] where, ∑ a w ia is the sum of edge weights between node i and all neighbor nodes a of the first associated network, and ∑ b w ib is the sum of edge weights between node i and all neighbor nodes b of the second associated network.

[0075] Optionally, in another implementation, the influence of the business scenario on the tightness of node relationships can be further considered. For example, in a financial scenario, the transfer relationship is a strong relationship and the call relationship is a weak relationship. Therefore, when calculating the label propagation probability of an edge, the strength of the edge relationship in different business scenarios of both parties can be considered for weighted aggregation.

[0076] In this case, both the first party and the second party can use the following formula to determine the weighted sum of edge weights ∑ j w ij :

[0077] ∑ jw ij = θ a ∑ a w ia + θ b ∑ b w ib ;

[0078] Among them, θ a is the graph weight corresponding to the first associated network, and θ b is the graph weight of the second associated network.

[0079] (3) Determine the label propagation probability P ij of neighbor node j for node i according to the ratio of edge weight w j to the sum of edge weights ∑ ij w ij .

[0080] Specifically, the label propagation probability of each edge in the federated associated network where w ij represents the weight value of edge ij. Here, for non-shared nodes, J represents the neighbor nodes of node i; for shared nodes, J represents all neighbor nodes of node i on both Party A and Party B sides.

[0081] ∑ j w ij The calculation logic of is that Party A calculates the sum of neighbor node weights of its own node i as ∑ a w ia ; Party B calculates the sum of neighbor node weights of its own node i as ∑ b w ib . The two parties interact with ∑ a w ia and ∑ b w ib to obtain the final denominator value of the weight calculation as ∑ j w ij = ∑ a w ia + ∑ b w ib .

[0082] Reference Figure 5, taking the node Vab2 as an example here. In Party A's local network, Vab2 has 1 neighbor node, and the label propagation probability of this party alone is calculated as P = 0.1 / 0.1 = 1; in Party B's local network, Vab2 has 3 neighbor nodes, and the label propagation probabilities of this party alone are calculated as P = 0.2 / (0.2 + 0.4 + 0.8) = 1 / 7, P = 0.4 / (0.2 + 0.4 + 0.8) = 2 / 7, P = 0.8 / (0.2 + 0.4 + 0.8) = 4 / 7 respectively; further, the two parties exchange the weight values of the target node Vab2 and its neighbor nodes, with Party A being 0.1 and Party B being 0.2 + 0.4 + 0.8 = 1.4. Combining with the federated association network, the label propagation probability of the target node is updated.

[0083] Reference Figure 6 , after the above calculations, in Party A, the label propagation probability of its neighbor nodes for this node i In Party B, similarly, the label propagation probabilities of its neighbor nodes for this node i are 2 / 15, 4 / 15, and 8 / 15 respectively.

[0084] Step 232, for each node, determine the current round label of each node according to the current round labels of its neighbor nodes and the label propagation probabilities of its neighbor nodes for this node.

[0085] Reference Figure 7 , continuing to take the federated association network formed by Vab2 as an example here, as follows. Node 5 is a risk node, shown as a white node with the label set to "1"; the remaining nodes are unknown nodes, shown as gray nodes with the label set to "0". During the label propagation process, the label of node 5 remains "1" all the time, while the labels of the remaining nodes are updated round by round until the labels of all nodes no longer change or exceed the update round threshold.

[0086] In one implementation, in the above step 232, the following steps are further included:

[0087] Step 2321, for node i, determine the current round label of each neighbor node of node i and the label propagation probability of each neighbor node for node i;

[0088] Step 2322, among all the neighbor nodes of node i, calculate the sum of the label propagation probabilities corresponding to each label to obtain the aggregated label propagation probability corresponding to each label;

[0089] Specifically, if node i is a non - shared node, only the party where node i is located calculates the aggregated label propagation probability corresponding to each label of its neighbor nodes.

[0090] Specifically, if node i is a common node, the following steps are executed: First, the first party calculates the first-party label propagation aggregation probability corresponding to the labels of all neighbor nodes of node i in the first associated network; the second party calculates the second label propagation aggregation probability corresponding to the labels of all neighbor nodes of node i in the second associated network; Second, the first party and the second party exchange the first label propagation aggregation probability and the second label propagation aggregation probability; Finally, the first party and the second party each perform label propagation probability aggregation again based on the exchanged information to obtain the label propagation aggregation probability corresponding to each label.

[0091] Step 2323, update the label of node i in this round according to the label with the maximum label propagation aggregation probability.

[0092] The node label update rule shown in the above Step 2321 - Step 2323 may include the following specific steps:

[0093] First, for the T-th update of node i, let its neighbor node set be J((J1, L1, P i1 ), (J2, L2, P i2 ), (J j , L j , P ij )..., (J n , L n , P in )>, where J j is the identifier of neighbor node j, L j is the label of neighbor node j, and P ij is the propagation probability of the edge <i, J j >.

[0094] Second, calculate the aggregated propagation probability of all labels in the neighbor node set. Specifically, P(L j ) = ∑P ij , where P ij is the label propagation probability of the neighbor node with label L j for the target node i. Among them, if node i is a non-common node, only the propagation probability of the labels of its own neighbor nodes needs to be calculated. If node i is a common node, Party A calculates the probability P(L aj ) = ∑P iaj corresponding to the labels of all its own neighbor nodes, Party B calculates the probability P(L bj ) = ∑P ibj corresponding to the labels of its own neighbor nodes, Party A and Party B exchange P(L aj ) and P(L bj ), and each performs label propagation probability aggregation again on its own side to obtain the final label propagation aggregation probability P(L j ) = P(Laj ) + P(L bj ).

[0095] Finally, select the largest P(L j ) corresponding to the label L j as the label of node i in this round. Repeat the above steps until the labels of all nodes no longer change.

[0096] In a specific example, refer to Figure 7 and Figure 8 to give a specific calculation example of label update.

[0097] Refer to Figure 7 , for the first round of propagation, for node Vab2, perform the following calculations:

[0098] (1) Calculate the label propagation aggregation probability of Party A's neighbor nodes as <"0", 1 / 15>, where "0" represents the risk-free label and 1 / 15 represents the label propagation aggregation probability corresponding to the label "0". It can be understood that since Party A's Vab2 has only one neighbor node 1 and its initial label value is "0", and the label propagation probability of node 1 to node Vab2 has been calculated as 1 / 15 in the above text, so for Party A's Vab2 node, there is only one propagable label "0", and the label propagation aggregation probability corresponding to this propagable label "0" is 1 / 15.

[0099] (2) Calculate the label propagation aggregation probabilities of Party B's neighbor nodes as <"0", 6 / 15>, <"1", 8 / 15>, where "0" represents the risk-free label, 6 / 15 represents the label propagation aggregation probability corresponding to the label "0", "1" represents the risky label, and 8 / 15 represents the label propagation aggregation probability corresponding to the label "1". Since Party B's Vab2 has three neighbor nodes (3, 4, 5), and the initial label values of nodes 3 and 4 are "0", and the initial label value of node 5 is "1". In the above text, the label propagation probability of node 3 to node Vab2 has been calculated as 2 / 15, the label propagation probability of node 4 to node Vab2 has been calculated as 4 / 15, and the label propagation probability of node 5 to node Vab2 has been calculated as 8 / 15. Therefore, for Party B's Vab2 node, there are 2 propagable labels "0" and "1", and the label propagation aggregation probability corresponding to this propagable label "0" is 6 / 15 = 2 / 15 + 4 / 15, and the label propagation aggregation probability corresponding to this propagable label "1" is 8 / 15.

[0100] (3) The two parties exchange the label propagation aggregation probabilities, and accumulate the label propagation aggregation probabilities corresponding to the same label. They can each calculate that the label propagation aggregation probabilities of node Vab2 are <"0", 7 / 15> and <"1", 8 / 15>, that is, the label propagation aggregation probability corresponding to the propagable label "0" is 7 / 15, and the label propagation aggregation probability corresponding to the propagable label "1" is 8 / 15.

[0101] (4) Select the label "1" corresponding to the maximum label propagation aggregation probability <"1", 8 / 15> as the label of node Vab2 in this round. Other nodes are similar to the above steps.

[0102] After the first round of propagation, the updated node label distribution diagram of this federated association network is as Figure 8 shown, where nodes Vab1 and Vab2 are both updated to the label "1". Continue the next round of label propagation until the labels of the nodes no longer change or the number of propagation rounds is greater than a certain threshold.

[0103] In one implementation, according to the closeness of the node relationships between the first association network and the second association network, determine the graph weights of the first association network and the second association network; and introduce the graph weights during the interaction between the first association network and the second association network.

[0104] For example, it is possible to determine whether the first association network and the second association network are strongly associated or weakly associated according to the business scenario. Furthermore, when calculating the label propagation probability of the edge, the graph weights of the first association network and the second association network can be introduced. Of course, it is also possible to introduce the graph weights of the first association network and the second association network when calculating the label propagation aggregation probability of each type of label. This application does not make specific restrictions on this.

[0105] In one implementation, introducing graph weights during the interaction between the first association network and the second association network includes at least the following two introduction methods:

[0106] (1) In the above step 231, if node i is a common node, use the following formula to determine the sum of edge weights ∑ j w ij :

[0107] ∑ j w ij =θ a ∑ a w ia +θ b ∑ b w ib ; where ∑ a w ia is the sum of the edge weights between node i and all neighbor nodes a of the first association network, ∑ b wib is the sum of the edge weights between node i and all neighbor nodes b of the second associated network, θ a is the graph weight of the first associated network, θ b is the graph weight of the second associated network.

[0108] (2) In step 232 above, after the first party and the second party exchange the first label propagation aggregation probability and the second label propagation aggregation probability, label propagation probability aggregation is performed again based on the graph weights of the first associated network and the second associated network to obtain the label propagation aggregation probability corresponding to each label. For example, if node i is a common node, Party A calculates the probability P(L aj ) = ∑P iaj corresponding to the labels of all its neighbor nodes, Party B calculates the probability P(L bj ) = ∑P ibj corresponding to the labels of its neighbor nodes, Party A and Party B exchange P(L aj ) and P(L bj ), and each performs label propagation probability aggregation again on its own side to obtain the final label propagation aggregation probability P(L j ) = θ a P(L aj ) + θ b P(L bj ).

[0109] In one implementation, it further includes: if the first associated network and the second associated network are directed graph networks, only the in-neighbor nodes of each node are regarded as neighbor nodes. For example, for a directed graph, when calculating the propagation probability of a node, only the in-neighbor nodes of the target node can be considered. Specifically, it can be judged in combination with the business scenario.

[0110] In the description of this specification, the descriptions referring to terms such as "some possible implementations", "some implementations", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the implementation or example are included in at least one implementation or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same implementation or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more implementations or examples. In addition, without contradiction, those skilled in the art can combine and combine the different implementations or examples described in this specification and the features of different implementations or examples.

[0111] Regarding the method flowchart of the embodiments of the present application, certain operations are described as different steps executed in a certain order. Such a flowchart is illustrative rather than restrictive. Certain steps described herein can be grouped together and executed in a single operation, certain steps can be split into multiple sub-steps, and certain steps can be executed in an order different from that shown herein. Each step shown in the flowchart can be implemented in any way by any circuit structure and / or tangible mechanism (e.g., software running on a computer device, hardware (e.g., a processor or chip-implemented logic function), etc., and / or any combination thereof).

[0112] Based on the same inventive concept, an embodiment of the present invention further provides a label propagation device for an associated network, which is used to execute the label propagation method for the associated network provided in any of the foregoing embodiments. Figure 9 Schematic diagram of the structure of a label propagation device for an associated network provided by an embodiment of the present invention.

[0113] As Figure 9 shown, the device 900 includes:

[0114] A graph construction module 910, configured to construct a first associated network based on first-party data and construct a second associated network based on second-party data;

[0115] A federated network module 920, configured to associate the first associated network and the second associated network based on a secure intersection protocol to obtain a federated associated network;

[0116] A label propagation module 930, configured to iteratively execute multiple rounds of label propagation on the nodes of the federated associated network; wherein, each round of the label propagation includes: determining the label propagation probability between adjacent nodes in the federated associated graph; for each node, determining the current-round label of each node according to the current-round labels of the neighbor nodes and the label propagation probability of the neighbor nodes for the node.

[0117] It should be noted that the device in the embodiments of the present application can implement each process of the foregoing method embodiments and achieve the same effects and functions, which will not be elaborated herein.

[0118] According to some embodiments of the present application, a non-volatile computer storage medium for the label propagation method of an associated network is provided, on which computer-executable instructions are stored, and the computer-executable instructions are set to execute: the method described in the foregoing embodiments when run by a processor.

[0119] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiments of the device, equipment, and computer-readable storage medium, since they are basically similar to the method embodiments, their descriptions are simplified, and reference can be made to the corresponding parts of the method embodiments for relevant content.

[0120] The device, equipment, and computer-readable storage medium provided by the embodiments of this application correspond one by one to the method. Therefore, the device, equipment, and computer-readable storage medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device, equipment, and computer-readable storage medium will not be elaborated here.

[0121] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device (equipment or system), or a computer-readable storage medium. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer-readable storage medium implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0122] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (equipment or systems), and computer-readable storage media according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process Figure 1 one process or a plurality of processes and / or Figure 1 steps for implementing the functions specified in a block or a plurality of blocks.

[0125] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0126] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0127] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. Additionally, although the operations of the method of the present invention are depicted in the figures in a particular order, this is not required or implied to perform the operations in that particular order, or to perform all of the illustrated operations to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step and performed, and / or one step may be decomposed into multiple steps and performed.

[0128] Although the spirit and principles of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for the convenience of expression. The present invention aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A method for label propagation in an associated network, characterized in that, Including: Construct a first association network based on first-party data and a second association network based on second-party data; Associate the first association network and the second association network based on a secure intersection protocol to obtain a federated association network; Iteratively perform multiple rounds of label propagation on the nodes of the federated association network; Among them, each round of the label propagation includes: determining the edge weight w of the edge ij between the node i and its neighbor node j in the federated association network ij ; Determine the sum of edge weights ∑ between the node i and all its neighbor nodes J j w ij ; According to the edge weight w ij and the ratio of the sum of the edge weights ∑ j w ij to determine the label propagation probability P of the neighbor node j for the node i ij ; For each node, determine the current round label of each node according to the current round label of its neighbor nodes and the label propagation probability of the neighbor nodes for this node.

2. The method according to claim 1, characterized in that, Associating the first association network and the second association network based on a secure intersection protocol to obtain a federated association network further includes: Perform encrypted intersection on the first-party data and the second-party data, determine the common nodes in the first association network and the second association network, and associate the first association network and the second association network according to the common nodes to obtain a federated association network.

3. The method according to claim 1, characterized in that, If the node i is a non-common node, all neighbor nodes J represent all neighbor nodes of the graph where the node i is located.

4. The method according to claim 1, characterized in that, If the node i is a common node, all neighbor nodes J represent the set of all neighbor nodes a of the node i in the first association network and all neighbor nodes b of the node i in the second association network.

5. The method according to claim 1, characterized in that, If the node i is a common node, the first party and the second party exchange the sum of the edge weights between the node i and all neighbor nodes a of the first association network and the sum of the edge weights between the node i and all neighbor nodes b of the second association network.

6. The method according to claim 1, characterized in that, Further including: If the node i is a shared node, use the following formula to determine the sum of the edge weights ∑ j w ij : ∑ j w ij =∑ a w ia +∑ b w ib ; Among them, ∑ a w ia is the sum of edge weights between the node i and all neighbor nodes a of the first associated network, and the ∑ b w ib is the sum of edge weights between the node i and all neighbor nodes b of the second associated network.

7. The method according to claim 1, characterized in that, Iteratively performing multiple rounds of label propagation on the nodes of the federated association network further includes: Determine the labeled nodes and unlabeled nodes of the federated association network; Update the labels of the unlabeled nodes round by round until the labels of the unlabeled nodes no longer change and / or exceed the update round threshold; and, Keep the labels of the labeled nodes unchanged.

8. The method according to claim 1, characterized in that, Determining the current round label of each node according to the current round label of its neighbor nodes and the label propagation probability of the neighbor nodes for this node includes: For each node, determine the current round label of each neighbor node of the node and the label propagation probability of each neighbor node for this node; Among all neighbor nodes of the node, calculate the sum of the label propagation probabilities corresponding to each label to obtain the aggregated label propagation probability corresponding to each label; Update the current round label of the node according to the label with the maximum aggregated label propagation probability.

9. The method according to claim 8, characterized in that, If the node is a non-common node, the method further includes: The party where the node is located calculates the aggregated label propagation probability corresponding to each neighbor node label of the node.

10. The method according to claim 8, characterized in that, If the node is a common node, the method further includes: The first party calculates the first aggregated label propagation probability corresponding to all neighbor node labels of the node in the first association network; The second party calculates the second aggregated label propagation probability corresponding to all neighbor node labels of the node in the second association network; The first party and the second party exchange the first aggregated label propagation probability and the second aggregated label propagation probability; The first party and the second party each perform label propagation probability aggregation again based on the interaction information to obtain the label propagation aggregation probability corresponding to each type of label.

11. The method according to claim 1, wherein, Further included: Determine the graph weights of the first association network and the second association network according to the tightness of the node relationships in the first association network and the second association network; And, Introduce the graph weights during the interaction process of the first association network and the second association network.

12. The method according to claim 11, wherein, Introducing the graph weights during the interaction process of the first association network and the second association network includes: If the node i is a shared node, use the following formula to determine the sum of the edge weights ∑ j w ij : ∑ j w ij = θ a ∑ a w ia + θ b ∑ b w ib ; Among them, ∑ a w ia is the sum of edge weights between the node i and all neighbor nodes a of the first associated network, and the ∑ b w ib is the sum of edge weights between the node i and all neighbor nodes b of the second associated network, and θ a is the graph weight of the first associated network, and θ b is the graph weight of the second associated network.

13. The method according to claim 11, wherein, Introducing the graph weights during the interaction process of the first association network and the second association network includes: After the first party and the second party exchange the first label propagation aggregation probability and the second label propagation aggregation probability, perform label propagation probability aggregation again based on the graph weights of the first association network and the second association network to obtain the label propagation aggregation probability corresponding to each type of label.

14. The method according to claim 1, wherein, Further included: If the first association network and the second association network are directed graph networks, only use the incoming neighbor nodes of each node as the neighbor nodes.

15. A label propagation device for an association network, wherein, Including: A graph construction module for constructing a first association network based on first-party data and a second association network based on second-party data; A federated network module for associating the first association network and the second association network based on a secure intersection protocol to obtain a federated association network; A label propagation module for iteratively performing multiple rounds of label propagation on the nodes of the federated association network; wherein, each round of the label propagation includes: Determine the edge weight w of the edge ij between node i and its neighbor node j in the federated association network ij ; Determine the sum of edge weights ∑ between the node i and all its neighbor nodes J j w ij ; According to the edge weight w ij and the sum of edge weights ∑ j w ij to determine the label propagation probability P of the neighbor node j for the node i ij ; For each node, determine the current round label of each node according to the current round labels of the neighbor nodes and the label propagation probability of the neighbor nodes for the node.

16. A label propagation device for an association network, wherein, Including: At least one processor; And a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute: the method according to any one of claims 1-14.

17. A computer-readable storage medium storing a program, which, when executed by a multi-core processor, causes the multi-core processor to execute the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Insurance customer recommendation method and system based on federal label propagation

    CN113095946A

  • Multi-dimensional graph network node clustering processing method, device and equipment

    CN113254717A