A multi-party computation method and apparatus based on graph computation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2026-08-14
AI Technical Summary
[0036]在以上实施例中,通过构建用于描述多方计算的数据协作关系的知识图谱,并针对构建得到的知识图谱进行图计算,可以求解出所述知识图谱中的各条边对应的权重值;其中,由于所述知识图谱中的各条边可以用于表示所述多方计算的数据协作关系,所述各条边对应的权重值可以用于指示各条边所表示的数据协作关系的紧密程度,因此,基于求解出的所述各条边的权重值,可以从与所述知识图谱中的主节点连接的多个从节点对应的多个数据提供方中,确定出满足预设条件的至少一个数据提供方来参与所述多方计算。
Smart Images

Figure CN115757571B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of multi-party computation technology, and more particularly to a multi-party computation method, apparatus, electronic device, and machine-readable storage medium based on graph computation. Background Technology
[0002] Enterprise users can leverage data sharing platforms to conduct multi-party collaborative data value mining without leaving their private original data domain.
[0003] For example, data providers can provide data description information corresponding to their own original data to the data sharing platform; data users can select data that they need based on the data description information and build privacy computing applications to realize their data usage needs through the privacy computing applications combined with the selected data; furthermore, after approval by the data provider, the privacy computing application can combine the data from all parties and use cryptographically secure methods to achieve multi-party collaborative data value mining without leaving the domain of each party's private original data.
[0004] Currently, a data collaboration network built upon the aforementioned data sharing platform is beginning to take shape. With the continuous development of privacy data protection technologies, the data sharing platform can offer users increasingly richer data collaboration functions, and more and more data providers are beginning to connect to the data collaboration network to solve the data silo problem. Therefore, in the process of building privacy computing applications, the number of data sources that can be used as input data for these applications is also increasing.
[0005] Therefore, it is evident that selecting suitable data providers from a vast pool of data providers to participate in multi-party computation has become an urgent problem to be solved. Summary of the Invention
[0006] This application provides a graph-based multi-party computation method, which is applied to a service device corresponding to the initiator of the multi-party computation; the method includes:
[0007] A knowledge graph is constructed to describe the data collaboration relationships of the multi-party computation; wherein the knowledge graph includes a master node corresponding to the initiator of the multi-party computation, slave nodes corresponding to each data provider among the multiple data providers of the multi-party computation, and edges connecting each slave node to the master node; each edge represents the data collaboration relationship of the multi-party computation; each data provider has registered on the service device corresponding to the initiator;
[0008] Graph computation is performed on the constructed knowledge graph to solve for the weight value corresponding to each edge in the knowledge graph; wherein, the weight value corresponding to each edge is used to indicate the tightness of the data collaboration relationship represented by each edge;
[0009] Based on the weight values corresponding to each edge, at least one data provider that meets the preset conditions is determined from the plurality of data providers, and the multi-party calculation is performed based on the user data provided by the at least one data provider.
[0010] Optionally, the data content corresponding to each slave node in the knowledge graph includes the identity identifier of each data provider corresponding to each slave node, and the data characteristics of the user data provided by each data provider;
[0011] The data content corresponding to each edge in the knowledge graph includes an operator identifier sequence corresponding to the multi-party computation; wherein, the operator identifier sequence includes operator identifiers of at least one operator concatenated according to the computation order specified by the initiator.
[0012] Optionally, the data characteristics of the user data include:
[0013] The encrypted data corresponding to the user data; or...
[0014] The hash value of the user data; or,
[0015] The data attribute information of the user data.
[0016] Optionally, determining at least one data provider that meets preset conditions from the plurality of data providers based on the weight values corresponding to each edge includes:
[0017] From the plurality of data providers, at least one data provider whose weight value corresponding to the edge between the slave node and the master node is greater than a preset threshold is determined as a data provider that meets the preset condition; or...
[0018] From the plurality of data providers, the data providers with the largest number of weight values corresponding to the edges between the slave nodes and the master nodes are determined as data providers that meet the preset conditions.
[0019] Optionally, the at least one data provider corresponds to a different data domain; the multi-party computation includes multi-party secure computation performed based on the ciphertext data corresponding to the user data provided by the at least one data provider;
[0020] The multi-party computation based on user data provided by the at least one data provider includes:
[0021] Multi-party secure computation is performed based on encrypted data transferred across domains by at least one data provider to the service device corresponding to the initiator.
[0022] Optionally, the initiator includes:
[0023] A data sharing platform corresponding to the aforementioned multiple data providers; or,
[0024] In a data collaboration network that includes the aforementioned multiple data providers, any data provider has multi-party computing needs.
[0025] Optionally, the data sharing platform includes a blockchain service platform; the blockchain nodes in the blockchain include service devices corresponding to the multiple data providers respectively;
[0026] The user data provided by each data provider is stored locally on the service device corresponding to each data provider; the data attribute information of the user data provided by each data provider is stored on the blockchain.
[0027] Optionally, the blockchain service platform includes a blockchain cloud service platform;
[0028] The service devices corresponding to each data provider include virtual service devices created for each data provider on the blockchain cloud service platform.
[0029] This application also provides a graph-based multi-party computation apparatus, which is applied to a service device corresponding to the initiator of the multi-party computation; the apparatus includes:
[0030] A construction unit is used to construct a knowledge graph describing the data collaboration relationships of the multi-party computation; wherein the knowledge graph includes a master node corresponding to the initiator of the multi-party computation, slave nodes corresponding to each data provider among the multiple data providers of the multi-party computation, and edges connecting each slave node to the master node; each edge represents the data collaboration relationship of the multi-party computation; each data provider has registered on the service device corresponding to the initiator;
[0031] The graph computation unit is used to perform graph computation on the constructed knowledge graph to solve for the weight value corresponding to each edge in the knowledge graph; wherein, the weight value corresponding to each edge is used to indicate the tightness of the data collaboration relationship represented by each edge;
[0032] A multi-party computation unit is used to determine at least one data provider that meets preset conditions from the plurality of data providers based on the weight values corresponding to each edge, and to perform the multi-party computation based on the user data provided by the at least one data provider.
[0033] This application also provides an electronic device, including a communication interface, a processor, a memory, and a bus, wherein the communication interface, the processor, and the memory are interconnected via the bus;
[0034] The memory stores machine-readable instructions, and the processor executes the above method by invoking the machine-readable instructions.
[0035] This application also provides a machine-readable storage medium storing machine-readable instructions, which, when called and executed by a processor, implement the above-described method.
[0036] In the above embodiments, by constructing a knowledge graph to describe the data collaboration relationships in multi-party computation, and performing graph computation on the constructed knowledge graph, the weight values corresponding to each edge in the knowledge graph can be solved. Since each edge in the knowledge graph can be used to represent the data collaboration relationships in the multi-party computation, and the weight values corresponding to each edge can be used to indicate the tightness of the data collaboration relationships represented by each edge, based on the solved weight values of each edge, at least one data provider that meets the preset conditions can be determined from multiple data providers corresponding to multiple slave nodes connected to the master node in the knowledge graph to participate in the multi-party computation.
[0037] In the above manner, in multi-party computation scenarios, the service device corresponding to the initiator can automatically identify the data providers with closer data collaboration relationships from multiple data providers, and perform the multi-party computation based on the user data provided by the data providers. This allows for the selection of input data for the multi-party computation that better meets the initiator's needs, while avoiding the leakage of private user data from each data provider, thus improving the user's multi-party data collaboration experience. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a schematic diagram illustrating the architecture of a data collaboration network, as shown in an exemplary embodiment.
[0040] Figure 2 This is a flowchart illustrating an exemplary embodiment of a multi-party computation method based on graph computation;
[0041] Figure 3 This is a schematic diagram of a knowledge graph, as illustrated in an exemplary embodiment.
[0042] Figure 4 This is an exemplary embodiment illustrating the structure of an electronic device containing a graph-based multi-party computing device;
[0043] Figure 5 This is a block diagram illustrating a graph-based multi-party computing device, as shown in an exemplary embodiment. Detailed Implementation
[0044] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0045] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.
[0046] With the continuous development of internet technology, enterprise users are accumulating more and more data. On the one hand, enterprise users hope to realize data assetization, that is, to monetize the data they hold, provided that they comply with relevant laws and regulations; on the other hand, enterprise users also hope to combine data held by other enterprises to conduct data mining, data collaboration, knowledge integration, and other activities from more dimensions to enrich the value of their own data.
[0047] In practical applications, because data is different from physical objects, it is easily copied and transferred arbitrarily during sharing and collaboration. Therefore, in order to protect their ownership of the data they hold, and for compliance reasons, enterprise users usually need to ensure that the original data they hold (such as the original plaintext data) does not leave the private domain.
[0048] In this context, enterprise users can leverage data sharing platforms to conduct collaborative data value mining within the domain of their private original data. For example, data providers can offer data descriptions corresponding to their own original data to the data sharing platform. Data users can then select data they need based on these descriptions and build privacy-preserving computation applications. These applications, combined with the selected data, fulfill their data usage requirements. Furthermore, after approval from the data provider, the privacy-preserving computation application can combine data from all parties and employ cryptographically secure methods to achieve collaborative privacy-preserving computation within the domain of each party's private original data.
[0049] The data sharing platform described herein can leverage privacy-preserving technologies such as Federated Learning (FL), Secure Multi-Party Computation (SMPC), Private Set Intersection (PSI), Trusted Execution Environment (TEE), and Differential Privacy (DP) to provide privacy-preserving computing services for data value analysis and mining while protecting privacy information. This data sharing platform may also be referred to as a data sharing service platform, a trusted data collaboration platform, or a privacy computing service platform, etc., and this specification does not impose any specific limitations on these terms.
[0050] Currently, a data collaboration network built upon the aforementioned data sharing platform is beginning to take shape. With the continuous development of privacy data protection technologies, the data sharing platform can offer users increasingly richer data collaboration functions, and more and more data providers are beginning to connect to the data collaboration network to solve the data silo problem. Therefore, in the process of building privacy computing applications, the number of data sources that can be used as input data for these applications is also increasing.
[0051] Therefore, it is evident that selecting suitable data providers from a vast pool of data providers to participate in multi-party computation has become an urgent problem to be solved.
[0052] For example, banking institutions typically hold user data related to financial services, such as personal asset information and personal loan repayment records; while internet companies typically hold user data related to online services, such as personal web browsing history and personal online shopping records. The banking institutions and internet companies can each act as data providers, performing multi-party secure computations through the data sharing platform, such as jointly training a credit model. This allows for the full utilization of multi-party data without exposing their respective private raw data to other participants in the multi-party secure computation, thereby extracting data value from multiple dimensions.
[0053] In the embodiments shown above, since there are a large number of banking institutions connected to the data collaboration network, internet companies hope to be able to perceive which banking institutions have higher user stickiness, that is, which banking institutions provide user data that is more suitable for participating in multi-party secure computation, so as to discover better cooperation opportunities. However, in the scenario of multi-party secure computation, internet companies cannot obtain the private original data of each banking institution, and it is difficult to identify which banking institutions are more suitable for data collaboration from among the many banking institutions.
[0054] It should be noted that the scenario of multi-party secure computation between banking institutions and internet companies shown above is merely an exemplary description and does not impose any particular limitations on this specification.
[0055] In view of this, this specification aims to propose a technical solution that first determines the closeness of the data cooperation relationship among multiple data providers in a multi-party computation based on graph computation, and then selects some data providers from the multiple data providers to participate in the multi-party computation.
[0056] In implementation, a knowledge graph describing the data collaboration relationships in multi-party computation can be constructed first. This knowledge graph may include a master node corresponding to the initiator of the multi-party computation, slave nodes corresponding to each data provider among the multiple data providers in the multi-party computation, and edges connecting the slave nodes to the master node. Each edge represents the data collaboration relationship in the multi-party computation. Each data provider is registered on a service device corresponding to the initiator. Further, the service device corresponding to the initiator can perform graph computation on the constructed knowledge graph to solve for the weight values corresponding to each edge in the knowledge graph. The weight values of each edge can indicate the closeness of the data collaboration relationship represented by each edge. Further, the service device corresponding to the initiator can determine at least one data provider that meets preset conditions from among the multiple data providers based on the weight values of each edge, and perform the multi-party computation based on user data provided by the at least one data provider.
[0057] Therefore, in the technical solution of this specification, by constructing a knowledge graph to describe the data collaboration relationships of multi-party computation, and performing graph computation on the constructed knowledge graph, the weight values corresponding to each edge in the knowledge graph can be solved. Since each edge in the knowledge graph can represent the data collaboration relationships of the multi-party computation, and the weight values corresponding to each edge can indicate the closeness of the data collaboration relationships represented by each edge, based on the solved weight values of each edge, at least one data provider that meets preset conditions can be determined from multiple data providers corresponding to multiple slave nodes connected to the master node in the knowledge graph to participate in the multi-party computation.
[0058] In the above manner, in multi-party computation scenarios, the service device corresponding to the initiator can automatically identify the data providers with closer data collaboration relationships from multiple data providers, and perform the multi-party computation based on the user data provided by the data providers. This allows for the selection of input data for the multi-party computation that better meets the initiator's needs, while avoiding the leakage of private user data from each data provider, thus improving the user's multi-party data collaboration experience.
[0059] The present application will now be described through specific embodiments and in conjunction with specific application scenarios.
[0060] In this specification, the multiple data providers of the multi-party computation have been registered on the service device corresponding to the initiator of the multi-party computation.
[0061] The multiple data providers in the multi-party computation may include multiple data providers that meet the data usage requirements determined from a number of registered data providers based on the data usage requirements of the multi-party computation; or, the multiple data providers in the multi-party computation may also include multiple data providers that may participate in the multi-party computation, designated by the initiator of the multi-party computation from a number of registered data providers.
[0062] In practical applications, in multi-party computation scenarios, in order to protect the privacy and security of user data, the private raw data of each data provider usually does not leave the domain. Therefore, the multiple data providers of the multi-party computation can register on the service device corresponding to the initiator of the multi-party computation, so that the initiator can use the registered data providers as optional data sources for the multi-party computation.
[0063] In this specification, the service device may include physical devices or virtual devices implemented in a server or server cluster.
[0064] For example, the service device corresponding to the initiator can be a physical host in a server cluster, or a virtual machine instance created by virtualizing the hardware resources on the server cluster based on virtualization technology.
[0065] In practical applications, the service devices corresponding to the initiator of the multi-party computation and the service devices corresponding to the multiple data providers of the multi-party computation can be coupled together through various types of communication networks (such as wired and / or wireless communication networks) and various types of communication methods (such as TCP / IP) to form a data collaboration network related to the multi-party computation. This data collaboration network can be a distributed network; after registering in the data collaboration network, each data provider can obtain its own DID (Decentralized Identity).
[0066] In one embodiment shown, the initiator of the multi-party computation may include a data sharing platform corresponding to the plurality of data providers.
[0067] For example, see Figure 1 , Figure 1 This is a schematic diagram illustrating the architecture of a data collaboration network, as shown in an exemplary embodiment. Figure 1 The data collaboration network shown may include a management node 101 corresponding to the data sharing platform, and data nodes 102, 103, and 104 corresponding to registered data providers A, B, and C, respectively. If data provider A, corresponding to data node 102, has a multi-party computation requirement, data provider A can also act as a data user and submit a data usage request related to the multi-party computation to the management node 101 corresponding to the data sharing platform. This allows the data sharing platform to determine, based on the data usage request, at least one data provider from among the registered data providers that needs to participate in the multi-party computation.
[0068] It should be noted that, Figure 1 Only three data nodes are shown as an example, which does not impose any special limitation on the technical solutions in this specification; in practical applications, any number of data providers can be registered on the data sharing platform, and any number of data nodes can be included in the data collaboration network.
[0069] The management node in the data collaboration network may include a service device corresponding to the initiator of the multi-party computation; each data node in the data collaboration network may include a service device corresponding to each data provider of the multi-party computation.
[0070] It should be noted that, in the embodiments shown above, the user data provided by each data provider can be stored locally on the data nodes corresponding to each data provider; while the data attribute information of the user data provided by each data provider can be stored on the management node corresponding to the data sharing platform, so that the data sharing platform can uniformly manage the data resources accessed to the data collaboration network even when it is unable to obtain the private original data (i.e., plaintext data of user data) held by each data provider.
[0071] In one possible embodiment, to improve the trustworthiness of data collaboration, the data sharing platform can utilize blockchain technology to provide data sharing services; alternatively, the data sharing platform can be a data sharing service provided by a blockchain service platform. In implementation, the data sharing platform may include a blockchain service platform; the blockchain nodes in the blockchain may include service devices corresponding to the multiple data providers; the user data provided by each data provider may be stored locally on the service device corresponding to each data provider; and the data attribute information of the user data provided by each data provider may be stored on the blockchain.
[0072] For example, such as Figure 1 As shown, data nodes 102, 103, and 104, which correspond to multiple data providers respectively, can be blockchain nodes; the user data provided by each data provider can be stored locally on the service device corresponding to each data provider, that is, it can be stored in the local database of the service device corresponding to each data user; and the data attribute information of the user data provided by each data provider can be stored on the blockchain.
[0073] In another possible embodiment, the data sharing platform may include a blockchain cloud service platform; the service devices corresponding to each data provider include virtual service devices created on the blockchain cloud service platform for each data provider. The virtual service devices may include virtual machines (i.e., virtual machine instances) created in a distributed cloud computing system.
[0074] For example, such as Figure 1 As shown, data nodes 102, 103, and 104, which correspond to multiple data providers respectively, can be virtual machine instances created on the blockchain cloud service platform for the multiple data providers respectively.
[0075] It should be noted that, in the embodiments shown above, when the data sharing platform may include a blockchain cloud service platform, data isolation between different data providers can be achieved through different virtual machine instances and different data domains. This can ensure the data privacy and security of each data provider, while fully utilizing the distributed cloud computing system to improve the performance of multi-party computing, and fully utilizing blockchain technology to ensure the trustworthiness of multi-party data collaboration.
[0076] In another embodiment shown, the initiator of the multi-party computation may include any data provider with a multi-party computation requirement in a data collaboration network that includes the plurality of data providers.
[0077] For example, such as Figure 1 As shown, if data provider A corresponding to data node 102 has a multi-party computation requirement, then data provider A can also act as a data user. Based on the data usage requirements related to the multi-party computation, at least one data provider that needs to participate in the multi-party computation is determined from among the registered data providers.
[0078] To enable those skilled in the art to better understand the technical solutions in the embodiments of this specification, the following description, in conjunction with... Figure 1 The schematic diagram shown illustrates an embodiment in this specification, using the data sharing platform as an example where the initiator of the multi-party computation is the data sharing platform. It should be noted that this is merely an exemplary description and does not specifically limit this specification; based on one or more embodiments shown in this specification, those skilled in the art can also obtain other embodiments where the initiator of the multi-party computation is any data provider.
[0079] Please see Figure 2 , Figure 2 This is a flowchart illustrating an exemplary embodiment of a multi-party computation method based on graph computation. The method can be applied to a service device corresponding to the initiator of the multi-party computation. The method may perform the following steps:
[0080] Step 202: Construct a knowledge graph to describe the data collaboration relationships of multi-party computation; wherein the knowledge graph includes a master node corresponding to the initiator of the multi-party computation, each slave node corresponding to each data provider among the multiple data providers of the multi-party computation, and each edge used to connect the slave nodes to the master node; each edge is used to represent the data collaboration relationships of the multi-party computation.
[0081] In step 202, each edge is used to represent the data collaboration relationship in the multi-party computation. That is, each edge in the knowledge graph can be used to represent the data collaboration relationship between various data providers in the multi-party computation.
[0082] For example, see Figure 3 , Figure 3 This is a schematic diagram illustrating a knowledge graph, as shown in an exemplary embodiment. Please refer to... Figure 1 and Figure 3 If the initiator of the multi-party computation is the data sharing platform corresponding to management node 101, and the multiple data providers of the multi-party computation are data provider A, data provider B, and data provider C corresponding to data nodes 102, 103, and 104, respectively, then a multi-party computation can be constructed as follows: Figure 3 The knowledge graph shown is used to describe the data collaboration relationships of the multi-party computation.
[0083] A knowledge graph is a semantic network used to reveal the relationships between entities, and can be used to assist in data analysis and decision-making. A knowledge graph typically includes several vertices and several edges; each node can represent a type of entity, and each edge can represent the relationship between two types of entities.
[0084] Continuing with the examples shown above, further illustrations will be provided in the following cases: Figure 3 The knowledge graph shown may include a master node 301 corresponding to the initiator of the multi-party computation, slave nodes 302, 303, and 304 corresponding to data providers A, B, and C of the multi-party computation, respectively, and edges connecting each slave node to the master node 301. Each edge can be used to represent the data collaboration relationship of the multi-party computation.
[0085] It should be noted that the specific implementation method for constructing the knowledge graph in step 202 is not particularly limited in this specification; for example, those skilled in the art can construct a knowledge graph to describe the data collaboration relationship of the multi-party computation through knowledge representation, knowledge fusion, knowledge reasoning and other methods.
[0086] The knowledge representation mentioned here refers to representing the semantic information of entities as dense low-dimensional vectors, that is, vectors obtained through encoding. This allows for efficient computation of entities, relations, and their complex semantic relationships in a low-dimensional space, which is of great significance for the construction, reasoning, fusion, and application of knowledge graphs. The knowledge fusion mentioned here refers to integrating, disambiguating, processing, reasoning, and updating heterogeneous data from different sources under the same framework and specification. This achieves the fusion of data, information, methods, experience, and human thought, contributing to the formation of a high-quality knowledge base.
[0087] In one embodiment, during the construction of the knowledge graph, each node in the knowledge graph can be characterized based on the identity information of each data provider and the data characteristics of the user data provided by each data provider. Furthermore, each edge in the knowledge graph can be characterized based on the computational process of the multi-party computation. In implementation, in step 202, the data content corresponding to each slave node in the knowledge graph may include the identity identifiers of each data provider corresponding to each slave node, and the data characteristics of the user data provided by each data provider. The data content corresponding to each edge in the knowledge graph may include a sequence of operator identifiers corresponding to the multi-party computation. The operator identifier sequence includes operator identifiers of at least one operator concatenated according to the computational order specified by the initiator.
[0088] The operator can be a packaged basic functional unit; the operator can have several input ports and output ports. In practical applications, based on the input data obtained from the input port of a certain operator and the operator parameter configuration information pre-configured for that operator, a pre-configured data analysis and calculation process can be executed, and the expected output data can be obtained from the output port of that operator.
[0089] It should be noted that the initiator of the multi-party computation can adopt a no-code orchestration approach, that is, without being aware of the internal processing of different operators, it can directly select one or more operators that meet the data usage requirements from different preset operator types, and specify the execution order for each selected operator to complete the multi-party computation process orchestration, thereby efficiently building a privacy computing application that meets the data usage requirements.
[0090] For example, in such Figure 3In the knowledge graph shown, the data content corresponding to node 302 can be "identity identifier of data provider A + encrypted data of user data provided by data provider A", and the data content corresponding to the edge connecting node 302 and master node 301 can be "[operator 1, ..., operator n]"; where "[operator 1, ..., operator n]" is the operator identifier sequence corresponding to the multi-party computation initiated by the data sharing platform, and "operator n" is the operator identifier of the nth operator included in the operator identifier sequence, where n is a positive integer; similarly, the data content corresponding to nodes 303 and 304, as well as the data content corresponding to the edge connecting nodes 303 and 304 and master node 301, can be determined, and will not be elaborated here.
[0091] It should be noted that, in cases such as Figure 3 The knowledge graph shown is merely an example of a multi-party computation relationship and does not imply any special limitation on this specification. In practical applications, depending on different multi-party computation processes or other privacy-preserving computation processes, the edges in the knowledge graph can also be used to represent data collaboration relationships corresponding to other computation processes.
[0092] In the embodiments shown above, the identity identifier of the data provider may specifically include the DID (Decentralized Identity) obtained by the data provider through registration.
[0093] In the embodiments shown above, the data features of the user data may specifically include any one of the following: the encrypted data corresponding to the user data, the hash value of the user data, and the data attribute information of the user data.
[0094] The encrypted data corresponding to the user data can be encrypted data obtained by performing encryption calculations on the original data held by the data provider; the specific encryption method of the encrypted data is not limited in this specification.
[0095] The user's hash value can be obtained by hashing the original data held by the data provider; this specification does not limit the specific hash algorithm.
[0096] The data attribute information of the user data is also the data description information of the user data; for example, the data attribute information can be a data directory.
[0097] Step 204: Perform graph computation on the constructed knowledge graph to solve for the weight value corresponding to each edge in the knowledge graph; wherein, the weight value corresponding to each edge is used to indicate the tightness of the data collaboration relationship represented by each edge.
[0098] For example, in constructing such Figure 3 Following the knowledge graph shown, graph computation can be performed on the knowledge graph to solve for the weight values data_A, data_B, and data_C corresponding to each edge connecting slave node 302, slave node 303, slave node 304 to master node 301. Thus, based on the weight values corresponding to each edge, the tightness of the data collaboration relationship represented by each edge can be determined.
[0099] The larger the weight value, the closer the data collaboration relationship represented by the edge corresponding to that weight value.
[0100] For example, the weight value data_B corresponding to the edge connecting slave node 303 to master node 301 can be used to indicate the degree of data collaboration between data provider B corresponding to slave node 303 and the multi-party computation; similarly, the weight value data_C corresponding to the edge connecting slave node 304 to master node 301 can be used to indicate the degree of data collaboration between data provider C corresponding to slave node 304 and the multi-party computation; if the weight value data_B is greater than the weight value data_C, it indicates that the data provider B has a closer data collaboration relationship than the data provider C in the multi-party computation.
[0101] Graph processing refers to the process of modeling data according to a graph data structure and performing computational analysis and data mining. The basic data structure can be expressed as G = (V, E, D), which is (Vertex, Edge, Weight). It should be noted that in the embodiments shown above, the specific graph algorithms used for graph processing can be flexibly configured by those skilled in the art according to their needs, and this specification does not impose any limitations.
[0102] In one possible embodiment, in step 204, performing graph computation on the constructed knowledge graph to solve for the weight values corresponding to each edge in the knowledge graph may specifically include: inputting the knowledge graph into a trained graph neural network for computation to obtain the weight values corresponding to each edge in the knowledge graph output by the graph neural network.
[0103] Step 206: Based on the weight values corresponding to each edge, determine at least one data provider that meets the preset conditions from the plurality of data providers, and perform the multi-party calculation based on the user data provided by the at least one data provider.
[0104] For example, in constructing such Figure 3 The knowledge graph shown is used to calculate the weight values (data_A, data_B, and data_C) of each edge connecting slave node 302, slave node 303, slave node 304, and master node 301. Based on the weight values of each edge, at least one data provider that meets the preset conditions can be determined from data provider A, data provider B, and data provider C. If at least one data provider that meets the preset conditions is determined to be data provider A and data provider B, the multi-party computation can be performed based on the user data provided by data provider A and data provider B.
[0105] In one embodiment shown, in step 206, determining at least one data provider that meets a preset condition from the plurality of data providers based on the weight values corresponding to each edge may specifically include: determining at least one data provider from the plurality of data providers whose weight value corresponding to the edge between the corresponding slave node and the master node is greater than a preset threshold as a data provider that meets the preset condition.
[0106] For example, in constructing such Figure 3 The knowledge graph shown is used to calculate the weight values (data_A, data_B, and data_C) of each edge connecting slave node 302, slave node 303, and slave node 304 to master node 301. Then, from data provider A, data provider B, and data provider C, at least one data provider whose weight value of the edge between the corresponding slave node and master node 301 is greater than a preset threshold is determined as a data provider that meets the preset conditions. That is, since data_A is greater than the preset threshold, data_B is greater than the preset threshold, and data_C is not greater than the preset threshold, the multi-party computation can be performed based on the user data provided by data provider A and data provider B.
[0107] In another embodiment shown, in step 206, determining at least one data provider that meets the preset conditions from the plurality of data providers based on the weight values corresponding to each edge may specifically include: determining a preset number of data providers with the largest weight values corresponding to the edges between the slave node and the master node as data providers that meet the preset conditions.
[0108] For example, in constructing such Figure 3 The knowledge graph shown is used to calculate the weight values (data_A, data_B, and data_C) of each edge connecting slave node 302, slave node 303, and slave node 304 to master node 301. Then, from data provider A, data provider B, and data provider C, the data provider with the largest number of weight values corresponding to the edges between the slave nodes and master node 301 is determined as the data provider that meets the preset condition. That is, if data_A>data_B>data_C, and the preset number is 2, then the multi-party computation can be performed based on the user data provided by data provider A and data provider B.
[0109] In one embodiment shown, the at least one data provider corresponds to a different data domain; the multi-party computation may include multi-party secure computation performed based on ciphertext data corresponding to user data provided by the at least one data provider.
[0110] In practical applications, to ensure the security of the private raw data of each data provider, each data provider can typically correspond to a different data domain, while data servers within the same data domain can share the same domain name access address. Since secure multi-party computation scenarios should require that private data not leave the domain, secret input data corresponding to the raw user data provided by each data provider can be generated, and secure multi-party computation can be performed based on this secret input data.
[0111] In this context, step 206, where the multi-party computation is performed based on user data provided by the at least one data provider, may specifically include: performing secure multi-party computation based on encrypted data transferred across domains by the at least one data provider to the service device corresponding to the initiator.
[0112] For example, after determining at least one data provider that meets the preset conditions from data provider A, data provider B, and data provider C, the encrypted data corresponding to the user data provided by data provider A and the encrypted data corresponding to the user data provided by data provider B can be transferred across domains to the data sharing platform, and then the data sharing platform can perform the multi-party computation based on the encrypted data transferred across domains.
[0113] It should be noted that the specific implementation method of cross-domain data transfer in the embodiments shown above is not particularly limited in this specification. For example, the data sharing platform can realize cross-domain transfer of encrypted data through HTTP calls.
[0114] As can be seen from the above technical solutions, by constructing a knowledge graph to describe the data collaboration relationships in multi-party computation, and performing graph computation on the constructed knowledge graph, the weight values corresponding to each edge in the knowledge graph can be solved. Since each edge in the knowledge graph can represent the data collaboration relationships in the multi-party computation, and the weight values corresponding to each edge can indicate the closeness of the data collaboration relationships represented by each edge, based on the solved weight values of each edge, at least one data provider that meets preset conditions can be determined from multiple data providers corresponding to multiple slave nodes connected to the master node in the knowledge graph to participate in the multi-party computation.
[0115] In the above manner, in multi-party computation scenarios, the service device corresponding to the initiator can automatically identify the data providers with closer data collaboration relationships from multiple data providers, and perform the multi-party computation based on the user data provided by the data providers. This allows for the selection of input data for the multi-party computation that better meets the initiator's needs, while avoiding the leakage of private user data from each data provider, thus improving the user's multi-party data collaboration experience.
[0116] Corresponding to the above embodiments of the graph-based multi-party computation method, this specification also provides an embodiment of a graph-based multi-party computation apparatus.
[0117] Please see Figure 4 , Figure 4 This is an exemplary embodiment illustrating the hardware structure of an electronic device housing a graph-based multi-party computing device. At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, memory 408, and non-volatile memory 410, and may also include other hardware required for various services. One or more embodiments of this specification can be implemented in software, for example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into memory 408 and then runs it. Of course, besides software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution entity of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0118] Please see Figure 5 , Figure 5 This is a block diagram illustrating an exemplary embodiment of a graph-based multi-party computing device. This graph-based multi-party computing device can be applied to, for example... Figure 4 The electronic device shown implements the technical solution of this specification. The graph-based multi-party computing device may include:
[0119] Construction unit 502 is used to construct a knowledge graph describing the data collaboration relationships of the multi-party computation; wherein, the knowledge graph includes a master node corresponding to the initiator of the multi-party computation, each slave node corresponding to each data provider among the multiple data providers of the multi-party computation, and each edge connecting the slave nodes to the master node; each edge represents the data collaboration relationship of the multi-party computation; each data provider has registered on the service device corresponding to the initiator;
[0120] Graph computation unit 504 is used to perform graph computation on the constructed knowledge graph to solve for the weight value corresponding to each edge in the knowledge graph; wherein, the weight value corresponding to each edge is used to indicate the tightness of the data collaboration relationship represented by each edge;
[0121] The multi-party calculation unit 506 is used to determine at least one data provider that meets preset conditions from the plurality of data providers based on the weight values corresponding to each edge, and to perform the multi-party calculation based on the user data provided by the at least one data provider.
[0122] In this embodiment, the data content corresponding to each slave node in the knowledge graph includes the identity identifier of each data provider corresponding to each slave node, and the data characteristics of the user data provided by each data provider;
[0123] The data content corresponding to each edge in the knowledge graph includes an operator identifier sequence corresponding to the multi-party computation; wherein, the operator identifier sequence includes operator identifiers of at least one operator concatenated according to the computation order specified by the initiator.
[0124] In this embodiment, the data characteristics of the user data include:
[0125] The encrypted data corresponding to the user data; or...
[0126] The hash value of the user data; or,
[0127] The data attribute information of the user data.
[0128] In this embodiment, the multi-party calculation unit 506 is specifically used for:
[0129] From the plurality of data providers, at least one data provider whose weight value corresponding to the edge between the slave node and the master node is greater than a preset threshold is determined as a data provider that meets the preset condition; or...
[0130] From the plurality of data providers, the data providers with the largest number of weight values corresponding to the edges between the slave nodes and the master nodes are determined as data providers that meet the preset conditions.
[0131] In this embodiment, the at least one data provider corresponds to a different data domain; the multi-party computation includes multi-party secure computation performed based on the encrypted data corresponding to the user data provided by the at least one data provider.
[0132] The multi-party computation unit 506 is specifically used for:
[0133] Multi-party secure computation is performed based on encrypted data transferred across domains by at least one data provider to the service device corresponding to the initiator.
[0134] In this embodiment, the initiator includes:
[0135] A data sharing platform corresponding to the aforementioned multiple data providers; or,
[0136] In a data collaboration network that includes the aforementioned multiple data providers, any data provider has multi-party computing needs.
[0137] In this embodiment, the data sharing platform includes a blockchain service platform; the blockchain nodes in the blockchain include service devices corresponding to the multiple data providers respectively;
[0138] The user data provided by each data provider is stored locally on the service device corresponding to each data provider; the data attribute information of the user data provided by each data provider is stored on the blockchain.
[0139] In this embodiment, the blockchain service platform includes a blockchain cloud service platform;
[0140] The service devices corresponding to each data provider include virtual service devices created for each data provider on the blockchain cloud service platform.
[0141] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0142] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0143] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.
[0144] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0145] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0146] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0147] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0148] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0149] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0150] It should be understood that although the terms first, second, third, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of one or more embodiments of this specification, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "in response to a determination," or "when," or "in the event of a determination."
[0151] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit the scope of one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the scope of protection of one or more embodiments of this specification.
Claims
1. A graph-based multi-party computation method, the method being applied to a service device corresponding to the initiator of the multi-party computation; the method comprising: A knowledge graph is constructed to describe the data collaboration relationships of the multi-party computation. The knowledge graph includes a master node corresponding to the initiator of the multi-party computation, slave nodes corresponding to each data provider among the multiple data providers of the multi-party computation, and edges connecting the slave nodes to the master node. Each edge represents a data collaboration relationship in the multi-party computation. Each data provider is registered on a service device corresponding to the initiator. The data content corresponding to each slave node in the knowledge graph includes data features of user data provided by each data provider. The data content corresponding to each edge in the knowledge graph includes operator identifiers for operators corresponding to the multi-party computation. Graph computation is performed on the constructed knowledge graph to solve for the weight value corresponding to each edge in the knowledge graph; wherein, the weight value corresponding to each edge is used to indicate the tightness of the data collaboration relationship represented by each edge; Based on the weight values corresponding to each edge, at least one data provider that meets the preset conditions is determined from the plurality of data providers, and the multi-party calculation is performed based on the user data provided by the at least one data provider.
2. The method according to claim 1, wherein the data content corresponding to each slave node in the knowledge graph includes the identity identifier of each data provider corresponding to each slave node, and the data characteristics of the user data provided by each data provider; The data content corresponding to each edge in the knowledge graph includes a sequence of operator identifiers corresponding to the multi-party computation; wherein... The operator identifier sequence includes operator identifiers of at least one operator concatenated according to the calculation order specified by the initiator.
3. The method according to claim 2, wherein the data characteristics of the user data include: The encrypted data corresponding to the user data; or, The hash value of the user data; or, The data attribute information of the user data.
4. The method according to claim 1, wherein determining at least one data provider satisfying a preset condition from the plurality of data providers based on the weight values corresponding to each edge includes: From the plurality of data providers, at least one data provider whose weight value corresponding to the edge between the slave node and the master node is greater than a preset threshold is determined as a data provider that meets the preset conditions; or, From the plurality of data providers, the data providers with the largest number of weight values corresponding to the edges between the slave nodes and the master nodes are determined as data providers that meet the preset conditions.
5. The method according to claim 1, wherein the at least one data provider corresponds to a different data domain; the multi-party computation includes multi-party secure computation performed based on ciphertext data corresponding to user data provided by the at least one data provider; The multi-party computation based on user data provided by the at least one data provider includes: Multi-party secure computation is performed based on encrypted data transferred across domains by at least one data provider to the service device corresponding to the initiator.
6. The method according to claim 1, wherein the initiator comprises: A data sharing platform corresponding to the aforementioned multiple data providers; or, In a data collaboration network that includes the aforementioned multiple data providers, any data provider has multi-party computing needs.
7. The method according to claim 6, wherein the data sharing platform includes a blockchain service platform; and the blockchain nodes in the blockchain include service devices corresponding to the plurality of data providers respectively; The user data provided by each data provider is stored locally on the service device corresponding to each data provider; the data attribute information of the user data provided by each data provider is stored on the blockchain.
8. The method according to claim 7, wherein the blockchain service platform includes a blockchain cloud service platform; The service devices corresponding to each data provider include virtual service devices created for each data provider on the blockchain cloud service platform.
9. A graph-based multi-party computation apparatus, the apparatus being applied to a service device corresponding to the initiator of the multi-party computation; the apparatus comprising: A construction unit is used to construct a knowledge graph describing the data collaboration relationships of the multi-party computation. The knowledge graph includes a master node corresponding to the initiator of the multi-party computation, slave nodes corresponding to each data provider among the multiple data providers of the multi-party computation, and edges connecting the slave nodes to the master node. Each edge represents a data collaboration relationship in the multi-party computation. Each data provider is registered on a service device corresponding to the initiator. The data content corresponding to each slave node in the knowledge graph includes data features of user data provided by each data provider. The data content corresponding to each edge in the knowledge graph includes operator identifiers corresponding to the operators in the multi-party computation. The graph computation unit is used to perform graph computation on the constructed knowledge graph to solve for the weight value corresponding to each edge in the knowledge graph; wherein, the weight value corresponding to each edge is used to indicate the tightness of the data collaboration relationship represented by each edge; A multi-party computation unit is used to determine at least one data provider that meets preset conditions from the plurality of data providers based on the weight values corresponding to each edge, and to perform the multi-party computation based on the user data provided by the at least one data provider.
10. An electronic device, comprising a communication interface, a processor, a memory, and a bus, wherein the communication interface, the processor, and the memory are interconnected via the bus; The memory stores machine-readable instructions, and the processor executes the method according to any one of claims 1 to 8 by invoking the machine-readable instructions.
11. A machine-readable storage medium storing machine-readable instructions that, when invoked and executed by a processor, implement the method of any one of claims 1 to 8.
Citation Information
Patent Citations
A group division method and device based on a knowledge graph
CN109710599A
Data sharing method, device and system based on consortium blockchain, electronic equipment and computer readable storage medium
CN112235360A